Video restoration method and device, related equipment and computer readable storage medium
By performing frame division and feature extraction on video, and combining with recurrent networks for detection and repair, the problem of difficulty in detecting video frame-level content in the prior art is solved, and more accurate video integrity detection and fluency repair are achieved.
Patent Information
- Application Number
- CN202411998081.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately detect whether there is a lack of frame-level content in the video, especially when the video playback time remains unchanged.
By partitioning the video to be detected, the pixel difference between each two adjacent frame images is extracted, and the target feature extraction network and recurrent network are used for feature extraction and mapping to determine the detection result of the video. If content is detected missing, a supplementary frame image is generated for repair.
Improve accurate detection of video integrity, accurately identify missing content between frames in the video, and repair it by generating supplementary frame images to ensure the smoothness of the repaired video.
Smart Images

Figure CN119941578A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video restoration technology, and in particular to a video restoration method, apparatus, related equipment and computer-readable storage medium. Background Art
[0002] With the development of Internet technology, more and more users will obtain videos online. Since the process of obtaining videos online may be affected by network fluctuations and other factors, the obtained videos may be incomplete. Currently, relevant technologies often determine whether the received video is complete by comparing the playback time of the video in the video information with the playback time of the received video. Although this method can detect whether there is any missing content in the video, when certain frames in the video are missing, it will not affect the playback time of the entire video. Therefore, it is difficult to detect whether the video is complete in the above way. Summary of the invention
[0003] In order to solve the technical problems existing in the related art, the embodiments of the present application provide a video repair method, apparatus, related equipment and computer-readable storage medium.
[0004] To achieve the above purpose, the technical solution of the embodiment of the present application is implemented as follows:
[0005] On the one hand, an embodiment of the present application provides a video repair method, the method comprising:
[0006] Dividing the video to be detected into frames to obtain a frame sequence of the video to be detected;
[0007] Determine the pixel difference between every two adjacent frame images in the frame sequence, and perform feature extraction on the pixel difference through a target feature extraction network to obtain a difference feature between every two adjacent frame images;
[0008] According to the correspondence between the frame sequence and the frame image, feature mapping is performed on the difference features between every two adjacent frame images through a cyclic network in turn to obtain mapping features corresponding to each of the difference features, and the detection result of the video to be detected is determined based on the mapping features corresponding to each of the difference features;
[0009] If the detection result indicates that there is content missing in the video to be detected, a supplementary frame image corresponding to the content missing is generated, and the video to be detected is repaired based on the supplementary frame image to obtain a target video.
[0010] On the other hand, an embodiment of the present application provides a video repair device, including:
[0011] An acquisition module, used for dividing the video to be detected into frames to obtain a frame sequence of the video to be detected;
[0012] A difference determination module is used to determine the pixel difference between every two adjacent frame images in the frame sequence, and perform feature extraction on the pixel difference through a target feature extraction network to obtain a difference feature between every two adjacent frame images;
[0013] A feature mapping module, used to perform feature mapping on the difference features between every two adjacent frame images in turn through a recurrent network according to the correspondence between the frame sequence and the frame image, obtain mapping features corresponding to each of the difference features, and determine the detection result of the video to be detected based on the mapping features corresponding to each of the difference features;
[0014] The video repair module is used to generate a supplementary frame image corresponding to the content missing if the detection result indicates that the video to be detected has content missing, and repair the video to be detected based on the supplementary frame image to obtain a target video.
[0015] On the other hand, an embodiment of the present application further provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps in the above method when running the computer program.
[0016] On the other hand, an embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.
[0017] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, which implements the steps in the above method when executed by a processor.
[0018] The video restoration method, apparatus, related equipment and computer-readable storage medium provided in the embodiments of the present application divide the video to be detected into frames to obtain a frame sequence of the video to be detected. By dividing the complete video into multiple frames of continuous images, each frame image is detected separately in the subsequent process, so as to accurately determine whether the video has any content missing, thereby improving the accuracy of detecting the integrity of the video. Then, the pixel difference between every two adjacent frame images in the frame sequence is determined, and feature extraction is performed on the pixel difference through a target feature extraction network to obtain the difference feature between every two adjacent frame images. According to the correspondence between the frame sequence and the frame image, feature mapping is performed on the difference feature between every two adjacent frame images through a cyclic network in turn to obtain the mapping feature corresponding to each difference feature, and based on the mapping feature corresponding to each difference feature The detection result of the video to be detected is determined by the projection features. Since the frame images in the video are continuous, the pixel differences between the consecutive frames can be used to determine whether the current frame image and the next frame image are continuous, and then determine whether there is content missing between the two frame images. By verifying whether there is content missing between any two frames in the video, it is determined whether there is content missing in the entire video, thereby improving the accuracy of detecting the integrity of the video; if the detection result indicates that there is content missing in the video to be detected, a supplementary frame image corresponding to the content missing is generated, and the video to be detected is repaired based on the supplementary frame image to obtain the target video. By generating a supplementary frame image, the video with content missing is repaired, which can make the previous and next frames of the repaired video correlated, thereby improving the fluency of the repaired video. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of a video repair method according to an embodiment of the present application;
[0020] Figure 2 It is a schematic diagram of the implementation process of video restoration provided by the embodiment of the present application;
[0021] Figure 3 It is a schematic diagram of an improved process of a neural network model provided in an embodiment of the present application;
[0022] Figure 4 A schematic diagram of the structure of a video restoration device according to an embodiment of the present application;
[0023] Figure 5 Schematic diagram of the hardware structure of the embodiment of the present application. DETAILED DESCRIPTION
[0024] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0026] When developing and testing functions related to video-on-demand products, a large number of video files are usually required to verify the video-related functions. Generally speaking, most of our test data files are directly downloaded from the Internet. When downloading, we cannot determine whether the downloaded videos are complete and available. We can only wait until the function has problems and then check whether it is a problem with the test video source file or a product function problem, which consumes a lot of time. Therefore, this solution discloses a test video integrity detection and repair method.
[0027] Video integrity means that the file structure of the video does not change during the transmission, storage, and copying process, and the files before and after the transmission can be used normally for video testing. Currently, most application scenarios for verifying video integrity are video transmission, verifying whether the video is complete before and after the video transmission, including:
[0028] (1) Video detection based on hash functions: by comparing the message digest algorithm (MD5) of the files before and after transmission, it is determined whether the transmitted video is complete. If the MD5 value of the video after transmission is consistent with that before transmission, the video is considered complete.
[0029] (2) Detection method based on video metadata: By reading the metadata information of the video, the duration information in the metadata is compared with the actual video duration to see if they are consistent. If they are consistent, the video is considered complete and can be played normally. If they are inconsistent, the video is considered incomplete.
[0030] In actual test application scenarios, there are many ways to obtain test videos, such as splicing video files of different formats, downloading directly from the Internet, etc. This type of acquisition method generally cannot know the source of the file, and it is impossible to verify whether the file video is complete and available by comparing the MD5 value. Therefore, verifying whether the file is complete by comparing the MD5 value of the file is not applicable to test files whose video source cannot be determined. The above-mentioned detection method based on video metadata determines whether the video is complete by comparing the video length in the file metadata information with the actual video length. However, in actual tests, it is found that when a frame of the video is damaged, it does not affect the length of the video.
[0031] In addition, the detection method based on hash functions and the detection method based on video metadata both simply verify whether the video is complete. However, in actual test scenarios, verifying whether the video is complete is not the ultimate goal of testing the video integrity. The ultimate goal is to repair the incomplete video so that it can support the test normally.
[0032] Based on this, an embodiment of the present application proposes a video repair method. In various embodiments of the present application, the video to be detected is divided into frames to obtain a frame sequence of the video to be detected. By dividing the complete video into multiple frames of continuous images and subsequently detecting each frame image separately, it is possible to accurately determine whether the video has content missing, thereby improving the accuracy of detecting video integrity; then, the pixel difference between every two adjacent frame images in the frame sequence is determined, and feature extraction is performed on the pixel difference through a target feature extraction network to obtain the difference feature between every two adjacent frame images; according to the correspondence between the frame sequence and the frame image, feature mapping is performed on the difference feature between every two adjacent frame images in turn through a cyclic network to obtain the mapping feature corresponding to each difference feature, and based on the mapping feature corresponding to each difference feature The features determine the detection result of the video to be detected. Since the frame images in the video are continuous, the pixel differences between the consecutive frames can be used to determine whether the current frame image and the next frame image are continuous, and then determine whether there is content missing between the two frame images. By verifying whether there is content missing between any two frames in the video, it is determined whether there is content missing in the entire video, thereby improving the accuracy of detecting the integrity of the video; if the detection result indicates that there is content missing in the video to be detected, a supplementary frame image corresponding to the content missing is generated, and the video to be detected is repaired based on the supplementary frame image to obtain the target video. By generating a supplementary frame image, the video with content missing is repaired, which can make the previous and next frames of the repaired video correlated, thereby improving the fluency of the repaired video.
[0033] The present application embodiment provides a video repair method, Figure 1 FIG. 1 is a flow chart of a video repair method according to an embodiment of the present application; Figure 1 As shown, the method includes:
[0034] Step 101: Divide the video to be detected into frames to obtain a frame sequence of the video to be detected.
[0035] In actual implementation, the video to be detected may be obtained by downloading the video to be detected through the Internet, or by reading the video to be detected stored in a storage medium.
[0036] In actual implementation, the process of dividing the video to be detected into frames can be the process of dividing the video into multiple frames of images, which usually involves video decoding and frame extraction. First, the file of the video to be detected needs to be read, and the video file can be opened using video processing software; then the video to be detected is decoded, and the file of the specific video to be detected usually contains compressed data, and the original image frame data needs to be obtained by decoding the video to be detected; then the frame rate of the video to be detected (for example, 30 frames / second) can be determined. The frame rate of the video defines how many frames of images there are per second. This information can usually be obtained from the header information of the video file; then the video frame of the video to be detected can be extracted, specifically, the video stream can be read frame by frame according to the frame rate of the video, and each time a video frame is read, it is extracted and saved. In some cases, it may be necessary to extract the video frame at a specific interval, such as extracting once every one or several frames, and finally the extracted video frames are sorted according to the position of the video frame in the video to be detected to obtain a frame sequence of the video to be detected.
[0037] Step 102: Determine the pixel difference between every two adjacent frame images in the frame sequence, and perform feature extraction on the pixel difference through a target feature extraction network to obtain the difference feature between every two adjacent frame images.
[0038] In some embodiments, determining the pixel difference between each two adjacent frame images in the frame sequence in step 102 can be achieved by the following technical solution: respectively determining the color value of each pixel point included in the two adjacent frame images; determining the absolute value of the color value difference of each group of target pixel points as the pixel difference, wherein each group of target pixel points includes two pixel points with the same coordinates in the two adjacent frame images.
[0039] In actual implementation, the color value of each pixel refers to the numerical value that represents the color information of the pixel in a digital image. For example, in the red, green and blue model, the color of each pixel is composed of three independent color channels: red, green and blue. Each channel is usually represented by an 8-bit unsigned integer with a value range from 0 (no color intensity) to 255 (maximum color intensity).
[0040] In actual implementation, the pixel difference is determined by the following formula (1):
[0041] F n =|f n+1 -f n | (1)
[0042] In formula (1), F n is the pixel difference, f n+1 The color value of the next frame image, f n It is the color value of the current frame image.
[0043] In some embodiments, before executing step 102 of extracting features of pixel differences through a target feature extraction network to obtain the difference features between every two adjacent frame images, the following technical solutions may also be executed: determining a first feature network, and performing network expansion on the initial layer of the first feature network to obtain a target initial layer; creating a first convolutional layer, the size of the first convolutional layer is smaller than the size of the target initial layer; and processing the first convolutional layer and the target initial layer in parallel to obtain a target feature extraction network.
[0044] In actual implementation, the first feature network for feature extraction of pixel differences can be a visual geometry group 16-layer network model (VGG16 model). In this model, the size of the initial layer of the first feature network is 5x5. In order to extract more feature information, the initial layer of the first feature network can be expanded to obtain a target initial layer with a size of 7x7. Since the larger the convolution kernel and the deeper the number of layers, the corresponding number of parameters will increase, which will increase the computational complexity of the network. In order to avoid this phenomenon, a new first convolution layer can be created after the target initial layer, and the first convolution layer and the target initial layer can be connected in parallel. The size of the first convolution layer can be 3x3.
[0045] Step 103: According to the correspondence between the frame sequence and the frame image, feature mapping is performed on the difference features between every two adjacent frame images through a cyclic network in turn to obtain mapping features corresponding to each difference feature, and the detection result of the video to be detected is determined based on the mapping features corresponding to each difference feature.
[0046] In actual implementation, the correspondence between the frame sequence and the frame image can be the position of the frame image in the frame sequence. For example, if the frame sequence is frame image A, frame image B and frame image C, frame image A can be input into the recurrent network first, then frame image B can be input into the recurrent network, and finally frame image C can be input into the recurrent network.
[0047] In some embodiments, the step 103 of sequentially performing feature mapping on the difference features between every two adjacent frame images through a recurrent network to obtain mapping features corresponding to each difference feature can be implemented by the following technical solution: for the difference features between every two adjacent frame images, performing the following processing: inputting the difference features into the initial layer of the recurrent network to obtain a second feature of the output of the initial layer; sequentially performing feature mapping on the second feature through multiple hidden layers included in the recurrent network to obtain a third feature; inputting the third feature into the output layer of the recurrent network to obtain a mapping feature output by the output layer.
[0048] In actual implementation, the recurrent network may include multiple hidden layers, each of which may include one or more recurrent units that can remember previous information in a sequence. In a stacked recurrent layer structure, multiple recurrent layers are stacked together, and each layer receives the output from the layer below as input. The obtained difference feature can be input into the initial layer of the recurrent network to obtain the first feature, and then the first feature can be input into the hidden layer according to the sorting order of the hidden layers in the recurrent network, and the output of the next hidden layer is used as the input of the next hidden layer, until the last hidden layer of the recurrent network inputs the third feature into the output layer of the recurrent network, and the output layer of the recurrent network outputs the mapping feature.
[0049] In some embodiments, determining the detection result of the video to be detected based on the mapping features corresponding to each difference feature in step 103 can be achieved by the following technical solution: performing feature fusion on the mapping features corresponding to each difference feature to obtain the target feature; performing feature mapping on the target feature to obtain the mapping probability; if the mapping probability is greater than or equal to the probability threshold, the detection result of the video to be detected is that there is no content missing in the video to be detected; if the mapping probability is less than the probability threshold, the detection result of the video to be detected is that there is content missing in the video to be detected.
[0050] In actual implementation, after each difference feature passes through the recurrent network, a mapping feature output by the recurrent network can be obtained. Finally, multiple mapping features can be fused to obtain the target feature. The process of fusing the mapping features can be feature splicing of the mapping features, and the target feature can be obtained by splicing the mapping features.
[0051] In actual implementation, the process of feature mapping the target feature can be to input the target feature into a pre-trained discriminator, and the discriminator outputs a mapping probability. When the mapping probability output by the discriminator is greater than or equal to a probability threshold, it indicates that there is no content missing in the video to be detected. Conversely, if the mapping probability output by the discriminator is less than the probability threshold, it indicates that the storage content of the video to be detected is missing.
[0052] In actual implementation, the process of training the discriminator can be to obtain video samples consisting of multiple complete video samples and content-missing video samples and to label these video samples, input the video samples into the discriminator to be trained, and the discriminator outputs the probability that each video sample is a complete video. When the probability output by the discriminator is greater than or equal to 0.5, it is determined that the result predicted by the discriminator is that the video sample is a complete video. When the probability output by the discriminator is less than 0.5, it is determined that the result predicted by the discriminator is that the video sample is a content-missing video. A loss function is constructed based on the result predicted by the discriminator and the labeling of the video samples, and the discriminator is trained, wherein the constructed loss function can be a cosine loss function, etc.
[0053] Step 104: If the detection result indicates that the video to be detected has missing content, a supplementary frame image corresponding to the missing content is generated, and the video to be detected is repaired based on the supplementary frame image to obtain a target video.
[0054] In some embodiments, the generation of the supplementary frame image corresponding to the missing content in step 104 can be achieved through the following technical solution: obtaining a noise image, inputting the noise image into an image generation network, and having the image generation network denoise the noise image to obtain a supplementary frame image output by the image generation network.
[0055] In actual implementation, the generative network usually accepts a random noise vector as input. This noise vector is usually a Gaussian or uniformly distributed random number, and its dimension depends on the specific network architecture and the complexity of the generated image. The noise vector is first fed into the first hidden layer of the generative network. This hidden layer maps the input noise to a higher-dimensional feature space; then, through a series of hidden layers, the noise vector is gradually converted into a more complex feature representation; in each hidden layer of the generative network, a nonlinear activation function (such as a linear rectification function) is usually used to introduce nonlinearity, so that the network can learn and represent more complex data distributions; in the generative network, the noise image is denoised to obtain a denoised feature map; finally, the generated feature map is fed into an output layer, which converts the feature map into image pixel values. These pixel values represent the output of the generative network, that is, the generated supplementary frame image.
[0056] In some embodiments, the step 104 of repairing the video to be detected based on the supplementary frame image to obtain the target video can be achieved by the following technical solution: determining the video frame rate of the target video, the video size of the target video and the resolution of the target video; determining the target playback time of the frame image in the target video based on the video frame rate, the video size and the video resolution; reading each frame image and the supplementary frame image in the video to be detected in turn, and synthesizing each frame image and the supplementary frame image in the video to be detected to obtain the target video, wherein the playback time of the frame image of each frame in the target video is the same as the target playback time.
[0057] In actual implementation, the video information of the video to be detected can be read to determine the video frame rate, video size and resolution of the video to be detected. Then, the video frame rate of the video to be detected can be used as the video frame rate of the target video, the video size of the video to be detected can be used as the video size of the target video, and the resolution of the video to be detected can be used as the resolution of the target video.
[0058] In actual implementation, the video frame rate, video size and video resolution are integrated to obtain the target playback time of each frame image in the target video, which can be seen in the following formula (2):
[0059]
[0060] In formula (2), T is the target playback time, W is the width of the video, H is the height of the video, FPS is the video frame rate, and Px is the video resolution.
[0061] In actual implementation, the generated supplementary frame image can be placed at the position where the content is missing in the video to be detected, and the complete images and supplementary frame images included in the video to be detected can be sorted in sequence according to the determined target playback time to obtain the target video.
[0062] It should be noted that after obtaining the target video, the target video can also be used as a video to be detected. The target video can be detected by the method shown in steps 101 to 103 above to obtain the detection result of the target video. When the detection result of the target video indicates that the target video is a complete video, the target video is displayed to the user.
[0063] The present application is described below in conjunction with application examples.
[0064] See also Figure 2 , Figure 2 It is a schematic diagram of the implementation process of video restoration provided in an embodiment of the present application.
[0065] In step 201, the server prepares and preprocesses the data set.
[0066] Prepare a synthetic test data set of normal videos and damaged videos, and preprocess the video files in the data set. First, convert the video to be tested into a frame sequence, and calculate the frame difference (pixel difference) between the previous and next frames of the video. For details, please refer to the following formula (3):
[0067] F n =|f n+1 -f n | (3)
[0068] In formula (3), F n is the pixel difference, f n+1 The color value of the next frame image, f n It is the color value of the current frame image.
[0069] In step 202, the server builds a detection model.
[0070] The basic framework of the original neural network model is a 16-layer network model of the visual geometry group, which is composed of 5 convolutional layers connected to pooling layers. The convolutional layers are connected using fully connected layers, and finally the activation network is used to classify the model features. Although more feature information can be extracted by stacking convolutions, this will also lead to too many model parameters, large storage capacity, and too long convergence time during training, which is easy to cause overfitting.
[0071] In order to improve the detection accuracy of the algorithm module without increasing the computational complexity of the algorithm, the original neural network model is improved. The improved model structure is as follows: Figure 3 As shown, Figure 3 It is a schematic diagram of the improved process of the neural network model provided in the embodiment of the present application.
[0072] In step 301, the server adjusts the network structure.
[0073] Replace the original 5x5 convolution layer from the input layer to the first convolution layer with a 7x7 network structure to obtain more feature information. At the same time, considering that the larger the convolution kernel and the deeper the number of layers, the more corresponding parameters will be, which will increase the computational complexity of the network. In order to avoid this phenomenon, the solution increases the width of the network model by connecting a 3x3 convolution layer in parallel and adding a lx1 convolution kernel. The 1x1 convolution kernel can increase the accuracy of the algorithm by increasing the model width without increasing the computational complexity of the original algorithm. The general process can be expressed as: Assume that the initial input of the model is l 1 =K*K*C, then the process from the input layer to the convolutional layer can be described as: 2 =ReLu(l 1 *w 2 *b 2 ), where b is the bias matrix, which is used to increase the translation ability of the network, and w is the weight of the network. The 3x3 convolution is followed by a pooling layer to reduce the number of model parameters. The result of the model after the pooling layer PxP can be expressed as: (K / P)*(K / P)*C, where PxP represents the size of the pooling layer. The output of the convolution layer with multiple layers of 3x3 and 1x1 in parallel can be expressed as: l n =ReLu(l n-1 *w n *b n ), where n represents the number of convolutional layers.
[0074] In step 302, the server determines mapping features through a recurrent network.
[0075] For the test video, it is composed of a series of continuous frames. The previous and next frames of the video are related. When processing and analyzing video files, the continuity of the previous and next frames of the video needs to be considered. In view of this, a loop structure (i.e., the above-mentioned loop network) is added after the output layer of the convolutional network layer to connect the input and output of the previous and next frames of the video. The specific steps are: take the output of the convolutional network (i.e., the above-mentioned difference feature) as the input of the loop network, and assume that the output of the convolutional network is the input x(n) of the loop network. Then the result of the loop output of the loop network is: H n =g(V*A n ), x(n)=f(Ux n-1 +wA n-1 ), where V and U represent the convolution matrices of the input layer and hidden layer respectively, and A n represents the output of the hidden layer, A n-1 Represents the output of the previous hidden layer. Finally, the output of the recurrent network (i.e., the mapping feature mentioned above) is concatenated with the sample data obtained through the fully connected layer to obtain the target feature, and then the target feature is input into the discriminator. The discriminator is based on the activation function. The activation function maps the output data of the recurrent network to (0, 1) and sets the judgment threshold: when the output result is less than 0.5, it is false, that is, the video is incomplete, and when it is greater than 0.5, the result is true, that is, the video is complete.
[0076] In step 203, the server generates a supplementary frame image.
[0077] The server uses a pre-trained generative adversarial network model to input the incomplete video frame determined by the video detection into the generative network and input a random noise x(t) (i.e. the noise image mentioned above). After passing through the generative network, the network output y(t) (i.e. the supplementary frame image mentioned above) is obtained.
[0078] In step 204, the server again detects whether the video to be detected with the supplementary frame image added is complete.
[0079] The improved neural network structure is used as an adversarial network to detect whether the generated image is complete and usable. The output of the generated network is used as the input of the above detection model to detect whether the video frame image generated by the generated network is complete. If it is complete, it is output for use. Its use process is mainly to merge a series of detected complete video frame images into a video file. The steps can be decomposed into: ① Calculate the playback time of each image. Let the width and height of the image (that is, the size of the above target video) be W and H respectively: the frame rate of the video (that is, the frame rate of the above target video) is: FPS, the resolution of the image (that is, the resolution of the above target video) is Px, then the playback time T (that is, the above target playback time) can be expressed as: Create a video encoder object, specify the output file format, encoder, audio encoding and other parameters; traverse the image files and read each image; output the merged video file.
[0080] If the result is judged to be incomplete, it is input into the generative network again for training output until a complete video frame image is finally output.
[0081] At present, the traditional methods of video integrity detection based on hash functions or methods based on video meta-information both require clear knowledge of the source file video information of the video being detected, and are not suitable for test video files with various acquisition methods. Compared with the traditional video detection method, the solution provided in the embodiment of the present application applies a deep convolutional network model to test video detection, which can detect the test video without paying attention to the video source file, and is more universal.
[0082] Secondly, the traditional video detection model can only simply detect whether the video is complete or defective, and cannot repair the defective video. In the application scenario of video testing, when the video is judged to be unavailable, the operation of downloading the video and detecting the video will be repeated, wasting a lot of time and cost. The solution provided by the embodiment of the present application combines the detection model with the repair model, which can automatically detect and repair incomplete videos, ensure that the video used for the test is a complete video, and effectively save time and cost.
[0083] Finally, compared with traditional detection models, the detection accuracy of the model is generally improved by increasing the depth of the model. Although this can improve the detection accuracy of the model to a certain extent, it also increases the number of model parameters and the computational complexity of the model. In response to this problem, the solution provided in the embodiment of the present application can improve the detection accuracy of the algorithm without changing the complexity of the model by increasing the width of the model. At the same time, combined with the correlation between the previous and next frames of the video, a loop structure is added to further improve the detection accuracy of the model.
[0084] In order to implement the video repair method on the server side of the embodiment of the present application, the embodiment of the present application also provides a video repair device, Figure 4 FIG. 1 is a schematic diagram of the composition structure of the video repair device according to an embodiment of the present application. Figure 4 As shown, the video repair device comprises:
[0085] An acquisition module 41 is used to divide the video to be detected into frames to obtain a frame sequence of the video to be detected;
[0086] A difference determination module 42 is used to determine the pixel difference between every two adjacent frame images in the frame sequence, and perform feature extraction on the pixel difference through a target feature extraction network to obtain a difference feature between every two adjacent frame images;
[0087] A feature mapping module 43 is used to perform feature mapping on the difference features between every two adjacent frame images through a recurrent network in accordance with the correspondence between the frame sequence and the frame image, obtain mapping features corresponding to each of the difference features, and determine the detection result of the video to be detected based on the mapping features corresponding to each of the difference features;
[0088] The video repair module 44 is used to generate a supplementary frame image corresponding to the content missing if the detection result indicates that the video to be detected has content missing, and repair the video to be detected based on the supplementary frame image to obtain a target video.
[0089] In one embodiment, the difference determination module 42 is further used to respectively determine the color value of each pixel included in the two adjacent frame images; and determine the absolute value of the color value difference of each group of target pixel points as the pixel difference, wherein each group of target pixel points includes two pixel points with the same coordinates in the two adjacent frame images.
[0090] In one embodiment, the difference determination module 42 is also used to determine a first feature network, and perform network expansion on the initial layer of the first feature network to obtain a target initial layer; create a first convolutional layer, the size of the first convolutional layer is smaller than the size of the target initial layer; and perform parallel processing on the first convolutional layer and the target initial layer to obtain the target feature extraction network.
[0091] In one embodiment, the feature mapping module 43 is also used to perform the following processing for the difference feature between every two adjacent frame images: input the difference feature into the initial layer of the recurrent network to obtain a second feature of the output of the initial layer; perform feature mapping on the second feature in turn through the multiple hidden layers included in the recurrent network to obtain a third feature; input the third feature into the output layer of the recurrent network to obtain the mapping feature output by the output layer.
[0092] In one embodiment, the feature mapping module 43 is also used to perform feature fusion on the mapping features corresponding to each difference feature to obtain the target feature; perform feature mapping on the target feature to obtain a mapping probability; if the mapping probability is greater than or equal to a probability threshold, the detection result of the video to be detected is that there is no content missing in the video to be detected; if the mapping probability is less than the probability threshold, the detection result of the video to be detected is that there is content missing in the video to be detected.
[0093] In one embodiment, the video restoration module 44 is further used to obtain a noisy image, input the noisy image into an image generation network, and perform denoising on the noisy image by the image generation network to obtain the supplementary frame image output by the image generation network.
[0094] In one embodiment, the video repair module 44 is also used to determine the video frame rate of the target video, the video size of the target video and the resolution of the target video; based on the video frame rate, the video size and the video resolution, determine the target playback time of the frame image in the target video; read each frame image and the supplementary frame image in the video to be detected in turn, and synthesize each frame image and the supplementary frame image in the video to be detected to obtain the target video, wherein the playback time of the frame image of each frame in the target video is the same as the target playback time.
[0095] It should be noted that: the video repair device provided in the above embodiment only uses the division of the above-mentioned program modules as an example when performing video repair. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above.
[0096] It should be noted that: the video repair device provided in the above embodiment only uses the division of the above program modules as an example when performing video repair. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the video repair device provided in the above embodiment and the video repair method embodiment on the second client side belong to the same concept. The specific implementation process is detailed in the video repair method embodiment, which will not be repeated here.
[0097] Based on the hardware implementation of the above program modules, and in order to implement the video repair method of the embodiment of the present application, the embodiment of the present application also provides a server, Figure 5 The hardware structure diagram of the embodiment of the present application is as follows: Figure 5 As shown, the server 50 includes:
[0098] A first communication interface 51, capable of exchanging information with other devices;
[0099] The first processor 52 is connected to the first communication interface 51 to implement information interaction with other devices and is used to execute the above-mentioned video repair method when running a computer program, and the computer program is stored in the first memory 53.
[0100] Specifically, the first processor 52 is used to perform frame division on the video to be detected to obtain a frame sequence of the video to be detected; determine the pixel difference between every two adjacent frame images in the frame sequence, and perform feature extraction on the pixel difference through a target feature extraction network to obtain a difference feature between every two adjacent frame images; according to the correspondence between the frame sequence and the frame image, perform feature mapping on the difference feature between every two adjacent frame images through a recurrent network in turn to obtain a mapping feature corresponding to each of the difference features, and determine a detection result of the video to be detected based on the mapping feature corresponding to each of the difference features; if the detection result indicates that the video to be detected has content missing, generate a supplementary frame image corresponding to the content missing, and repair the video to be detected based on the supplementary frame image to obtain a target video;
[0101] In one embodiment, the first processor 52 is further used to respectively determine the color value of each pixel included in the two adjacent frame images; and determine the absolute value of the color value difference of each group of target pixel points as the pixel difference, wherein each group of target pixel points includes two pixel points with the same coordinates in the two adjacent frame images.
[0102] In one embodiment, the first processor 52 is further used to determine a first feature network, and perform network expansion on an initial layer of the first feature network to obtain a target initial layer; create a first convolutional layer, the size of the first convolutional layer is smaller than the size of the target initial layer; and perform parallel processing on the first convolutional layer and the target initial layer to obtain the target feature extraction network.
[0103] In one embodiment, the first processor 52 is further used to perform the following processing for the difference feature between every two adjacent frame images: input the difference feature into the initial layer of the recurrent network to obtain a second feature of the output of the initial layer; perform feature mapping on the second feature in turn through the multiple hidden layers included in the recurrent network to obtain a third feature; input the third feature into the output layer of the recurrent network to obtain the mapping feature output by the output layer.
[0104] In one embodiment, the first processor 52 is further used to perform feature fusion on the mapping features corresponding to each difference feature to obtain a target feature; perform feature mapping on the target feature to obtain a mapping probability; if the mapping probability is greater than or equal to a probability threshold, the detection result of the video to be detected is that there is no content missing in the video to be detected; if the mapping probability is less than the probability threshold, the detection result of the video to be detected is that there is content missing in the video to be detected.
[0105] In one embodiment, the first processor 52 is further used to obtain a noisy image, input the noisy image into an image generation network, and perform denoising on the noisy image by the image generation network to obtain the supplementary frame image output by the image generation network.
[0106] In one embodiment, the first processor 52 is further used to determine the video frame rate of the target video, the video size of the target video and the resolution of the target video; determine the target playback time of the frame image in the target video based on the video frame rate, the video size and the video resolution; read each frame image and the supplementary frame image in the video to be detected in sequence, and synthesize each frame image and the supplementary frame image in the video to be detected to obtain the target video, wherein the playback time of the frame image of each frame in the target video is the same as the target playback time.
[0107] It should be noted that the specific processing process of the first communication interface 51 and the first processor 52 can be understood by referring to the above-mentioned video repair method.
[0108] Of course, in actual application, the various components in the server 50 are coupled together through the first bus system 54. It is understandable that the first bus system 54 is used to realize the connection and communication between these components. In addition to the data bus, the first bus system 54 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 In the figure, various buses are labeled as a first bus system 54 .
[0109] The first memory 53 in the embodiment of the present application is used to store various types of data to support the operation of the server 40. Examples of such data include: any computer program used to operate on the server 50.
[0110] The video repair method disclosed in the above embodiment of the present application can be applied to the first processor 52, or implemented by the first processor 52. The first processor 52 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above video repair method can be completed by the hardware integrated logic circuit or software instructions in the first processor 52. The above first processor 52 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The first processor 52 can implement or execute the video repair method, steps and logic block diagram disclosed in the embodiment of the present application. The general processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the video repair method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the first memory 53. The first processor 52 reads the information in the first memory 53 and completes the steps of the above video repair method in combination with its hardware.
[0111] In an exemplary embodiment, the server 50 can be implemented by one or more application-specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field-programmable gate array (FPGA), general-purpose processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned video repair method.
[0112] In an exemplary embodiment, the embodiment of the present application further provides an electronic device, including a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps of any of the above methods when running the computer program.
[0113] The embodiment of the present application further provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a memory 603 storing a computer program, and the computer program can be executed by a processor 602 of an electronic device 600 to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.
[0114] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0115] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. The term "and / or" herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "one or more" herein represents any combination of at least two of any one or more of a plurality of items. For example, including one or more of A, B, and C can represent including any one or at least two or more elements selected from the set consisting of A, B, and C.
[0116] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0117] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A video repair method, characterized in that: The method comprises: Divide the video to be detected into frames to obtain a frame sequence of the video to be detected; Determine the pixel difference between every two adjacent frame images in the frame sequence, and perform feature extraction on the pixel difference through a target feature extraction network to obtain a difference feature between every two adjacent frame images; According to the correspondence between the frame sequence and the frame image, feature mapping is performed on the difference features between every two adjacent frame images through a cyclic network in turn to obtain mapping features corresponding to each of the difference features, and the detection result of the video to be detected is determined based on the mapping features corresponding to each of the difference features; If the detection result indicates that there is content missing in the video to be detected, a supplementary frame image corresponding to the content missing is generated, and the video to be detected is repaired based on the supplementary frame image to obtain a target video.
2. The method according to claim 1, characterized in that The determining of the pixel difference between every two adjacent frame images in the frame sequence comprises: Respectively determining the color value of each pixel point included in the two adjacent frame images; The absolute value of the color value difference of each group of target pixel points is determined as the pixel difference, wherein each group of target pixel points includes two pixel points with the same coordinates in the two adjacent frame images.
3. The method according to claim 1, characterized in that Before extracting features from the pixel differences through a target feature extraction network to obtain difference features between every two adjacent frame images, the method includes: Determine a first characteristic network, and perform network expansion on an initial layer of the first characteristic network to obtain a target initial layer; Creating a first convolutional layer, wherein the size of the first convolutional layer is smaller than the size of the target initial layer; The first convolutional layer and the target initial layer are processed in parallel to obtain the target feature extraction network.
4. The method according to claim 1, characterized in that: The step of sequentially performing feature mapping on the difference features between every two adjacent frame images through a cyclic network to obtain mapping features corresponding to each of the difference features includes: For the difference features between every two adjacent frame images, the following processing is performed: Inputting the difference feature into the initial layer of the cyclic network to obtain a second feature of the output of the initial layer; Performing feature mapping on the second features in sequence through the hidden layer of the recurrent network to obtain a third feature; The third feature is input into the output layer of the cyclic network to obtain the mapping feature.
5. The method according to claim 1, characterized in that The determining the detection result of the video to be detected based on the mapping features corresponding to each of the difference features includes: Performing feature fusion on the mapping features corresponding to each difference feature to obtain the target feature; Performing feature mapping on the target feature to obtain a mapping probability; If the mapping probability is greater than or equal to the probability threshold, determining that there is no detection result of content missing in the video to be detected; If the mapping probability is less than the probability threshold, a detection result indicating that the video to be detected has missing content is determined.
6. The method according to claim 1, characterized in that The generating of the supplementary frame image corresponding to the missing content includes: A noisy image is acquired, and the noisy image is input into an image generation network, and the image generation network performs denoising on the noisy image to obtain the supplementary frame image output by the image generation network.
7. The method according to claim 1, characterized in that The repairing of the to-be-detected video based on the supplementary frame image to obtain a target video includes: Determine the video frame rate of the target video, the video size of the target video, and the resolution of the target video; Determining a target playback time of a frame image in the target video based on the video frame rate, the video size, and the video resolution; Each frame image and the supplementary frame image in the video to be detected are read in sequence, and each frame image and the supplementary frame image in the video to be detected are synthesized to obtain the target video, wherein the playback time of the frame image of each frame in the target video is the same as the target playback time.
8. A video restoration device, characterized in that: include: An acquisition module, used for dividing the video to be detected into frames to obtain a frame sequence of the video to be detected; A difference determination module is used to determine the pixel difference between every two adjacent frame images in the frame sequence, and extract features of the multiple pixel differences through a target feature extraction network to obtain a difference feature between every two adjacent frame images; A feature mapping module, used to perform feature mapping on the difference features between every two adjacent frame images in turn through a recurrent network according to the correspondence between the frame sequence and the frame image, obtain mapping features corresponding to each of the difference features, and determine the detection result of the video to be detected based on the mapping features corresponding to each of the difference features; The video repair module is used to generate a supplementary frame image corresponding to the content missing if the detection result indicates that the video to be detected has content missing, and repair the video to be detected based on the supplementary frame image to obtain a target video.
9. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be executed on the processor, wherein: The processor is used to execute the steps of the method according to any one of claims 1 to 7 when running a computer program.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.