A black and white video colorization method based on neural network and motion information
By using a convolutional neural network in video coloring combined with self-attention mechanism and motion information, the problem of difficult handling of moving objects in video coloring is solved, and a more accurate and stable video coloring effect is achieved.
Patent Information
- Application Number
- CN202210298382.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-24
AI Technical Summary
The prior art is difficult to accurately process moving objects in video coloring, resulting in inaccurate coloring results or artifacts.
A video coloring method based on a convolutional neural network is adopted, combining the self-attention mechanism and motion information, and the chromaticity component of the target video frame is generated by extracting the motion information between video frames and the brightness and chromaticity conversion relationship of the reference frame.
It improves the accuracy of video coloring, reduces the flickering phenomenon of moving objects, and maintains the temporal continuity and spatial consistency of the video.
Smart Images

Figure CN114782571B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a black-and-white video coloring method based on a neural network and motion information, and belongs to the technical field of image processing. Background Art
[0002] The term "colorization" was proposed as early as 1970. When movies and film cameras first appeared, due to the limitations of the technical conditions at the time, movies and photos were all black and white. With the continuous development of market demand, black and white movies and photos gradually failed to meet people's needs, while color movies and photos with rich colors were very popular. How to recolor these black and white videos and photos is a problem worth studying. Colorization is not only used in the field of film art, but also in many fields, such as: in the medical field, coloring the black and white images of X-ray imaging can help doctors diagnose the disease; in the field of military aviation, after coloring satellite remote sensing images, the target can be distinguished from the background, increasing the readability of satellite images, etc. In the early days of the development of colorization technology, it was mainly to hire professionals to manually colorize videos or use media production tools to colorize videos frame by frame. This is not only labor-intensive but also costly. With the development of deep learning, the combination of convolutional neural networks and the image field has broadened the ideas for solving problems. A series of image colorization methods based on convolutional networks have emerged. These methods have achieved good colorization effects and greatly saved manpower and time.
[0003] Video colorization is a challenging problem. Compared with image colorization, since the video is composed of multiple video frames, when coloring the video, it is necessary not only to ensure the rationality of the colorization, but also to maintain the spatial consistency and temporal continuity between frames. In video colorization, the motion in the video often affects the result of video colorization. The more objects move in a video and the faster the motion, the more difficult the colorization is. If the image colorization method is used to colorize the video, each frame in the black-and-white video is regarded as an image, and the corresponding color reference image is selected for each frame in the black-and-white video for matching to generate a color video frame. Finally, each frame of the video frame that has been colored is connected to complete the entire colorization process. However, coloring each frame of the image separately does not take into account the connection between the video frames, and eventually the coloring difference between the frames often causes visual flickering when the video is played. Summary of the invention
[0004] In the prior art video colorization, it is difficult to colorize moving objects in the video, the colorization results are inaccurate, and even artifacts appear in the colorization results. The present invention proposes a video colorization method based on a self-attention mechanism and motion information.
[0005] Terminology explanation:
[0006] The source and reference attention module is essentially the same as the self-attention mechanism, except that the self-attention mechanism only has one input, the source feature, and the self-attention mechanism only focuses on its own internal connections. The source and reference attention module takes two different features as input, one corresponding to the source feature and the other corresponding to the reference feature. The source and reference attention module can mine the non-local similarity between the source feature and the reference feature, allowing the network to find the area in the reference feature that is similar to the source feature and act on the source feature.
[0007] The self-attention mechanism was first proposed by the Google team in 2017 and was initially used in the Transformer language model. Compared with the attention mechanism, the self-attention mechanism focuses on the internal connections.<Key,Value> In the form of key-value pairs, according to the query value Query in the given task goal, the similarity coefficient between Key and Query is calculated to obtain the weight coefficient corresponding to the Value value, and then the weight coefficient is used to weight the Value value to get the output. Q, K, and V are used to represent Query, Key, and Value respectively. The Q, K, and V of the self-attention mechanism all come from the same data source, as shown in formula (), where It is a scaling factor used to prevent the inner product value from being too large and affecting network learning.
[0008]
[0009] The technical solution of the present invention is:
[0010] A black-and-white video colorization method based on a neural network and motion information comprises: inputting a black-and-white video frame to be colored, i.e., a target black-and-white video frame and a reference video frame, into a trained video colorization model, extracting motion information between the brightness component of the reference video frame and the brightness component of the target black-and-white video frame, combining the motion information with the obtained conversion relationship between brightness and chromaticity between the reference frames, obtaining the conversion relationship between brightness and chromaticity between the target black-and-white video frames, applying the obtained conversion relationship to the target black-and-white video frame, obtaining the chromaticity component of the target black-and-white video frame, and completing the black-and-white video colorization.
[0011] Preferably, according to the present invention, the training process of the trained video colorization model is as follows:
[0012] Obtain the data set, preprocess the data set, and split it into training set and test set;
[0013] A video colorization model is constructed, and the obtained training set is input into the video colorization model for training, and the test set is input into the trained black and white video colorization model for testing to obtain a trained video colorization model.
[0014] Preferably, according to the present invention, the video coloring model includes a motion information extraction network, a reference feature extraction network, and a coloring network; the motion information extraction network extracts features from the brightness components of the black-and-white video frame and the reference frame respectively, combines the features of the black-and-white video frame with the features of the brightness components of the reference frame, and obtains motion information between the reference frame and the black-and-white video frame;
[0015] The reference feature extraction network extracts the features of the luminance and chrominance components in the reference frame, fuses the extracted features with the motion information, and sends them to the colorization network;
[0016] The colorization network fuses the extracted features and motion information and restores the features to their original size, predicting the chrominance components of the black-and-white video frame to be colored, thereby realizing the colorization of the black-and-white video frame.
[0017] Preferably, according to the present invention, the motion information extraction network includes an input-end feature extraction module, a reference-end brightness component feature extraction module, and a source and reference attention module; the features of the input black-and-white video frame to be colored are extracted by the input-end feature extraction module, the features of the brightness component of the reference frame are extracted by the reference-end brightness component feature extraction module, and the features of the black-and-white video frame to be colored are fused with the features of the brightness component of the reference frame by the source and reference attention mechanism module to obtain the motion information between the reference frame and the black-and-white video frame.
[0018] Further preferably, the input-end feature extraction module and the reference-end brightness component feature extraction module both include an input layer, a convolution layer, a BN layer, and an activation function layer;
[0019] The convolution layer is used to extract features from the input video frame, obtain the features of the video frame, and reduce the size of the features of the video frame; the BN layer is used for normalization; and the activation layer is used to realize nonlinear mapping of the features of the video frame.
[0020] Further preferably, the convolution layer uses 3D convolution, and the convolution kernel size is 1×3×3.
[0021] Further preferably, the input feature extraction module is as shown in formula (I):
[0022] y in =σ 1 (w 1 ×y input ) (I)
[0023] In formula (I), w 1 represents the weight, y in represents the features of the extracted black-and-white video frame to be colored, σ 1 represents the activation function, w 1 By back-propagation update, Represents the i-th black-and-white video frame of the input, where i represents the frame number of the input black-and-white video frame.
[0024] Further preferably, the reference end brightness component feature extraction module is as shown in formula (II):
[0025] y ref =σ 1 (w 2 ×y reference ) (II)
[0026] In formula (II), w 2 represents the weight, y ref represents the features of the extracted reference frame, σ 1 represents the activation function, w 2 By back-propagation update, Represents the xth reference frame of the input, where x represents the number of the reference frame.
[0027] Preferably, according to the present invention, the final output of the motion information extraction network is shown in formula (III):
[0028] M=A 1 (y in ,y ref ) (III)
[0029] In formula (III), M represents the extracted motion information, A 1 (·,·) denotes the source and reference attention modules.
[0030] Preferably, according to the present invention, the reference feature extraction network includes an input layer, a convolutional layer, a BN layer, and an activation function layer;
[0031] It includes two feature extraction branches. The first branch extracts features of 1 / 8 of the original size of the reference frame, which are then combined with the motion information through the source and reference attention module. The second branch extracts features of 1 / 16 of the original size of the reference frame, and then downsamples the motion information to a size that matches the reference frame features. After the self-attention mechanism, attention is focused on more effective information. The first branch and the second branch are combined to obtain the final output of the reference frame extraction network.
[0032] Further preferably, the first branch is as shown in formula (IV):
[0033]
[0034] In formula (IV), w 3 represents the weight, represents the extracted features of the reference frame at 1 / 8 of its original size, σ1 represents the activation function, w 3 By back-propagation update,
[0035] x represents the frame number of the reference frame, refers to the xth reference frame;
[0036] The features of the reference frame at 1 / 8 of the original size are combined with the motion information as shown in formula (V):
[0037]
[0038] In formula (V), a 1 represents the characteristics of the first branch output, A 1 (·,·) denotes the source and reference attention modules.
[0039] Further preferably, the second branch is as shown in formula (VI):
[0040]
[0041] In formula (VI), w 4 represents the weight, represents the extracted features of the reference frame at 1 / 16 of its original size, σ 1 represents the activation function, w 4 Update via back-propagation;
[0042] The motion information is downsampled once to a size that matches the reference frame features as shown in formula (VII):
[0043] M down =σ 1 (w 5 ×M) (VII)
[0044] In formula (VII), w 5 represents the weight, M down represents the feature of motion information after downsampling to 1 / 16 size, σ 1 represents the activation function, w 5 Update via back-propagation;
[0045] The features of the reference frame at 1 / 16 of the original size are combined with the motion information as shown in formula (VIII):
[0046]
[0047] In formula (VIII), b represents the feature extracted from the second branch, A 1 (·,·) denotes the source and reference attention modules.
[0048] The features extracted by the second branch are expressed as follows after the self-attention mechanism:
[0049] a 2 =S 1 (b,b) (IX)
[0050] In formula (IX), a 2 represents the characteristics of the second branch output, S 1 (·,·) represents the self-attention mechanism;
[0051] The first branch and the second branch are combined to obtain the final output O(yuv reference ), as shown in formula (X):
[0052] O(yuv reference )=a 1 +a 2 (X).
[0053] Preferably, according to the present invention, the coloring network includes an upsampling layer, a convolution layer, a BN layer, an activation function layer and a self-attention mechanism;
[0054] The upsampling layer restores the features to the original video frame size, the convolution layer is used to predict the chrominance components of the black and white video frames; the BN layer is used for normalization to accelerate the training process; the activation function layer is used to realize the nonlinear mapping of features; and the self-attention mechanism is used to obtain more effective information.
[0055] According to the preferred embodiment of the present invention, the coloring network is as shown in formula (XI) and formula (XII):
[0056] b uv =S 1 (O(yuv reference ),O(yuv reference )) (XI)
[0057] O uv =σ 1 (w 6 ×b uv ) (XII)
[0058] In formula (XI) and formula (XII), w 6 represents weight, O(yuv reference ) represents the features extracted by the reference feature extraction network, σ 1 represents the activation function, w 6 By back-propagation update, S 1 represents the self-attention mechanism, O uv represents the chrominance component of the final predicted black-and-white video frame to be colored, b uvRepresents the weighted features obtained after the feature passes through the self-attention module. uv is the target chroma component obtained after passing through the shading network.
[0059] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a black-and-white video colorization method based on a neural network and motion information when executing the computer program.
[0060] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a black-and-white video colorization method based on a neural network and motion information.
[0061] The beneficial effects of the present invention are:
[0062] The present invention proposes a video colorization method based on a convolutional neural network and combined with a self-attention mechanism and motion information. The method extracts motion information between the input black-and-white video frame and the brightness component of the reference frame, which can help the colorization network to color the moving objects. At the same time, multi-branch feature extraction is used to obtain a more accurate coloring effect in terms of details. The self-attention mechanism and the source and reference attention mechanism are also used to help the network obtain more effective feature information and improve the accuracy of coloring. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flow chart of a video colorization method based on a convolutional neural network and combined with a self-attention mechanism and motion information of the present invention;
[0064] Figure 2 Schematic diagram of the network structure of the video coloring model of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.
[0066] Example 1
[0067] A black-and-white video colorization method based on a neural network and motion information comprises: inputting a black-and-white video frame to be colored, i.e., a target black-and-white video frame and a reference video frame, into a trained video colorization model, extracting motion information between the brightness component of the reference video frame and the brightness component of the target black-and-white video frame, combining the motion information with the obtained conversion relationship between brightness and chromaticity between the reference frames, obtaining the conversion relationship between brightness and chromaticity between the target black-and-white video frames, applying the obtained conversion relationship to the target black-and-white video frame, obtaining the chromaticity component of the target black-and-white video frame, and completing the black-and-white video colorization.
[0068] There is a conversion relationship between the brightness component and the chromaticity component in the video. Since the reference video frame has complete brightness components and chromaticity components, it is hoped that the conversion relationship between the brightness component and the chromaticity component can be fitted from the reference video frame. However, since the reference video frame and the target black-and-white video frame are not exactly the same, the conversion relationship between the reference frame and the target frame is not exactly the same, and this difference is often caused by motion. The present invention extracts the motion information between the brightness component of the reference video frame and the brightness component of the target black-and-white video frame. After combining the motion information with the obtained brightness and chromaticity conversion relationship between the reference frames, the brightness and chromaticity conversion relationship between the target black-and-white video frames is obtained. After the conversion relationship is obtained, the chromaticity component of the target black-and-white video frame can be obtained by applying it to the target black-and-white video frame. The principle is expressed by the following formula:
[0069] f:Y input →U output
[0070] f reference :Y reference →U reference
[0071] motion=M(Y input ,Y referemce )
[0072]
[0073]
[0074] Among them, Y input Represents the brightness component of the target black and white video frame, U output Y represents the chrominance component of the target black-and-white video frame, and f represents the conversion relationship between the brightness component and the chrominance component of the target black-and-white video frame. reference Represents the brightness component of the reference frame, U reference represents the chrominance component of the reference frame, f reference It represents the conversion relationship between the luminance component and the chrominance component of the reference frame. Motion represents the motion information in the video frame, and M(·,·) represents the operation of extracting motion information.
[0075] Example 2
[0076] The difference between the black-and-white video colorization method based on neural network and motion information described in Example 1 is that:
[0077] The training process of the trained video colorization model is as follows:
[0078] Obtain the data set, preprocess the data set, and split it into training set and test set;
[0079] Construct a video colorization model, input the obtained training set into the video colorization model for training, input the test set into the trained black and white video colorization model for testing, and obtain a trained video colorization model. The training process of the trained video colorization model specifically includes:
[0080] The dataset size is scaled to 256×256; the grayscale video frame after removing the chrominance component is used as input, the reference frame is the first frame of the video to be colored, and 1 to 5 frames are randomly selected from the dataset; an end-to-end training method is adopted, the batch size is set to 5, and the L1 loss function is used. During training, the weights in the network are continuously updated and optimized by gradient descent. The Adadelta optimization algorithm is used to adaptively adjust the learning rate according to the gradient, instead of manually changing the learning rate parameters.
[0081] like Figure 1 As shown in, the video coloring model includes a motion information extraction network, a reference feature extraction network, and a coloring network; Figure 2 As shown, the motion information extraction network extracts features from the brightness component of the black-and-white video frame and the reference frame through the convolution layer, where the target black-and-white video frame and the reference frame are multi-frame inputs. In the process of motion information extraction, the size of the feature is continuously reduced to reduce the memory occupied during training. After the convolution layer, the features of the black-and-white video frame are combined with the features of the brightness component of the reference frame through the source and reference attention mechanism to obtain the motion information between the reference frame and the black-and-white video frame;
[0082] The reference feature extraction network extracts the features of the luminance and chrominance components in the reference frame through a convolutional layer. The extraction process is divided into two branches. One branch reduces the reference features to 1 / 8 of the original size, and the other branch further reduces the reference features to 1 / 16 of the original size. The extracted features and motion information are then fused together through the source and reference attention mechanism and sent to the colorization network.
[0083] The colorization network is composed of a self-attention mechanism module and a convolutional layer. The colorization network fuses the extracted features and motion information and restores the features to their original size. Since a large number of features are extracted, the self-attention mechanism can help the network focus on more important information among many features, and finally predict the chroma component of the black-and-white video frame to be colored, thereby realizing the colorization of the black-and-white video frame.
[0084] The motion information extraction network includes an input-end feature extraction module, a reference-end brightness component feature extraction module, and a source and reference attention module; the features of the input black-and-white video frame to be colored are extracted through the input-end feature extraction module, the features of the brightness component of the reference frame are extracted through the reference-end brightness component feature extraction module, and the features of the black-and-white video frame to be colored are fused with the features of the brightness component of the reference frame through the source and reference attention mechanism module to obtain the motion information between the reference frame and the black-and-white video frame.
[0085] The input-end feature extraction module and the reference-end brightness component feature extraction module both include an input layer, a convolution layer, a BN (Batch Normalization) layer, and an activation function layer;
[0086] The input layer of the input feature extraction module is used to input black and white video frames Where T represents the black and white video frame y input The number of frames, H represents the black and white video frame y input The length of W represents the black and white video frame y input The width of the convolution layer is 1, which represents a single channel (i.e., a grayscale image). The convolution layer is used to extract features from the input video frames, obtain the features of the video frames, and reduce the size of the features of the video frames; it can adapt to different input sizes and frame numbers. The key to the convolution operation is the convolution kernel (Kernel Size) and the stride (Stride). In this embodiment, since the input is multiple frames, the convolution layer uses 3D convolution, and the convolution kernel of the convolution layer is 1×3×3, and the stride is 1×1×1 or 1×2×2, where the stride of 1×2×2 is to reduce the size of the features. The small size can reduce the complexity of network calculations.
[0087] Similarly, the input layer of the reference end luminance component feature extraction module is used to input the reference video frame Where T represents the reference video frame y reference The number of frames, H represents the reference video frame y reference The length of W represents the reference video frame y reference The width of the convolution layer is 3, and 3 represents 3 channels (i.e., color image). The convolution layer is used to extract the features of the reference video frame to obtain the features of the reference video frame. Specifically, the main purpose of the convolution operation performed by the convolution layer is to extract and map the features of the reference video frame. In this embodiment, the convolution kernel of the convolution layer is 1×3×3, and the step size is 1×1×1 or 1×2×2, where the step size of 1×2×2 is to reduce the size of the feature. In this embodiment, 8 groups of convolution operations are used to extract black and white video features, and the parameter details are set as shown in Table 1.
[0088] Table 1
[0089]
[0090] Since the different distributions of training data and test data will lead to a decrease in the generalization performance of the network during the neural network training process, in order to increase the generalization of the network and improve the training speed, a BN layer is set in the feature extraction network. The BN layer normalizes the features to prevent gradient explosion.
[0091] The activation layer is used to implement nonlinear mapping of the features of the video frame. In this embodiment, the ELU function is used as the activation function:
[0092]
[0093] The ELU function can bring the output mean of the activation function closer to 0, making the gradient closer to the natural gradient and improving the robustness to noise.
[0094] The input feature extraction module is shown in formula (I):
[0095] y in =σ 1 (w 1 ×y input ) (I)
[0096] In formula (I), w 1 represents the weight, y in represents the features of the extracted black-and-white video frame to be colored, σ 1 represents the activation function, w 1 By back-propagation, w 1 is a matrix that corresponds to each input feature one by one, representing the importance of each feature. The weight is updated through back propagation. During the training process of the convolutional layer, the gradient descent algorithm will gradually change w in order to make the output value of the loss function smaller. 1 , thereby gradually making the prediction more accurate. Represents the i-th black-and-white video frame of the input, where i represents the frame number of the input black-and-white video frame.
[0097] The reference end brightness component feature extraction module is shown in formula (II):
[0098] y ref =σ 1 (w 2 ×y reference ) (II)
[0099] In formula (II), w 2 represents the weight, y ref represents the features of the extracted reference frame, σ 1 represents the activation function, w 2 By back-propagation update, Represents the xth reference frame of the input, where x represents the number of the reference frame.
[0100] The self-attention mechanism is a type of attention mechanism. The source and reference attention module is essentially the same as the self-attention mechanism, except that the self-attention mechanism only has one input, the source feature, and only focuses on its own internal connections. The source and reference attention module takes two features as input, one corresponding to the source feature and the other corresponding to the reference feature. The source and reference attention module can mine the non-local similarity between the source feature and the reference feature, allowing the network to use the region in the reference feature that is similar to the source feature to act on the source feature.
[0101] The features of the input brightness component and the reference brightness component are connected through the source and reference attention module, using A sr (·,·) represents the source and reference attention operations. For features S and R of C×T×H×W, C represents the number of channels, T represents the number of frames, H and W represent the length and width respectively. The source and reference attention mechanism is expressed as:
[0102] A sr (S,R)=S+γd(e t (R)softmak(e r (R) T e s (S)))
[0103] Where γ is the learning rate, Represents the reduction of the dimension of the feature map. It represents increasing the dimension of the feature.
[0104] Use A 1 (·,·) represents the source and reference attention modules. The features of the input brightness component and the reference brightness component are combined through the source and reference attention modules to obtain motion information. The final output of the motion information extraction network is shown in formula (III):
[0105] M=A 1 (y in ,y ref ) (III)
[0106] In formula (III), M represents the extracted motion information, A 1 (·,·) denotes the source and reference attention modules.
[0107] The reference feature extraction network includes input layer, convolution layer, BN (Batch Normalization) layer, and activation function layer;
[0108] It includes two feature extraction branches. The first branch extracts features of 1 / 8 of the original size of the reference frame, and then combines them with the motion information through the source and reference attention module. The second branch extracts features of 1 / 16 of the original size of the reference frame, and then downsamples the motion information to a size that matches the reference frame features. After the self-attention mechanism, the attention is focused on more effective information. After combining the source and reference attention modules, since this branch extracts more features, a self-attention mechanism is added to this branch to help the network focus on more effective information. The structure of using two branches to extract features in parallel helps to utilize the detailed features of the reference frame. The parameter details are shown in Table 2.
[0109] Table 2
[0110]
[0111] The first branch and the second branch are combined to obtain the final output of the reference frame extraction network.
[0112] The first branch is shown in formula (IV):
[0113]
[0114] In formula (IV), w 3 represents the weight, represents the extracted features of the reference frame at 1 / 8 of its original size, σ 1 represents the activation function, w 3 By back-propagation update,
[0115] x represents the frame number of the reference frame, refers to the xth reference frame;
[0116] The features of the reference frame at 1 / 8 of the original size are combined with the motion information as shown in formula (V):
[0117]
[0118] In formula (V), a 1 represents the characteristics of the first branch output, A 1 (·,·) denotes the source and reference attention modules.
[0119] The second branch is shown in formula (VI):
[0120]
[0121] In formula (VI), w 4 represents the weight, represents the extracted features of the reference frame at 1 / 16 of its original size, σ 1 represents the activation function, w 4 Update via back-propagation;
[0122] The motion information is downsampled once to a size that matches the reference frame features as shown in formula (VII):
[0123] M down =σ 1 (w 5 ×M) (VII)
[0124] In formula (VII), w 5 represents the weight, M down represents the feature of motion information after downsampling to 1 / 16 size, σ 1 represents the activation function, w 5 Update via back-propagation;
[0125] The features of the reference frame at 1 / 16 of the original size are combined with the motion information as shown in formula (VIII):
[0126]
[0127] In formula (VIII), b represents the feature extracted from the second branch, A 1 (·,·) denotes the source and reference attention modules.
[0128] The features extracted by the second branch are expressed as follows after the self-attention mechanism:
[0129] a 2 =S 1 (b,b) (IX)
[0130] In formula (IX), a 2 represents the characteristics of the second branch output, S 1 (·,·) represents the self-attention mechanism;
[0131] The first branch and the second branch are combined to obtain the final output O(yuv reference ), as shown in formula (X):
[0132] O(yuv reference )=a 1 +a 2 (X).
[0133] The coloring network includes upsampling layer, convolution layer, BN (Batch Normalization) layer, activation function layer and self-attention mechanism;
[0134] The upsampling layer uses a trilinear interpolation function to restore the features to the size of the chrominance component of the original video frame. Since the video frame used in the example is in YUV 4:2:0 format, the size of the chrominance component is half the size of the input luminance component. The convolution layer is used to predict the chrominance component of the black and white video frame. The BN layer is used for normalization to accelerate the training process. The activation function layer is used to implement nonlinear mapping of features. The self-attention mechanism helps the network obtain more effective information. The coloring network predicts the input features as the chrominance components (UV components) of the black and white video, and finally combines them with the corresponding black and white video frames to generate multiple color video frames. In this embodiment, 10 groups of convolution operations are used to fuse the black and white video frame features with the reference color frame features and finally predict the chrominance components of the black and white video frame. The parameter details are set as shown in Table 3.
[0135] Table 3
[0136]
[0137]
[0138] The coloring network is shown in formula (XI) and formula (XII):
[0139] b uv =S 1 (O(yuv reference ),O(yuv reference )) (XI)
[0140] O uv =σ 1 (w 6 ×b uv ) (XII)
[0141] In formula (XI) and formula (XII), w 6 represents weight, O(yuv reference ) represents the features extracted by the reference feature extraction network, σ 1 represents the activation function, w 6 By back-propagation update, S 1 represents the self-attention mechanism, O uv represents the chrominance component of the final predicted black-and-white video frame to be colored, b uv Represents the weighted features obtained after the features pass through the self-attention module. uv is the target chroma component obtained after passing through the shading network.
[0142] The loss function of the video colorization model adopts the minimum absolute error function L 1 :
[0143]
[0144] k, Represent the true value of the chrominance component and the predicted value of the network output respectively.
[0145] The effects of the present invention are described below through experiments.
[0146] This experiment uses Youku-VESR and Videvo as training sets. Youku-VESR includes 998 videos, each of which takes the first 90 frames. Videvo selects 50 videos from the Videvo video website. Both parts are converted to YUV4:2:0 video format, a total of 1343 videos, and 119,527 frames are used for training. The test set uses 30 videos selected from the natural attribute classification of the DAVIS dataset and the Videvo video website. During the experiment, the brightness components of 5 consecutive video frames are selected as input each time. During the training process, 5 video frames are selected as references. The true value of the first frame of each video and the previous frame of the current frame are used as two reference frames, and the remaining three reference frames are randomly selected from the dataset.
[0147] The black and white video colorization effect of the present invention is reasonable in visual observation, and the colored video can maintain good temporal continuity and spatial consistency, can maintain better colorization effect on moving objects, and reduce the occurrence of flickering. In addition, the results obtained by the method of the present invention are compared with the current advanced video colorization methods: the method of Iizuka et al. (Zhang B, He M, J Liao, et al. Deep Exemplar-based Video Colorization [C] / / 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019.) and the method of Zhang et al. (Iizuka S, Simo-Serra E. Deep Remaster: Temporal Source-Reference Attention Networks for Comprehensive Video Enhancement [J]. ACM Transactions on Graphics, 2019, 38(6): 176.1-176.13.) During the test, the brightness components of 5 consecutive video frames are selected as input each time, and the true value of the first frame of each video is selected as a reference. The test is performed under the same conditions. The results show that the method of Iizuka et al. and the method of Zhang et al. have the problem of background blur, and the color of moving objects will be lost in long sequence video frames. The colorization result obtained by the method of the present invention is visually closer to the true value, and can better maintain the colorization effect on moving objects in long video sequences. The colorization effect is clear without blurring, and the color is reasonable.
[0148] The present invention also uses quantitative indicators PSNR and SSIM to compare with other methods, as shown in Table 4:
[0149] Table 4
[0150]
[0151] In Table 4, Zhang et al. refers to the method of Zhang et al., and Iizuka refers to the method of Iizuka et al. It can be seen from Table 4 that the PSNR and SSIM results of the present invention are better than those of the other two methods. This also shows that the coloring effect of the present invention is more stable.
[0152] Example 3
[0153] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the black-and-white video colorization method based on a neural network and motion information in embodiment 1 or 2 are implemented.
[0154] Example 4
[0155] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the black-and-white video colorization method based on a neural network and motion information in embodiment 1 or 2.
Claims
1. A black-and-white video coloring method based on neural network and motion information, characterized in that, it includes: Input the black-and-white video frame to be colored, i.e., the target black-and-white video frame, and the reference video frame into the trained video coloring model. Extract the motion information between the luminance components of the reference video frame and the target black-and-white video frame. After combining the motion information with the conversion relationship between luminance and chrominance of the obtained reference frame, obtain the conversion relationship between luminance and chrominance of the target black-and-white video frame. After obtaining the conversion relationship, apply it to the target black-and-white video frame to obtain the chrominance component of the target black-and-white video frame, that is, the black-and-white video coloring is completed; The video coloring model includes a motion information extraction network, a reference feature extraction network, and a coloring network; The motion information extraction network extracts features from the luminance components of the black-and-white video frame and the reference frame respectively, combines the features of the black-and-white video frame with the features of the luminance component of the reference frame, and obtains the motion information between the reference frame and the black-and-white video frame; The reference feature extraction network extracts the features of the luminance component and the chrominance component in the reference frame, fuses the extracted features and the motion information together, and sends them into the coloring network; the coloring network fuses the extracted features and the motion information and restores the features to the original size, and predicts the chrominance component of the black-and-white video frame to be colored, that is, the coloring of the black-and-white video frame is realized.
2. The black-and-white video coloring method based on neural network and motion information according to claim 1, characterized in that, The training process of the trained video coloring model is as follows: Obtain a data set, preprocess the data set, and divide it into a training set and a test set; Construct a video coloring model, input the obtained training set into the video coloring model for training, input the test set into the trained black-and-white video coloring model for testing, and obtain the trained video coloring model.
3. The black-and-white video coloring method based on neural network and motion information according to claim 1, characterized in that, The motion information extraction network includes an input end feature extraction module, a reference end luminance component feature extraction module, and a source and reference attention module; Extract the features of the input black-and-white video frame to be colored through the input end feature extraction module, extract the features of the reference frame luminance component through the reference end luminance component feature extraction module, and fuse the features of the black-and-white video frame to be colored with the features of the reference frame luminance component through the source and reference attention mechanism module to obtain the motion information between the reference frame and the black-and-white video frame.
4. The black-and-white video coloring method based on neural network and motion information according to claim 3, characterized in that, Both the input end feature extraction module and the reference end luminance component feature extraction module include an input layer, a convolutional layer, a BN layer, and an activation function layer; The convolutional layer is used to extract features from the input video frame, obtain the features of the video frame, and reduce the size of the features of the video frame; the BN layer is used for normalization; the activation layer is used to realize the non-linear mapping of the features of the video frame.
5. The black-and-white video coloring method based on neural network and motion information according to claim 3, characterized in that, The convolution layer uses 3D convolution with a kernel size of 1×3×3.
6. The black and white video colorization method based on neural network and motion information according to claim 3, It is characterized in that The input feature extraction module is shown in formula (I): y in =s 1 (w 1 ×y input (I) In formula (I), w 1 represents the weight, y in represents the features of the extracted black-and-white video frame to be colored, σ 1 represents the activation function, w 1 By back-propagation update, Represents the i-th black-and-white video frame of the input, where i represents the frame number of the input black-and-white video frame.
7. The black and white video colorization method based on neural network and motion information according to claim 3, It is characterized in that The reference end brightness component feature extraction module is shown in formula (II): y ref =s 1 (w 2 ×y reference (II) In formula (II), w 2 represents the weight, y ref represents the features of the extracted reference frame, σ 1 represents the activation function, w 2 By back-propagation update, Represents the xth reference frame of the input, where x represents the number of the reference frame.
8. The black and white video colorization method based on neural network and motion information according to claim 3, It is characterized in that The final output of the motion information extraction network is shown in formula (III): M=A 1 (the in ,the ref )(III) In formula (III), M represents the extracted motion information, A 1 (·,·) denotes the source and reference attention modules.
9. The black and white video colorization method based on neural network and motion information according to claim 1, It is characterized in that The reference feature extraction network includes input layer, convolution layer, BN layer, and activation function layer; It includes two feature extraction branches. The first branch extracts features of 1 / 8 of the original size of the reference frame, which are then combined with the motion information through the source and reference attention module. The second branch extracts features of 1 / 16 of the original size of the reference frame, and then downsamples the motion information to a size that matches the reference frame features. After the self-attention mechanism, attention is focused on more effective information. The first branch and the second branch are combined to obtain the final output of the reference frame extraction network.
10. A black and white video colorization method based on neural network and motion information according to claim 9, It is characterized in that The first branch is shown in formula (IV): In formula (IV), w 3 represents the weight, represents the extracted features of the reference frame at 1 / 8 of its original size, σ 1 represents the activation function, w 3 By back-propagation update, x represents the frame number of the reference frame, refers to the xth reference frame; The features of the reference frame at 1 / 8 of the original size are combined with the motion information as shown in formula (V): In formula (V), a 1 represents the characteristics of the first branch output, A 1 (·,·) denotes the source and reference attention modules.
11. The black and white video colorization method based on neural network and motion information according to claim 9, It is characterized in that The second branch is shown in formula (VI): In formula (VI), w $ represents the weight, represents the extracted features of the reference frame at 1 / 16 of its original size, σ 1 represents the activation function, w $ Update via back-propagation; The motion information is downsampled once to a size that matches the reference frame features as shown in formula (VII): M %o′n =σ 1 (In ( ×M)(VII) In formula (VII), w ( represents the weight, M %o′n represents the feature of motion information after downsampling to 1 / 16 size, σ 1 represents the activation function, w ( Update via back-propagation; The features of the reference frame at 1 / 16 of the original size are combined with the motion information as shown in formula (VIII): In formula (VIII), b represents the feature extracted from the second branch, A 1 (·,·) denotes source and reference attention modules; The features extracted by the second branch are expressed as follows after the self-attention mechanism: a 2 =S 1 (b,b)(IX) In formula (IX), a 2 represents the characteristics of the second branch output, S 1 (·,·) represents the self-attention mechanism; The first branch and the second branch are combined to obtain the final output O(yuv reference ), as shown in formula (X): O(yuv reference )=a 1 +a 2 (X)。 12. The black and white video colorization method based on neural network and motion information according to claim 1, It is characterized in that The coloring network includes upsampling layer, convolution layer, BN layer, activation function layer and self-attention mechanism; The upsampling layer restores the features to the original video frame size, the convolution layer is used to predict the chrominance components of the black and white video frame; the BN layer is used for normalization to accelerate the training process; The activation function layer is used to realize nonlinear mapping of features; The self-attention mechanism is used to obtain more effective information.
13. A black and white video colorization method based on neural network and motion information according to claim 12, It is characterized in that The coloring network is shown in formula (XI) and formula (XII): b u- =S 1 .O(yuv reference ),O(yuv reference ) / (XI) O u- =σ 1 (w 0 ×b u- )(XII) In formula (XI) and formula (XII), w 0 represents weight, O(yuv reference ) represents the features extracted by the reference feature extraction network, σ 1 represents the activation function, w 0 By back-propagation update, S 1 represents the self-attention mechanism, O u- represents the chrominance component of the final predicted black-and-white video frame to be colored, b u- Represents the weighted features obtained after the features pass through the self-attention module.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the black and white video colorization method based on neural network and motion information are implemented as described in any one of claims 1-13.
15. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the black and white video colorization method based on neural network and motion information described in any one of claims 1-13 are implemented.
Citation Information
Patent Citations
Video frame supplementing method based on neural network and training method of model thereof
CN110324664A
Black-and-white video coloring method and device, storage medium and terminal
CN113421312A