Video Frame Interpolation Method, Electronic Device, Chip System and Readable Storage Medium

By canceling non-convolution operations in the video interpolation model and using the interpolation method of convolution layer, segmentation module and calculation module, the problem of low video playback fluency in the prior art is solved, and faster intermediate frame image output and higher playback fluency are achieved.

CN119478614BActive Publication Date: 2025-06-20HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411942868.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-20
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

In the existing video interpolation method, the non-convolution operation of the pooling layer and the upsampling layer is performed on the neural network processor of the electronic device for a long time, resulting in the video not being updated in time and the playback fluency is not high.

Method used

By canceling non-convolution operations in the interpolation model, using an interpolation model including a convolution layer, a segmentation module and a computing module, the pixel features are extracted using the convolution layer, the segmentation module divides the feature map into a sub-feature map, and the computing module predicts the intermediate frame image through the differential data.

Benefits of technology

The time for the interpolation model to output intermediate frame images is shortened, the smoothness of video playback is improved, and the picture can be updated in time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478614B_ABST
    Figure CN119478614B_ABST
Patent Text Reader

Abstract

The present application provides a video frame interpolation method, an electronic device, a chip system, and a readable storage medium. The method is applied to the electronic device and includes: obtaining a first image and a second image; inputting the first image and the second image into a first module, and extracting pixel features through the first module to obtain a first feature map; inputting the first feature map into a second module, and through the second module, splitting the first feature map into a first sub-feature map and a second sub-feature map, determining the difference data between the first sub-feature map and the second sub-feature map, and predicting a third image according to the difference data, where the difference data is used to represent the movement of pixels in the time series from the first image to the second image. Thus, the present application uses a frame interpolation model that does not include non-convolution operations such as pooling operations and interpolation upsampling operations to realize obtaining an intermediate image between two images, and can shorten the time for the frame interpolation model to output the intermediate image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a video frame insertion method, an electronic device, a chip system and a readable storage medium. Background Art

[0002] Video interpolation is the process of predicting and inserting one or more intermediate frames between the previous and next frames of the original video to shorten the display time span between frames and thus obtain a video with a higher frame rate.

[0003] Currently, video interpolation can be achieved through neural networks. Specifically, two frames of images in the video are input into a neural network consisting of a convolutional layer, a pooling layer, and an upsampling (interpolation type upsampling) layer, and the intermediate frame image between the two frames is output.

[0004] However, in the above implementation, the non-convolution operations of the pooling layer and the upsampling (interpolation upsampling) layer take a long time to execute on the neural-network processing unit (NPU) of the electronic device, which makes it take a long time to obtain the intermediate frame image through two frames of images each time, so that the picture is not updated in time when the video is played, resulting in low smoothness of the video playback. Summary of the invention

[0005] The present application provides a video interpolation method, an electronic device, a chip system and a readable storage medium, which can shorten the output time of intermediate frame images in an interpolation model as much as possible and improve the smoothness of video playback.

[0006] In a first aspect, the present application provides a video frame insertion method, which is applied to an electronic device, wherein the electronic device includes a frame insertion model, the frame insertion model includes a first module and a second module, the first module includes a convolution layer, and the second module includes a first segmentation module and a first calculation module; the method includes:

[0007] A first image and a second image are obtained from a target video, wherein the first image is a frame in the target video that is earlier than the second image in a time series; the first image and the second image are input into a first module to obtain a first feature map, wherein the first module is used to extract pixel features of the first image and the second image, wherein the first feature map includes pixel features of the first image and pixel features of the second image; the first feature map is input into a second module to obtain a third image, wherein a first segmentation module in the second module is used to segment the first feature map into a first sub-feature map and a second sub-feature map, wherein a first calculation module is used to determine difference data between the first sub-feature map and the second sub-feature map, and predict a third image based on the difference data, wherein the difference data is used to represent the movement of pixels from the first image to the second image in a time series, and wherein the third image is an intermediate image to be inserted between the first image and the second image.

[0008] In the above method, when performing frame interpolation processing, the first image and the second image in the target video can be input into the first module, and pixel features can be extracted through the aforementioned convolutional layer to obtain a first feature map. The first feature map is input into the second module, and the first segmentation module in the second module segments the first feature map to obtain a first sub-feature map and a second sub-feature map. Then, the first calculation module in the second module can determine the difference data between the first sub-feature map and the second sub-feature map through subtraction calculation. Since the difference data is used to represent the pixel motion of the first image and the second image in the time series, the first calculation module can obtain a third image to be inserted between the first image and the second image based on the difference data. In this way, the electronic device can obtain the third image without using an interpolation model that includes non-convolution operations such as pooling operations and interpolation-based upsampling operations. Instead, after obtaining the first feature map, the first feature map is segmented into two sub-feature maps, and the difference data obtained through the subtraction operation between the two sub-feature maps is used to predict the third image to be inserted between the first image and the second image, thereby being able to shorten the duration of the interpolation model to output the third image. When playing the video, the picture can be updated in a timely manner, and further, the video playback smoothness of the target video can be ensured.

[0009] Combined with the first aspect, in some implementation manners of the first aspect, inputting the first image and the second image into the first module to obtain a first feature map includes:

[0010] Splicing the first image and the second image to obtain a spliced image; inputting the spliced image into the first module to obtain a first feature map.

[0011] In the above method, before inputting the first image and the second image into the first module, the electronic device splices the first image and the second image to obtain a spliced image. After that, the first module can directly process the spliced image, and in the intermediate process, the second module can directly splice the first difference data, the second difference data, and the spliced image, without having to splice the first image and the second image before performing the step of splicing the first difference data, the second difference data, and the spliced image, and the fourth module can directly splice the first transformed feature map, the second transformed feature map, and the spliced image, without having to splice the first image and the second image again before performing the step of splicing the first transformed feature map, the second transformed feature map, and the spliced image. Thus, the splicing time of the first image and the second image can be saved, and the duration of the interpolation model to output the third image can be shortened.

[0012] In combination with the first aspect, in some implementations of the first aspect, the frame interpolation model further includes a third module and a fourth module. The third module includes a convolutional layer, and the fourth module includes a second segmentation module and a second calculation module. Inputting the first feature map into the second module to obtain a third image includes:

[0013] Inputting the first feature map and the spliced image into the second module to obtain a second feature map. The first calculation module is used to determine the first difference data and the second difference data between the first sub-feature map and the second sub-feature map, and splice the first difference data, the second difference data, and the spliced image. The first difference data is used to represent the movement of pixels in the time series from the first image to the second image, and the second difference data is used to represent the movement of pixels in the time series from the second image to the first image. Inputting the second feature map into the third module to obtain a third feature map. The third module is used to extract features from the second feature map. Inputting the third feature map, the first difference data, the second difference data, the first image, and the second image into the fourth module to obtain a third image. The second segmentation module in the fourth module is used to segment the third feature map into a third sub-feature map and a fourth sub-feature map. The second calculation module is used to fuse the first difference data and the third sub-feature map to obtain a first fusion feature map, fuse the second difference data and the fourth sub-feature map to obtain a second fusion feature map, perform a warping transformation on the first image according to the first fusion feature map to obtain a first transformed feature map, perform a warping transformation on the second image according to the second fusion feature map to obtain a second transformed feature map, and predict the third image according to the first transformed feature map and the second transformed feature map. The accuracy of the first transformed feature map and the second transformed feature map for predicting the third image is higher than the accuracy of the first difference data and the second difference data for predicting the third image.

[0014] In the above method, after obtaining the first feature, in addition to the first feature map, the electronic device can also input the spliced image into the second module, and after obtaining the third feature map, in addition to the third feature map, the electronic device can also input the spliced image into the fourth module, so that the intermediate layer of the frame interpolation model can also learn the original image information (referring to the spliced image), avoid the frame interpolation model losing important feature information, and ensure the performance of the model.

[0015] Moreover, after the first feature map and the spliced image are input into the second module to obtain the second feature map, the electronic device can further input the second feature map into the third module to perform feature extraction on the second feature map to obtain the third feature map, and input the third feature map, the first difference data, the second difference data, the first image, and the second image into the fourth module, so as to obtain the first transformed feature map and the second transformed feature map. The first transformed feature map is equivalent to the optimized first difference data, and the second transformed feature map is equivalent to the optimized second difference data. In this way, the accuracy of predicting the third image by the first transformed feature map and the second transformed feature map can be higher than the accuracy of predicting the third image by the first difference data and the second difference data.

[0016] Combined with the first aspect, in some implementation manners of the first aspect, the interpolation model further includes a fifth module and a sixth module. The fifth module includes a convolutional layer, and the sixth module includes a transformation module, a third segmentation module, and a third calculation module. Inputting the third feature map, the first difference data, and the second difference data into the fourth module to obtain the third image includes:

[0017] Inputting the third feature map, the first difference data, the second difference data, the first image, the second image, and the spliced image into the fourth module to obtain a fourth feature map, where the fourth feature map is obtained by splicing the first transformed feature map, the second transformed feature map, and the spliced image; inputting the fourth feature map into the fifth module to obtain a fifth feature map, where the fifth module is used to perform feature extraction on the fourth feature map; inputting the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module to obtain the third image. The transformation module in the sixth module is used to normalize each feature data in the fifth feature map to obtain a sixth feature map. The third segmentation module is used to segment the sixth feature map into a fifth sub-feature map and a sixth sub-feature map. The third calculation module is used to generate the third image including target information according to the fifth sub-feature map, the sixth sub-feature map, the first transformed feature map, and the second transformed feature map. The target information includes light and shadow information and / or occlusion information. The target information included in the third image is the target information included in the first image and the second image.

[0018] In the above method, after obtaining the fifth feature map, the electronic device can further input the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module, so that through the processing of the third segmentation module and the third calculation module, the obtained third image has the light and shadow information and / or occlusion information included in the first image and the second image. In this way, it can be ensured that there will be no abrupt changes when the first image, the third image, and the second image are played in sequence, the visual jump during playback can be avoided, and the continuity during playback can be ensured.

[0019] In combination with the first aspect, in certain implementations of the first aspect, the frame interpolation model further includes a seventh module and a first activation function. The seventh module includes a convolutional layer. Inputting the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module to obtain a third image, including:

[0020] Inputting the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module to obtain a seventh feature map; concatenating the seventh feature map and the concatenated image to obtain an eighth feature map; inputting the eighth feature map into the seventh module to obtain a ninth feature map, where the seventh module is used to extract features from the eighth feature map; fusing the seventh feature map and the ninth feature map to obtain a third fused feature map; inputting the third fused feature map into the first activation function to obtain a third image, where the first activation function is used to convert the third fused feature map into an image with pixel values maintained within a preset range.

[0021] In the above method, the seventh feature map and the ninth feature map are fused to form a residual connection, so as to retain as many image detail features including textures and clear edges as possible, thereby facilitating the generation of a third image with more complete feature information.

[0022] Moreover, through the processing of the first activation function, the pixel values of the third image output by the electronic device are all maintained within a preset range, so that the electronic device can correctly recognize the third image and display the third image normally.

[0023] In combination with the first aspect, in certain implementations of the first aspect, the first module includes a first sub-layer and a second sub-layer. The first sub-layer includes a first convolutional layer and a second activation function layer, and the second sub-layer includes a second convolutional layer. Inputting the concatenated image into the first module to obtain a first feature map, including:

[0024] Inputting the concatenated image into the first sub-layer to obtain a first intermediate feature map. The first convolutional layer in the first sub-layer is used to extract pixel features from the concatenated image to obtain a second intermediate feature map. The second activation function layer is used to perform a non-linear transformation on the second intermediate feature map to obtain the first intermediate feature map. During the process of the first convolutional layer extracting pixel features from the concatenated image, the convolution stride used is related to the target output duration of the third image, and the convolution stride is an even number greater than or equal to 2. The size of the first intermediate feature map is related to the convolution stride; inputting the first intermediate feature map into the second sub-layer to obtain a first feature map, and the second convolutional layer in the second sub-layer is used to adjust the size of the first intermediate feature map to be the same as the size of the concatenated image.

[0025] In the above method, the convolution stride is adjustable. The smaller the convolution stride, the greater the overall computational amount of the first module, the longer the overall execution duration of the first module, and the longer the output duration of the third image. The larger the convolution stride, the smaller the overall computational amount of the first module, the shorter the overall execution duration of the first module, and the shorter the output duration of the third image. By setting the size of the convolution stride to a value related to the target output duration of the third image, the output duration of the third image can be ensured.

[0026] Combined with the first aspect, in some implementation manners of the first aspect, the first module further includes a padding layer and a cropping layer. Inputting the spliced image into the first sub-layer to obtain a first intermediate feature map, including:

[0027] Input the spliced image into the padding layer, and determine whether the length and width of the spliced image are odd numbers respectively through the padding layer; if the length and / or width of the spliced image is odd, perform a first expansion on the pixels of the spliced image to obtain an updated spliced image, and the first expansion is used to make the length and width of the spliced image both even numbers; input the updated spliced image into the first sub-layer to obtain a first intermediate feature map; input the first intermediate feature map into the second sub-layer to obtain a first feature map, including: input the first intermediate feature map into the second sub-layer to obtain a second intermediate feature map, and the second convolutional layer of the second sub-layer is used to adjust the size of the first intermediate feature map to the same size as the updated spliced image; input the second intermediate feature map into the cropping layer to obtain a first feature map, and the cropping layer is used to adjust the size of the second intermediate feature map to the same size as the spliced image.

[0028] In the above method, considering that when the length and / or width of the spliced image is odd, the spliced image cannot be processed by the convolutional layer corresponding to an even convolution stride. When the length and / or width of the spliced image is odd, first, by padding with 0, the length and width of the spliced image are both made even numbers to obtain an updated spliced image, and then input it into the first convolutional layer for processing. The size of the first intermediate feature map obtained after processing by the first convolutional layer and the second activation function layer is a smaller size. Based on this, by padding with 0, the second convolutional layer of the second sub-layer can also adjust the size of the first intermediate feature map to be the same as the updated spliced image to obtain a second intermediate feature map, and then the second intermediate feature map is cropped by the cropping layer to obtain a first feature map. In this way, the size of the first feature map is the same as the size of the spliced image, which can avoid the situation that the size of the third image obtained from the first feature map is different from the sizes of the first image and the second image.

[0029] Combined with the first aspect, in some implementation manners of the first aspect, the first image and the second image are two adjacent frames of images in the target video, or the first image and the second image are two frames of images corresponding to N frames apart in the target video, where N is a positive integer greater than or equal to 1.

[0030] For example, the target video includes a total of 9 frames of images. The first image is the first frame among the 9 frames of images, and the second image is the second frame among the 9 frames of images.

[0031] For example, the target video includes a total of 9 frames of images. The second frame, the fourth frame, the sixth frame, and the eighth frame are extracted from the target video. The remaining images include the first frame, the third frame, the fifth frame, the seventh frame, and the ninth frame. The value of N is 1. The first image is the first frame among the 9 frames of images, and the second image is the third frame among the 9 frames of images.

[0032] In combination with the first aspect, in some implementation manners of the first aspect, the generation process of the interpolation model includes:

[0033] Obtain multiple groups of training images. Each group of training images in the multiple groups of training images includes three images that are consecutive in time series. The second image in each group of training images is the label intermediate image corresponding to each group of training images; input the first image and the third image in each group of training images into the first sample module of the neural network to obtain a sample feature map. The first sample module is used to extract pixel features of the first image and the third image through the convolutional layer included in the first sample module. The sample feature map includes the pixel features of the first image and the pixel features of the third image; input the sample feature map into the second sample module of the neural network to obtain the estimated intermediate image corresponding to each group of training images. The second sample module includes a first sample segmentation module and a first sample calculation module. The first sample segmentation module is used to segment the sample feature map into a first sub-sample feature map and a second sub-sample feature map. The first sample calculation module is used to determine the sample difference data between the first sub-sample feature map and the second sub-sample feature map, and determine the estimated intermediate image according to the sample difference data. The sample difference data is used to represent the movement of pixels from the first image to the third image in time series; train the neural network according to the difference between the label intermediate image and the estimated intermediate image to obtain the interpolation model.

[0034] In this application, the generation process of the interpolation model and the video interpolation method of the embodiments of this application can be executed by different devices. For example, the generation process of the interpolation model is executed by a server, and the video interpolation method of the embodiments of this application is executed by an electronic device.

[0035] In the above method, during the training process, according to the first image and the third image in multiple groups of three consecutive images, the intermediate image can be predicted. Then, according to the difference between the predicted intermediate image and the second image in the three consecutive images, that is, the loss between the predicted intermediate image and the second image in the three consecutive images, through the backpropagation method, the neural network can update the model parameters of the neural network according to the foregoing loss, so as to obtain the interpolation model.

[0036] In a second aspect, the present application provides a video frame interpolation device, which includes a module for executing the method in the first aspect and any possible implementation manner of the first aspect.

[0037] In a third aspect, the present application provides an electronic device, which includes: one or more processors, and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0038] In a fourth aspect, the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0039] In a fifth aspect, the present application provides a computer-readable storage medium, which includes instructions, and when the instructions run on an electronic device, the electronic device is enabled to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0040] In a sixth aspect, the present application provides a computer program product, and when the computer program product runs on a computer, the computer is enabled to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0041] It can be understood that the beneficial effects of the above second aspect to the sixth aspect can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings

[0042] Figure 1 It is a schematic diagram of frame interpolation for a video or animation effect provided by an embodiment of the present application;

[0043] Figure 2 It is a schematic diagram of the classification of various layers in an interpolation model provided by the prior art;

[0044] Figure 3 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0045] Figure 4 It is a schematic diagram of the software architecture of an electronic device provided by an embodiment of the present application;

[0046] Figure 5 It is a schematic diagram of the process of a video frame interpolation method provided by an embodiment of the present application;

[0047] Figure 6Schematic diagram of the structure of a first module provided by an embodiment of the present application;

[0048] Figure 7 Schematic flowchart of a video frame interpolation method provided by an embodiment of the present application;

[0049] Figure 8 Schematic diagram of the structure of a second module provided by an embodiment of the present application;

[0050] Figure 9 Schematic diagram of the structure of a fourth module provided by an embodiment of the present application;

[0051] Figure 10 Schematic diagram of the structure of a sixth module provided by an embodiment of the present application;

[0052] Figure 11 Schematic diagram of the structure of a frame interpolation model provided by an embodiment of the present application. Detailed implementation manners

[0053] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the first conversion relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a alone, b alone, or c alone may represent: a alone, b alone, c alone, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b, and c, where a, b, and c may be single or multiple. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0054] For ease of understanding, some examples related to the concepts of the embodiments of the present application are given for reference.

[0055] 1. Convolutional layer

[0056] The convolutional layer can consist of many convolutional operators, also known as kernels, which act as filters for extracting specific information from the input image in image processing. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix typically processes the input image pixel by pixel (or two pixels at a time... depending on the value of the stride) along the horizontal direction, thus completing the task of extracting specific features from the image.

[0057] 2. Loss Function

[0058] During the training of a neural network, since we hope that the output of the deep neural network is as close as possible to the value we truly want to predict, we can compare the predicted value of the current network with the true target value, and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the true target value or a value very close to the true target value. Therefore, we need to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then, the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0059] 3. Backpropagation Algorithm

[0060] A neural network can use the backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial model during training, making the reconstruction error loss of the model smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial model parameters are updated by backpropagating the error loss information, thereby converging the error loss. The backpropagation algorithm is a reverse propagation movement dominated by the error loss, aiming to obtain the optimal model parameters, such as the weight matrix.

[0061] 4. Pixel Value

[0062] The pixel value of an image can be a red-green-blue (RGB) color value, and the pixel value can be a long integer representing a color. For example, the pixel value is 256*Red + 100*Green + 76*Blue, where Blue represents the blue component, Green represents the green component, and Red represents the red component. Among the respective color components, the smaller the value, the lower the brightness, and the larger the value, the higher the brightness. For a grayscale image, the pixel value can be a grayscale value.

[0063] Video frame interpolation is to predict and insert one or more intermediate frames between the front and back frames of the original video, shortening the display duration span between frames, so as to obtain a video with a higher frame rate.

[0064] In actual application scenarios, frame interpolation technology can be applied to scenarios such as videos and animations (i.e., dynamic pictures). As Figure 1 shown, if the original video or animation is 9 frames, when applying frame interpolation technology, there may be two situations:

[0065] Situation 1: The 9-frame images rendered by the graphics processing unit (GPU) in the original video or animation can be reduced to 5 frames. For the 5-frame images, the images between every two images can be interpolated using frame interpolation technology, and the interpolation operation is executed on the NPU, which can reduce the power consumption of generating intermediate images, as Figure 1 shown in (a) of

[0066] Situation 2: The original video or animation can be directly interpolated, that is, one frame is inserted between two adjacent frames, so that the original 9 frames become 17 frames, improving the smoothness of the video or animation, as Figure 1 shown in (b) of

[0067] Currently, frame interpolation technology can be implemented through traditional interpolation algorithms or through neural networks. For the implementation method of obtaining intermediate frame images through a frame interpolation model based on a neural network, two frame images in the video can be input into a frame interpolation model including a convolutional layer, a pooling layer, and an upsampling layer (such as nearest neighbor interpolation and / or bilinear interpolation), and the intermediate frame image between the two frame images is output.

[0068] Each time frame interpolation is performed, 2 images can be input into the frame interpolation model, and through the processing of the convolutional layer, the pooling layer, and the interpolation layer, the intermediate frame image is obtained, so that the intermediate frame image can be inserted between the aforementioned 2 images.

[0069] Combined with Figure 2It can be seen that in the implementation of obtaining the intermediate frame image by the neural network-based frame interpolation model, the frame interpolation model includes a convolutional layer, a pooling layer, and an interpolation layer, so that there are many non-convolution operations in the processing of the frame interpolation model, such as pooling operations, size adjustment operations, interpolation operations, etc. The number of parameters and the amount of calculation of non-convolution operations are relatively large, and the execution time on the NPU may be close to or exceed the execution time of one convolutional layer. That is to say, the execution time of non-convolution operations on the NPU is relatively long. In this way, the time to obtain the intermediate frame image from two frame images will be relatively long. When playing a video, the picture update is not timely, resulting in low smoothness of the video picture playback.

[0070] In the face of the above problems, the present application can provide a video frame interpolation method, an electronic device, a chip system, a computer-readable storage medium, and a computer program product, which can cancel non-convolution operations, so that the frame interpolation model includes a first module and a second module. The first module includes a convolutional layer, and the second module includes a segmentation module and a calculation module. When performing frame interpolation processing, two images in the target video can be input into the first module, and pixel features are extracted through the aforementioned convolutional layer to obtain a feature map, and the feature map is input into the second module. The segmentation module divides the aforementioned feature map into two sub-feature maps corresponding to the two images respectively, and the calculation module determines the difference between the two sub-feature maps, so as to determine the motion of pixels in the two images in the time series according to the aforementioned difference. Through the processing of the second module, an intermediate image for insertion between the two images can be output. In this way, through a frame interpolation model that does not include non-convolution operations such as pooling operations and interpolation-based upsampling operations, the intermediate image between the two images can be obtained, which can shorten the time for the frame interpolation model to output the intermediate image and further ensure the smoothness of the target video playback.

[0071] Among them, the above-mentioned electronic device can be an electronic device with neural network processor hardware and corresponding software support.

[0072] Among them, the above-mentioned electronic device can be a mobile phone, a tablet computer, a vehicle-mounted device, a notebook computer, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart car, a smart TV, a robot, etc.

[0073] It should be noted that in some possible implementation manners, the electronic device may also be referred to as a terminal device (station), a user equipment (UE), etc., and the embodiments of the present application do not limit this.

[0074] For the sake of convenience of description, Figure 3Among them, the electronic device 100 is taken as a mobile phone as an example for illustration.

[0075] As Figure 3 shown, in some embodiments, the electronic device 100 may include a processor 101, a communication module 102, a display screen 103, a camera 104, a sensor 105, an internal memory 106, a USB interface 107, an external memory interface 108, a charging management module 109, a power management module 110, and a battery 111, etc.

[0076] Among them, the processor 101 may include one or more processing units. For example, the processor 101 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video stream codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors 101.

[0077] The controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0078] A memory may also be provided in the processor 101 for storing instructions and data.

[0079] The neural-network processing unit is a hardware component specifically designed to accelerate artificial intelligence (AI) and machine learning tasks. The neural-network processing unit is mainly used to process computations related to neural networks, such as image recognition, speech recognition, natural language processing, image interpolation processing, etc.

[0080] The communication module 102 may include antenna 1 and antenna 2, a mobile communication module, and / or a wireless communication module.

[0081] Optionally, the electronic device 100 may further include peripheral devices, such as a mouse, a button, an indicator light, a keyboard, a speaker, a microphone, etc.

[0082] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100.

[0083] In some other embodiments, the electronic device 100 may include more or fewer components than shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figures may be implemented in hardware, software, or a combination of software and hardware.

[0084] Please refer to Figure 4 , which is a schematic diagram of the software architecture of an electronic device provided by an embodiment of this application. When the video frame interpolation method provided by an embodiment of this application is applied to Figure 3 the electronic device 100 shown, the software in the electronic device 100 can be divided into an application layer 201, an application framework layer (FWK) 202, a hardware abstraction layer (HAL) 203, and a driver layer 204 as shown in Figure 4 .

[0085] Multiple applications can be installed in the application layer 201.

[0086] The application framework layer 202 provides a set of basic functions and services for the application layer 201 to call and use.

[0087] The hardware abstraction layer 203 is a software layer located between the operating system kernel and the hardware circuit, and is usually used to abstract the hardware to achieve the interaction between the operating system and the hardware circuit at the logical layer.

[0088] Multiple drivers for driving the hardware to work can be installed in the driver layer 204.

[0089] It should be noted that the application layer 201, the application framework layer 202, the hardware abstraction layer 203, and the driver layer 204 may also include other content, which is not specifically limited here.

[0090] Among them, the frame interpolation model is involved in the embodiments of this application. The frame interpolation model is obtained by training a neural network through a frame interpolation model generation process. The method provided by this application will be described below from the model training side and the model application side of the frame interpolation model:

[0091] The interpolation model generation process provided by the embodiments of the present application involves computer vision processing, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs intelligent information modeling, extraction, preprocessing, training, etc. on training data (such as multiple training image groups described later in this application), and finally obtains a trained interpolation model. Moreover, the video interpolation method provided by the embodiments of the present application can use the above-trained interpolation model to input input data (such as two images in the target video in this application) into the trained interpolation model to obtain output data (such as the intermediate image in this application). It should be noted that the interpolation model generation process and the video interpolation method provided by the embodiments of the present application are inventions based on the same concept, and can also be understood as two parts of a system or two stages of an overall process: such as the model training stage and the model application stage.

[0092] Based on the above description, first, the training process of the interpolation model involved in the video interpolation method of the embodiments of the present application will be elaborated in detail below.

[0093] Specifically, the system architecture involved in the generation process of the interpolation model includes a first sample module and a second sample module, and the generation process of the interpolation model includes:

[0094] Step 11: Obtain multiple training image groups. Each training image group in the multiple training image groups includes three consecutive images in a time series, and the second image in each training image group is the label intermediate image corresponding to each training image group.

[0095] Step 12: Input the first image and the third image in each training image group into the first sample module of the neural network to obtain a sample feature map. The first sample module is used to extract pixel features of the first image and the third image through the convolutional layer included in the first sample module. The sample feature map includes the pixel features of the first image and the pixel features of the third image.

[0096] Step 13: Input the sample feature map into the second sample module of the neural network to obtain the estimated intermediate image corresponding to each training image group. The second sample module includes a first sample segmentation module and a first sample calculation module. The first sample segmentation module is used to segment the sample feature map into a first sub-sample feature map and a second sub-sample feature map. The first sample calculation module is used to determine the sample difference data between the first sub-sample feature map and the second sub-sample feature map, and determine the estimated intermediate image according to the sample difference data. The sample difference data is used to represent the pixel movement situation from the first image to the third image in the time series.

[0097] Step 14: Train the neural network according to the difference between the label intermediate image and the estimated intermediate image to obtain an interpolation model.

[0098] Combined with the above description, during the training process, based on the first and third images in multiple sets of three consecutive images, the predicted intermediate image can be obtained. Then, according to the difference between the predicted intermediate image and the second image in the three consecutive images, that is, the loss between the predicted intermediate image and the second image in the three consecutive images, through the backpropagation method, the neural network can update the model parameters of the neural network according to the aforementioned loss, thereby obtaining the interpolation model.

[0099] Based on this, the obtained interpolation model can predict the intermediate image that can be inserted between the two images in the target video.

[0100] It should be noted that the neural network used in the generation process of the interpolation model in the embodiments of the present application is usually the same as the structure of the interpolation model used in the video interpolation method in the embodiments of the present application. The structure of the model will not be specifically described here. For details, please refer to the following description.

[0101] In addition, the generation process of the interpolation model in the present application and the video interpolation method in the embodiments of the present application can be executed by the same device. For example, both are executed by Figure 3 the electronic device 100 shown; the generation process of the interpolation model in the present application and the video interpolation method in the embodiments of the present application can also be executed by different devices. For example, the generation process of the interpolation model is executed by a training device (such as the training device can be a server).

[0102] Based on the above description, the video interpolation method in the application stage of the interpolation model in the embodiments of the present application will be elaborated in detail below in combination with the accompanying drawings and application scenarios.

[0103] Among them, the video interpolation method provided in the embodiments of the present application is applied to an electronic device. The electronic device includes an interpolation model. The interpolation model includes a first module and a second module. The first module includes a convolutional layer, and the second module includes a first segmentation module and a first calculation module. Among them, the aforementioned electronic device can be Figure 3 the electronic device 100 shown.

[0104] Please refer to Figure 5 , Figure 5 which shows a schematic flowchart of the video interpolation method provided in an embodiment of the present application.

[0105] As Figure 5 shown, the video interpolation method provided in the present application may include:

[0106] S301. Obtain a first image and a second image from the target video. The first image is a frame in the target video that is earlier than the second image in the time series.

[0107] Among them, there are usually two situations for the first image and the second image.

[0108] In some embodiments, the first image and the second image are two adjacent frames of images in the target video.

[0109] Generally, the electronic device can perform frame interpolation between two adjacent frames of images in the target video, that is, insert one frame of image between every two adjacent frames in the target video, which can make the number of frames of the target video more, so that the updated target video obtained. When the electronic device plays the updated target video, compared with the original target video, the playing smoothness of the updated target video is higher. The embodiments of the present application are applicable to scenarios where the playing smoothness of the target video needs to be improved.

[0110] For example, the target video includes a total of 9 frames of images. The first image is the first frame among the 9 frames of images, and the second image is the second frame among the 9 frames of images.

[0111] In other embodiments, the first image and the second image are two frames of images corresponding to N frames apart in the target video, where N is a positive integer greater than or equal to 1.

[0112] Generally, the electronic device can perform frame extraction on the target video every N frames of images, and then perform frame interpolation on the remaining two frames of images corresponding to every N frames to obtain the same number of images as the original target video. Since the frame interpolation process is executed on the NPU, the generation of a smaller number of intermediate images can reduce the power consumption of the electronic device. The embodiments of the present application are applicable to scenarios where the power consumption of the electronic device needs to be reduced.

[0113] For example, the target video includes a total of 9 frames of images. The second frame, the fourth frame, the sixth frame, and the eighth frame are extracted from the target video. The remaining images include the first frame, the third frame, the fifth frame, the seventh frame, and the ninth frame. The value of N is 1. The first image is the first frame among the 9 frames of images, and the second image is the third frame among the 9 frames of images.

[0114] S302. Input the first image and the second image into the first module to obtain a first feature map. The first module is used to extract pixel features of the first image and the second image through a convolutional layer. The first feature map includes pixel features of the first image and pixel features of the second image.

[0115] Since the first module includes a convolutional layer, after the electronic device inputs the first image and the second image into the first module, the electronic device can extract pixel features of the first image and the second image through the convolutional layer, thereby obtaining a first feature map. The first feature map includes pixel features of the first image and pixel features of the second image.

[0116] In some embodiments, the first module includes a first sub-layer and a second sub-layer. The first sub-layer includes a first convolutional layer and a second activation function layer, and the second sub-layer includes a second convolutional layer. Inputting the spliced image into the first module to obtain a first feature map includes:

[0117] Input the spliced image into the first sub-layer to obtain a first intermediate feature map. The first convolutional layer in the first sub-layer is used to extract pixel features of the spliced image to obtain a second intermediate feature map. The second activation function layer is used to perform a non-linear transformation on the second intermediate feature map to obtain the first intermediate feature map. During the process of the first convolutional layer extracting pixel features of the spliced image, the convolution stride adopted is related to the target output duration of the third image. The convolution stride is an even number greater than or equal to 2, and the size of the first intermediate feature map is related to the convolution stride. Input the first intermediate feature map into the second sub-layer to obtain a first feature map. The second convolutional layer in the second sub-layer is used to adjust the size of the first intermediate feature map to the same size as the spliced image.

[0118] It should be understood that the convolution stride can be adjusted. The smaller the convolution stride, the greater the overall computational amount of the first module, the longer the overall execution duration of the first module, and the longer the output duration of the third image. The larger the convolution stride, the smaller the overall computational amount of the first module, the shorter the overall execution duration of the first module, and the shorter the output duration of the third image.

[0119] Based on this, by setting the size of the convolution stride to a value related to the target output duration of the third image, the output duration of the third image can be guaranteed.

[0120] Among them, the convolution stride can be determined according to the actual situation. For example, the convolution stride can be set to 2.

[0121] In addition, in order to further reduce the overall computational amount of the first module, the convolution stride is set to an even number greater than or equal to 2. Considering that when the length and / or width of the spliced image is odd, the spliced image (i.e., the first image and the second image) cannot be processed by the convolutional layer corresponding to the even convolution stride, the spliced image can be first expanded and then further processed by the first convolutional layer.

[0122] Specifically, the convolution stride is an even number greater than or equal to 2. The first module further includes a padding layer and a cropping layer. Inputting the spliced image into the first sub-layer to obtain a first intermediate feature map includes:

[0123] Input the spliced image into the padding layer, and determine whether the length and width of the spliced image are odd respectively through the padding layer. If the length and / or width of the spliced image is odd, perform a first expansion on the pixels of the spliced image to obtain an updated spliced image. The first expansion is used to make the length and width of the spliced image both even. Input the updated spliced image into the first sub-layer to obtain a first intermediate feature map.

[0124] Moreover, inputting the first intermediate feature map into the second sub-layer to obtain a first feature map, including:

[0125] Inputting the first intermediate feature map into the second sub-layer to obtain a second intermediate feature map. The second convolutional layer (which can also be called a transposed convolutional layer) of the second sub-layer is used to adjust the size of the first intermediate feature map to the same size as the updated spliced image; inputting the second intermediate feature map into the cropping layer to obtain a first feature map, and the cropping layer is used to adjust the size of the second intermediate feature map to the same size as the spliced image.

[0126] Among them, as Figure 6 shown in the first module with a convolutional stride of 2, the first module includes a padding layer, a first sub-layer (a convolutional layer and an activation function layer, the number of first sub-layers is P, and the value of P here is 2), a second sub-layer (a convolutional layer), and a cropping layer. If a convolutional layer with a convolutional stride of 2 is adopted, when the length and width of the input of the subsequent layer become 1 / 2 of the original, the computational cost of the subsequent convolutional layer can become on the order of 1 / 4. When the original input length and width are odd, the output size may be smaller than the original output. Therefore, a padding layer is added before the convolutional layer with a convolutional stride equal to 2 to zero-pad the size of the image input to the current layer to an even number. Finally, instead of using an interpolation-based upsampling layer, a transposed convolution operation is adopted, and for the final output, the image corresponding to the original odd width and height is cropped. In this way, a stacking method of pure convolutional layers can be achieved.

[0127] Among them, if a convolutional layer with a convolutional stride of 2 is adopted, in order to maintain the performance of the network, in the existing solutions, a pooling layer is first used to reduce the input size of the subsequent convolutional layer, and then additional convolutional layers are added for preprocessing or postprocessing, and additional interpolation-based upsampling layer operations are required for the subsequent layer to maintain the same size of the output image as the original input image. This may increase a large amount of computational overhead and latency; while in the embodiments of the present application, here an interpolation-based upsampling layer is replaced, and a convolutional layer with an even stride replacing the pooling layer is adopted, and then a transposed convolution operation is used to make the output image have the same size as the original input image. In this way, it can be further ensured that the first module is a pure convolutional module. Using a pure convolutional module as a feature extraction unit can reduce the overall number of model parameters and computational cost.

[0128] For example, if the dimension of the current input stitching image i is [b, c1, n, m], the length of the stitching image i is n, and n is odd, then it is necessary to pad after the rightmost column of the stitching image i. After padding, the length is n + 1, and then it is sent into a convolutional layer with a convolutional stride of 2 for convolution processing. Finally, the dimension of the output feature map after the transpose operation is [b, c2, n + 1, m], and then it is intercepted by the interception layer as the final output of the first module.

[0129] Based on the above description, Figure 6 For the first module with a convolutional stride of 2, compared with the first module with a convolutional stride of 1, the output time of the first feature map of the first module with a convolutional stride of 2 is shorter.

[0130] Combined with Figure 6 the structure of the first module shown, and the above description, when the length and / or width of the stitching image is odd, the stitching image can be first expanded by padding with 0s, that is, padding, so that the length and width of the stitching image are both even, obtaining an updated stitching image, and then input into the first convolutional layer (the convolutional stride can be 2) for processing. After processing through the first convolutional layer and the second activation function layer, the size of the first intermediate feature map obtained is a smaller size. Based on this, the second convolutional layer of the second sublayer adjusts the size of the first intermediate feature map to be the same as the updated stitching image through a transposed convolutional operation (ConvTranspose) with a convolutional stride of 2, obtaining a second intermediate feature map, and then the second intermediate feature map is intercepted by the interception layer to obtain the first feature map. In this way, the size of the first feature map is the same as the size of the stitching image, and the situation where the size of the third image obtained from the first feature map is different from the sizes of the first image and the second image can be avoided.

[0131] Among them, the number of the first sublayers is P, and the value of P is related to the target output duration of the third image. Each first sublayer is used for extracting pixel features. The application does not make a specific limitation on the value of P.

[0132] When the value of P is 2, the specific structure of the first module can be referred to Figure 6 the structure shown, and it can be seen that the first module includes a padding layer, two first sublayers (a convolutional layer + a second activation function layer), a second sublayer (a second convolutional layer), and an interception layer. Each first sublayer includes a convolutional layer and an activation function layer.

[0133] It should be understood that the larger the value of P, the greater the overall computational amount of the first module, the longer the overall execution duration of the first module, and the longer the output duration of the third image. The smaller the value of P, the smaller the overall computational amount of the first module, the shorter the overall execution duration of the first module, and the shorter the output duration of the third image.

[0134] Based on this, if it is desired to shorten the output duration of the third image, it is necessary to shorten the overall execution duration of the first module. The value of P can be set to a smaller value. Specifically, the value of P can be set according to the target output duration of the third image. The shorter the target output duration of the third image, the smaller the value of P set, that is, the fewer the number of the first sub-layers.

[0135] In some embodiments, the first convolutional layer includes M output channels, the second convolutional layer includes R output channels. The values of M and R are related to the target output duration of the third image. The first intermediate feature map is a feature map with M dimensions, and the first feature map is a feature map with R dimensions.

[0136] In some embodiments, the second activation function layer is a rectified linear unit (ReLU) function.

[0137] In the embodiments of the present application, the calculation process of the ReLU function is simple. By using the ReLU function, on the premise of ensuring the running performance of the interpolation model, the forward inference time of the interpolation model can be shortened, that is, the time required from inputting the first image and the second image to outputting the third image can be shortened.

[0138] Among them, the input channels of the first convolutional layer in the first sub-layer are determined according to the input image. For example, it can be three RGB channels, and the number of input channels of each other convolutional layer is the same as the number of output channels of the previous convolutional layer.

[0139] For example, the number of input channels of the first convolutional layer is 3, the number of output channels of the first convolutional layer is 1, and the number of input channels of the second convolutional layer is 1.

[0140] Among them, the output channels of each convolutional layer can be adjusted. After the number of output channels of a single convolutional layer reaches a certain threshold, at this time, the larger the number of output channels, the greater the amount of computation that the convolutional layer needs to perform, and the smaller the number of output channels, the smaller the amount of computation that the convolutional layer needs to perform. It can be seen that the size of the number of output channels can affect the amount of computation that the convolutional layer needs to perform. The present application does not make specific limitations on the number of output channels of each convolutional layer.

[0141] Moreover, after the number of output channels of a single convolutional layer reaches a certain threshold, at this time, the greater the amount of computation that each convolutional layer needs to perform, the longer the overall execution duration of the first module, and the longer the output duration of the third image. The smaller the amount of computation that each convolutional layer needs to perform, the shorter the overall execution duration of the first module, and the longer the output duration of the third image.

[0142] Based on this, to shorten the output duration of the third image, it is necessary to shorten the overall execution duration of the first module. The number of output channels of the first convolutional layer and / or the second convolutional layer can be set to a smaller value. Specifically, the number of output channels of the first convolutional layer and / or the second convolutional layer can be set according to the target output duration of the third image; the shorter the target output duration of the third image, within a certain range, the smaller the value of the number of output channels of the first convolutional layer and / or the second convolutional layer is set.

[0143] It should be noted that the frame interpolation model runs on the neural network processor of the electronic device. Considering the performance of the neural network processor, the number of output channels of each convolutional layer cannot be adjusted without limit, and only needs to be adjusted to the number of output channels corresponding to the target output duration of the third image.

[0144] S303. Input the first feature map into the second module to obtain a third image. The first segmentation module in the second module is used to segment the first feature map into a first sub-feature map and a second sub-feature map. The first calculation module is used to determine the difference data between the first sub-feature map and the second sub-feature map. The difference data is used to represent the movement of pixels in the time series from the first image to the second image. The third image is an intermediate image to be inserted between the first image and the second image.

[0145] The second module includes a first segmentation module. After the electronic device inputs the first feature map into the second module, the first segmentation module in the second module can segment the first feature map into a first sub-feature map and a second sub-feature map, thereby preparing data for the electronic device to determine the third image that can be inserted between the first image and the second image according to the difference between the first sub-feature map and the second sub-feature map.

[0146] The second module also includes a first calculation module. After the first segmentation module segments the first feature map into a first sub-feature map and a second sub-feature map, the first calculation module can calculate the difference data between the first sub-feature map and the second sub-feature map. Since the difference data is used to represent the movement of pixels in the time series between the first image and the second image, the first calculation module can predict the third image according to the difference data.

[0147] In some embodiments, the difference data includes first difference data and second difference data. The first difference data is used to represent the movement of pixels in the time series from the first image to the second image, and the second difference data is used to represent the movement of pixels in the time series from the second image to the first image.

[0148] The video frame interpolation method of the present application, the frame interpolation model includes a first module and a second module. The first module includes a convolutional layer, and the second module includes a first segmentation module and a first calculation module. When performing frame interpolation processing, the first image and the second image in the target video can be input into the first module, and pixel features are extracted through the aforementioned convolutional layer to obtain a first feature map. The first feature map is input into the second module, and the first segmentation module in the second module segments the first feature map to obtain a first sub-feature map and a second sub-feature map. Then, the first calculation module in the second module can determine the difference data between the first sub-feature map and the second sub-feature map through subtraction calculation. Since the difference data is used to represent the pixel motion situation of the first image and the second image in the time series, the first calculation module can obtain a third image for insertion between the first image and the second image. In this way, the constructed first module is a pure convolutional module, and as an extraction unit, it can reduce the overall number of parameters and computational amount of the model. Moreover, the electronic device can obtain the third image without using a frame interpolation model including non-convolutional operations such as pooling operations and interpolation-based upsampling operations. Instead, after obtaining the first feature map, the first feature map is segmented into two sub-feature maps, and the difference data obtained through the subtraction operation between the two sub-feature maps is used to predict the third image for insertion between the first image and the second image, thereby being able to shorten the output duration of the third image by the frame interpolation model. When playing the video, the picture can be updated in time, and further, the video playback smoothness of the target video can be ensured.

[0149] Furthermore, the convolutional stride is adjustable. The smaller the convolutional stride, the greater the overall computational amount of the first module, the longer the overall execution duration of the first module, and the longer the output duration of the third image. The larger the convolutional stride, the smaller the overall computational amount of the first module, the shorter the overall execution duration of the first module, and the shorter the output duration of the third image. By setting the size of the convolutional stride to a value related to the target output duration of the third image, the output duration of the third image can be ensured. And, the convolutional stride can be further set to an even number. Considering that when the length and / or width of the spliced image is odd, the spliced image cannot be processed by the convolutional layer corresponding to the even convolutional stride, when the length and / or width of the spliced image is odd, first, by padding with 0, the length and width of the spliced image are made even, and then processed through the convolutional layer. Subsequently, by padding with 0 and cropping, the final output first feature map can be made the same size as the spliced image, which can avoid the situation where the size of the third image obtained from the first feature map is different from the sizes of the first image and the second image.

[0150] Based on the above Figure 5 description of the illustrated embodiment, the frame interpolation model further includes a third module, a fourth module, a fifth module, a sixth module, a seventh module, and a first activation function.

[0151] The third module includes a convolutional layer, the fourth module includes a second segmentation module and a second calculation module. The fifth module includes a convolutional layer, the sixth module includes a transformation module, a third segmentation module and a third calculation module. The seventh module includes a convolutional layer, and a first activation function is used to convert the feature map into an image with pixel values maintained within a preset range.

[0152] Next, in combination with Figure 7 , the video frame interpolation method of this application will be introduced in detail.

[0153] Please refer to Figure 7 , Figure 7 which shows a schematic flowchart of the video frame interpolation method provided by an embodiment of this application.

[0154] As Figure 7 shown, the video frame interpolation method provided by this application may include:

[0155] S401. Obtain a first image and a second image from the target video, where the first image is a frame in the target video that is earlier than the second image in the time series.

[0156] Among them, the implementation of S401 is similar to that of S301 in the Figure 5 shown embodiment, and will not be elaborated here.

[0157] S402. Stitch the first image and the second image to obtain a stitched image.

[0158] Among them, the stitching of the first image and the second image is stitching in the channel direction, specifically, the first image and the second image are stitched in a front-back stacking manner.

[0159] For example, the pixel point at position A of the first image is (3, 255, 255), and the pixel point at position A of the second image is (3, 255, 255). After stitching the first image and the second image, the pixel point at position A of the first image is (3, 255, 255), and the value of the pixel point at position A of the second image is (6, 255, 255).

[0160] S403. Input the stitched image into the first module to obtain a first feature map. The first module is used to extract pixel features of the first image and the second image through a convolutional layer, and the first feature map includes pixel features of the first image and pixel features of the second image.

[0161] Among them, the implementation of S403 is similar to that of S302 in the Figure 5 shown embodiment, and will not be elaborated here.

[0162] S404. Input the first feature map and the spliced image into the second module to obtain a second feature map. The first segmentation module in the second module is used to segment the first feature map into a first sub-feature map and a second sub-feature map. The first calculation module is used to determine first difference data and second difference data between the first sub-feature map and the second sub-feature map, and to splice the first difference data, the second difference data, and the spliced image. The first difference data is used to represent the movement of pixels in the time series from the first image to the second image, and the second difference data is used to represent the movement of pixels in the time series from the second image to the first image. The second feature map is a feature map obtained by splicing the first difference data, the second difference data, and the spliced image.

[0163] Combined with Figure 8 , specifically, after the electronic device inputs the first feature map and the spliced image into the second module, the first segmentation (split) module in the second module can segment the first feature map to obtain a first sub-feature map and a second sub-feature map. Taking the first feature map as a feature map with 4 dimensions as an example, from Figure 8 it can be seen that the first segmentation module can segment between the first two dimensions and the last two dimensions of the first feature map to obtain a first sub-feature map and a second sub-feature map ( Figure 8 in [:,:2,:,:] and[:,2:4,:,:], and a comma is used to separate between the two dimensions).

[0164] Next, the first calculation module in the second module can subtract the first sub-feature map and the second sub-feature map to obtain first difference data ( Figure 8 Flow in _ 0 _ 0), subtract the constant 0 from the first difference data to obtain second difference data ( Figure 8 Flow in _ 0 _ 1), and the first difference data and the second difference data are opposite difference data.

[0165] After obtaining the first difference data, the first calculation module subtracts the constant 0 from the first difference data to obtain second difference data, which is equivalent to obtaining two-way flow information, that is, the movement of pixels in the time series from the first image to the second image, and the movement of pixels in the time series from the second image to the first image. Thus, the second module can obtain a second feature map according to the two-way flow information. The second feature map carries two-way flow information, which can ensure the accuracy of the subsequent obtained third image.

[0166] Finally, in some embodiments, the first calculation module in the second module splices the first difference data and the second difference data to obtain a second feature map.

[0167] In some other embodiments, the first computing module splices the first difference data, the second difference data, and the spliced image to obtain a second feature map. Therefore, the original input information is included in the spliced second feature map. In this way, the subsequent intermediate layers of the interpolation model can learn the original input information, ensuring that the interpolation model does not lose important feature information, thereby improving the performance of the model.

[0168] The first computing module subtracts the first sub-feature map from the second sub-feature map to obtain r _ 1, and r _ 1 can also be multiplied by a constant C (multiply) to obtain the first difference data ( Figure 8 Flow in _ 0 _ 0).

[0169] The first computing module multiplies r _ 1 by the constant C, which can play a role in mutually adjusting the contribution degrees of the first difference data and the second difference data.

[0170] It should be noted that S405 - S412 are optional steps. When S405 - S412 are not executed, after the first module outputs the first feature map, the electronic device can directly input the first feature map into the second module to obtain a third image.

[0171] S405: Input the second feature map into the third module to obtain a third feature map. The third module is used for feature extraction of the second feature map.

[0172] Among them, the third module includes a convolutional layer, and the structure of the third module can be the same as or different from that of the first module.

[0173] When the structure of the third module is the same as that of the first module, if the first module includes a first sub-layer and a second sub-layer, the first sub-layer includes a first convolutional layer and a second activation function layer, and the second sub-layer includes a second convolutional layer, the first module includes one first sub-layer and one second sub-layer; the third module includes one first sub-layer and one second sub-layer.

[0174] When the structure of the third module is different from that of the first module, if the first module includes a first sub-layer and a second sub-layer, the first sub-layer includes a first convolutional layer and a second activation function layer, and the second sub-layer includes a second convolutional layer, the first module includes one first sub-layer and one second sub-layer; for example, the third module can include two first sub-layers and one second sub-layer; for example, the third module can include three first sub-layers and one second sub-layer.

[0175] This application does not make specific limitations on the structure of the third module.

[0176] The third module is used to extract features from the second feature map. The third feature map obtained by inputting the second feature map into the third module includes the features of the second feature map.

[0177] S406. Input the third feature map, the first difference data, the second difference data, the first image, the second image, and the spliced image into the fourth module to obtain a fourth feature map. The second segmentation module in the fourth module is used to segment the third feature map into a third sub-feature map and a fourth sub-feature map. The second calculation module is used to fuse the first difference data and the third sub-feature map to obtain a first fused feature map, fuse the second difference data and the fourth sub-feature map to obtain a second fused feature map, perform a warping transformation on the first image according to the first fused feature map to obtain a first transformed feature map, and perform a warping transformation on the second image according to the second fused feature map to obtain a second transformed feature map. The fourth feature map is obtained by splicing the first transformed feature map, the second transformed feature map, and the spliced image. The accuracies of the first transformed feature map and the second transformed feature map for predicting the third image are higher than the accuracies of the first difference data and the second difference data for predicting the third image.

[0178] Combined with Figure 9 , specifically, after the electronic device inputs the third feature map, the first difference data, the second difference data, and the spliced image into the fourth module, the second segmentation module in the fourth module can segment the third feature map to obtain a third sub-feature map and a fourth sub-feature map. Taking the feature map with 4 dimensions of the third feature map as an example, from Figure 9 it can be seen that the second segmentation module can perform segmentation between the first two dimensions and the last two dimensions of the third feature map in the channel direction to obtain a third sub-feature map and a fourth sub-feature map ( Figure 9 [:,:2,:,:] and[:,2:4,:,:] in ).

[0179] Next, the second calculation module in the fourth module can fuse (add) the first difference data and the third sub-feature map to obtain a first fused feature map, and fuse the second difference data and the fourth sub-feature map to obtain a second fused feature map.

[0180] Finally, in some embodiments, the second calculation module can also splice the first fused feature map and the second fused feature map to obtain a fourth feature map.

[0181] In other embodiments, the second calculation module can splice the first fused feature map, the second fused feature map, and the spliced image to obtain a fourth feature map. Therefore, the fourth feature map obtained by splicing includes the original input information. In this way, the subsequent intermediate layers of the interpolation model can learn the original input information, ensuring that the interpolation model does not lose important feature information, thereby ensuring the performance of the model.

[0182] Optionally, the second computing module may further perform a warp transformation on the first image according to the first fused feature map to obtain a first transformed feature map, and perform a warp transformation on the second image according to the second fused feature map to obtain a second transformed feature map.

[0183] Among them, the warp transformation can be understood as an image transformation, which realizes the deformation of the image by changing the positions of the pixels in the image; through the warp transformation, the interpolation model can better learn the key features in the first image and the second image, such as perspective transformation and pose transformation.

[0184] In some embodiments, the second computing module may further splice the first transformed feature map and the second transformed feature map to obtain a fourth feature map.

[0185] In some other embodiments, the second computing module may splice the first transformed feature map, the second transformed feature map, and the spliced image to obtain a fourth feature map. Therefore, the fourth feature map obtained by splicing includes the original input information. In this way, the subsequent intermediate layers of the interpolation model can learn the original input information, as well as the key features such as perspective transformation and pose transformation in the first image and the second image, ensuring that the interpolation model does not lose important feature information, and thus ensuring the performance of the model.

[0186] Based on this, the first transformed feature map is equivalent to the optimized first difference data, and the second transformed feature map is equivalent to the optimized second difference data. In this way, the accuracy of predicting the third image by the first transformed feature map and the second transformed feature map is higher than the accuracy of predicting the third image by the first difference data and the second difference data.

[0187] It should be noted that S407-S412 are optional steps. When S407-S412 are not executed, after the third module outputs the third feature map, the electronic device may directly input the third feature map, the first difference data, and the second difference data into the fourth module to obtain the third image.

[0188] S407: Input the fourth feature map into the fifth module to obtain a fifth feature map, where the fifth module is used to extract features from the fourth feature map.

[0189] Among them, the fifth module includes a convolutional layer. The structure of the fifth module may be the same as that of the first module and the third module, or may be different from at least one of the first module and the third module.

[0190] When the structure of the fifth module is the same as that of the first module and the third module, if the first module includes a first sub-layer and a second sub-layer, the first sub-layer includes a first convolutional layer and a second activation function layer, the second sub-layer includes a second convolutional layer, the first module includes one first sub-layer and one second sub-layer; the third module includes one first sub-layer and one second sub-layer; the fifth module includes one first sub-layer and one second sub-layer.

[0191] When the structure of the fifth module is different from that of the first module and the same as that of the third module, if the first module includes a first sub-layer and a second sub-layer, the first sub-layer includes a first convolutional layer and a second activation function layer, the second sub-layer includes a second convolutional layer, the first module includes one first sub-layer and one second sub-layer; for example, the third module includes two first sub-layers and one second sub-layer, and the fifth module includes two first sub-layers and one second sub-layer; for example, the third module includes three first sub-layers and one second sub-layer, and the fifth module may include three first sub-layers and one second sub-layer.

[0192] When the structure of the fifth module is different from that of the first module and also different from that of the third module, if the first module includes a first sub-layer and a second sub-layer, the first sub-layer includes a first convolutional layer and a second activation function layer, the second sub-layer includes a second convolutional layer, the first module includes one first sub-layer and one second sub-layer; for example, the third module includes two first sub-layers and one second sub-layer, and the fifth module includes four first sub-layers and one second sub-layer.

[0193] This application does not specifically limit the structure of the fifth module.

[0194] The fifth module is used for feature extraction of the fourth feature map, and the fifth feature map obtained by inputting the fourth feature map into the fifth module includes the image features of the fourth feature map.

[0195] S408. Input the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module to obtain a seventh feature map. The transformation module in the sixth module is used to normalize each feature data in the fifth feature map to obtain a sixth feature map. The third segmentation module is used to segment the sixth feature map into a fifth sub-feature map and a sixth sub-feature map. The third calculation module is used to generate a seventh feature map including target information according to the fifth sub-feature map, the sixth sub-feature map, the first transformed feature map, and the second transformed feature map. The target information includes light and shadow information and / or occlusion information. The target information included in the seventh feature map is the target information included in the first image and the second image.

[0196] Combine Figure 10, specifically, after the electronic device inputs the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module, the transformation module in the sixth module can normalize each feature data in the sixth feature map through a logical function (which can be called the sigmoid function), that is, map each feature data in the sixth feature map to the interval [0, 1] to obtain the sixth feature map.

[0197] Based on the above description, the transformation module normalizes each feature data in the sixth feature map through a logical function, and the obtained sixth feature map is a feature map including multiple weight coefficients, so that target information (light and shadow information and / or occlusion information) can be introduced into the seventh feature map.

[0198] Next, the third segmentation module in the sixth module can segment the sixth feature map to obtain a fifth sub-feature map and a sixth sub-feature map. Taking the sixth feature map as a feature map with two dimensions as an example, from Figure 10 it can be seen that the third segmentation module can segment between the two dimensions of the sixth feature map to obtain a fifth sub-feature map and a sixth sub-feature map ( Figure 10 in [:,:1,:,:] and[:,1:2,:,:].

[0199] Next, the third calculation module in the sixth module can fuse (add) the fifth sub-feature map and the sixth sub-feature map to obtain a fourth fused feature map. The third calculation module can also multiply the fifth sub-feature map and the first transformed feature map to obtain a first enhanced feature map. The third calculation module can also multiply the sixth sub-feature map and the second transformed feature map to obtain a second enhanced feature map, fuse (add) the first enhanced feature map and the second enhanced feature map to obtain a fifth fused feature map, and divide the fifth fused feature map by the fourth fused feature map to obtain a seventh feature map.

[0200] Based on the above description, the sixth module can output a seventh feature map, and the output seventh feature map includes the target information in the first image and the second image, and the target information includes light and shadow information and / or occlusion information.

[0201] Optionally, the third calculation module can also multiply the infinitesimal C by the scale factor k to obtain a first product, and add the first product to the fused feature map of the fifth sub-feature map and the sixth sub-feature map to obtain a fourth fused feature map, so as to avoid the situation where the numerator is equal to 0 when dividing the fifth fused feature map by the fourth fused feature map and prevent the interpolation model from running abnormally.

[0202] It should be noted that S409 - S412 are optional steps. When S409 - S412 are not executed, after the fifth module outputs the fifth feature map, the electronic device can directly input the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module to obtain the third image.

[0203] S409: Concatenate the seventh feature map and the concatenated image to obtain the eighth feature map.

[0204] Among them, the concatenation of the seventh feature map and the concatenated image is a concatenation in the channel direction.

[0205] S410: Input the eighth feature map into the seventh module to obtain the ninth feature map. The seventh module is used to perform feature extraction on the eighth feature map.

[0206] Among them, the seventh module includes a convolutional layer. The structure of the seventh module can be the same as that of the first module, the third module, and the fifth module, or can be different from at least one of the first module, the third module, and the fifth module.

[0207] This application does not make specific limitations on the structure of the seventh module.

[0208] S411: Fuse the seventh feature map and the ninth feature map to obtain the third fused feature map.

[0209] S412: Input the third fused feature map into the first activation function to obtain the third image. The first activation function is used to convert the third fused feature map into an image with pixel values maintained within a preset range.

[0210] Among them, the preset range is [0, 255]; it can be understood that usually, the RGB values corresponding to each pixel in an image are all between [0, 255]. Therefore, it is necessary to convert the third fused feature map so that the RGB values corresponding to each pixel are all between [0, 255].

[0211] In the video frame interpolation method of this application, after obtaining the first feature map, in addition to the first feature map, the electronic device can also input the concatenated image into the second module, and after obtaining the third feature map, in addition to the third feature map, the electronic device can also input the concatenated image into the fourth module, so that the intermediate layer of the frame interpolation model can also learn the original image information (referring to the concatenated image), avoid the frame interpolation model losing important feature information, and ensure the performance of the model.

[0212] Moreover, before inputting the first image and the second image into the first module, the electronic device stitches the first image and the second image to obtain a stitched image. After that, the first module can directly process the stitched image. And in the intermediate process, the second module can directly stitch the first difference data, the second difference data, and the stitched image, without having to stitch the first image and the second image before performing the step of stitching the first difference data, the second difference data, and the stitched image. Also, the fourth module can directly stitch the first transformed feature map, the second transformed feature map, and the stitched image, without having to stitch the first image and the second image again before performing the step of stitching the first transformed feature map, the second transformed feature map, and the stitched image. Thus, the stitching time of the first image and the second image can be saved, and the duration for the interpolation model to output the third image can be shortened.

[0213] Further, after inputting the first feature map and the stitched image into the second module to obtain a second feature map, the electronic device can also input the second feature map into the third module to perform feature extraction on the second feature map to obtain a third feature map. Then, the third feature map, the first difference data, the second difference data, the first image, the second image, and the stitched image are input into the fourth module, so as to obtain a first transformed feature map and a second transformed feature map, and the first transformed feature map, the second transformed feature map, and the stitched image are stitched to obtain a fourth feature map. The first transformed feature map is equivalent to the optimized first difference data, and the second transformed feature map is equivalent to the optimized second difference data. In this way, by correcting the first difference data and the second difference data, the accuracy of the first transformed feature map and the second transformed feature map for predicting the third image can be higher than that of the first difference data and the second difference data for predicting the third image.

[0214] In addition, after obtaining the fifth feature map, the electronic device can also input the fifth feature map, the first transformed feature map, and the second transformed feature map into the sixth module. Thus, through the processing of the third segmentation module and the third calculation module, a seventh feature map can be obtained. The seventh feature map can include the light and shadow information and / or occlusion information in the first image and the second image. In this way, it can be ensured that there will be no abrupt changes when the first image, the third image obtained through the seventh feature map, and the second image are played in sequence, the visual jump during playback can be avoided, and the continuity during playback can be guaranteed.

[0215] In addition, after obtaining the ninth feature map, the electronic device can also fuse the seventh feature map and the ninth feature map to form a residual connection, so as to obtain a third fused feature map, in order to retain as many image detail features including texture and clear edges as possible and ensure the performance of the model.

[0216] Based on Figure 7In the description of a specific embodiment, the following assumptions are made:

[0217] 1. The electronic device is a mobile phone.

[0218] 2. The first module is a pure convolutional stacking module 1, the second module is a redundant simplified post-processing module 1, the third module is a pure convolutional stacking module 2, the fourth module is a redundant simplified post-processing module 2, the fifth module is a pure convolutional stacking module 3, the sixth module is a redundant simplified post-processing module 3, the seventh module is a pure convolutional stacking module 4, and the first activation function is a truncated activation function.

[0219] 3. The first image is input 1, the second image is input 2, and the spliced image is input C.

[0220] Based on the above assumptions, as Figure 11 shown, when the mobile phone executes the video frame interpolation method of the embodiment of the present application, the following steps may be included:

[0221] Step 21: Obtain the first image and the second image from the target video. The first image is a frame in the target video that is earlier than the second image in the time series.

[0222] Step 22: Splice the first image and the second image to obtain a spliced image.

[0223] Step 23: Input the spliced image into the pure convolutional stacking module 1 to obtain a first feature map.

[0224] Step 24: Input the first feature map and the spliced image into the redundant simplified post-processing module 1 to obtain a second feature map.

[0225] Step 25: Input the second feature map into the pure convolutional stacking module 2 to obtain a third feature map.

[0226] Step 26: Input the third feature map, the first difference data, the second difference data, the first image, the second image, and the spliced image into the redundant simplified post-processing module 2 to obtain a fourth feature map.

[0227] Step 27: Input the fourth feature map into the pure convolutional stacking module 3 to obtain a fifth feature map.

[0228] Step 28: Input the fifth feature map, the first transformed feature map, and the second transformed feature map into the redundant simplified post-processing module 3 to obtain a seventh feature map (intermediate state output 1).

[0229] Step 29: Splice the seventh feature map and the spliced image to obtain an eighth feature map.

[0230] Step 30: Input the eighth feature map into the pure convolutional stacking module 4 to obtain a ninth feature map.

[0231] Step 31: Fuse the seventh feature map and the ninth feature map to obtain the third fused feature map.

[0232] Step 32: Input the third fused feature map into a truncation activation function to obtain the third image.

[0233] Exemplarily, the present application provides a video frame interpolation device, and the video frame interpolation device includes a module for executing the video frame interpolation method in the foregoing embodiments.

[0234] Exemplarily, the present application provides an electronic device, including a processor; when the processor executes computer code or instructions in a memory, the electronic device is caused to execute the video frame interpolation method in the foregoing embodiments.

[0235] Exemplarily, the present application provides an electronic device, including: a memory and a processor; the memory is coupled to the processor, and the memory is used to store program code or instructions; the processor is used to call the program code or instructions in the memory to cause the electronic device to execute the video frame interpolation method in the foregoing embodiments.

[0236] Exemplarily, the present application provides a chip system, and the chip system is applied to an electronic device including a memory, a display screen, and a sensor; the chip system includes: one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory and send signals to the processors, and the signals include computer code or instructions stored in the memory; when the electronic device (which may be the NPU in the electronic device) executes the computer code or instructions, the electronic device executes the video frame interpolation method in the foregoing embodiments.

[0237] Exemplarily, the present application provides a computer-readable storage medium, and code or instructions are stored in the computer-readable storage medium. When the code or instructions run on an electronic device, the electronic device is caused to execute the video frame interpolation method in the foregoing embodiments when executed.

[0238] Exemplarily, the present application provides a computer program product. When the computer program product runs on a computer, the electronic device is caused to implement the video frame interpolation method in the foregoing embodiments.

[0239] In the above embodiments, all or part of the functions may be implemented by software, hardware, or a combination of software and hardware. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer codes or instructions. When the computer program code or instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer code or instructions may be stored in a computer-readable storage medium. The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.

[0240] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The foregoing storage medium includes: various media that can store program codes such as a read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

Claims

1. A video frame insertion method, characterized in that: The method is applied to an electronic device, wherein the electronic device includes an interpolation model, the interpolation model includes a first module, a second module, a third module and a fourth module, the first module includes a convolution layer, the second module includes a first segmentation module and a first calculation module, the third module includes a convolution layer, and the fourth module includes a second segmentation module and a second calculation module; the method includes: Acquire a first image and a second image from a target video, wherein the first image is a frame in the target video that is earlier than the second image in time sequence; Inputting a spliced ​​image of the first image and the second image into the first module to obtain a first feature map, wherein the first module is used to extract pixel features of the first image and the second image, and the first feature map includes pixel features of the first image and pixel features of the second image; Inputting the first feature map and the stitched image into the second module to obtain a second feature map, the first segmentation module is used to segment the first feature map into a first sub-feature map and a second sub-feature map, the first calculation module is used to determine first difference data and second difference data between the first sub-feature map and the second sub-feature map, and stitching the first difference data, the second difference data and the stitched image, the first difference data is used to represent the movement of pixels from the first image to the second image in a time series, and the second difference data is used to represent the movement of pixels from the second image to the first image in a time series; Inputting the second feature map into the third module to obtain a third feature map, wherein the third module is used to extract features from the second feature map; The third feature map, the first difference data, the second difference data, the first image and the second image are input into the fourth module to obtain a third image, the second segmentation module is used to segment the third feature map into a third sub-feature map and a fourth sub-feature map, the second calculation module is used to fuse the first difference data and the third sub-feature map to obtain a first fused feature map, fuse the second difference data and the fourth sub-feature map to obtain a second fused feature map, warp the first image according to the first fused feature map to obtain a first transformed feature map, warp the second image according to the second fused feature map to obtain a second transformed feature map, and predict the third image according to the first transformed feature map and the second transformed feature map, the third image being an intermediate image to be inserted between the first image and the second image.

2. The method according to claim 1, characterized in that The interpolation model further includes a fifth module and a sixth module, the fifth module includes a convolution layer, the sixth module includes a transformation module, a third segmentation module and a third calculation module; the third feature map, the first difference data and the second difference data are input into the fourth module to obtain the third image, including: Inputting the third feature map, the first difference data, the second difference data, the first image, the second image and the spliced ​​image into the fourth module to obtain a fourth feature map, where the fourth feature map is obtained by splicing the first transformed feature map, the second transformed feature map and the spliced ​​image; Inputting the fourth feature map into the fifth module to obtain a fifth feature map, wherein the fifth module is used to extract features from the fourth feature map; The fifth feature map, the first transformed feature map and the second transformed feature map are input into the sixth module to obtain the third image, the transformation module in the sixth module is used to normalize each feature data in the fifth feature map to obtain a sixth feature map, the third segmentation module is used to segment the sixth feature map into a fifth sub-feature map and a sixth sub-feature map, the third calculation module is used to generate the third image including target information according to the fifth sub-feature map, the sixth sub-feature map, the first transformed feature map and the second transformed feature map, the target information including light and shadow information and / or occlusion information, and the target information included in the third image is the target information included in the first image and the second image.

3. The method according to claim 2, characterized in that The interpolation model further includes a seventh module and a first activation function, the seventh module includes a convolution layer, and the fifth feature map, the first transformation feature map, and the second transformation feature map are input into the sixth module to obtain the third image, including: Inputting the fifth feature map, the first transformed feature map and the second transformed feature map into the sixth module to obtain a seventh feature map; Splicing the seventh feature map and the spliced ​​image to obtain an eighth feature map; Inputting the eighth feature map into the seventh module to obtain a ninth feature map, wherein the seventh module is used to extract features from the eighth feature map; Fusing the seventh feature map and the ninth feature map to obtain a third fused feature map; The third fused feature map is input into the first activation function to obtain the third image, and the first activation function is used to convert the third fused feature map into an image in which pixel values ​​are kept within a preset range.

4. The method according to any one of claims 1 to 3, characterized in that The first module includes a first sublayer and a second sublayer, the first sublayer includes a first convolutional layer and a second activation function layer, and the second sublayer includes a second convolutional layer; the step of inputting the spliced ​​image of the first image and the second image into the first module to obtain a first feature map includes: Inputting the stitched image into the first sublayer to obtain a first intermediate feature map, the first convolution layer in the first sublayer is used to extract pixel features of the stitched image to obtain a second intermediate feature map, the second activation function layer is used to perform a nonlinear transformation on the second intermediate feature map to obtain the first intermediate feature map, and in the process of extracting pixel features of the stitched image by the first convolution layer, a convolution step length used is related to the target output duration of the third image, the convolution step length is an even number greater than or equal to 2, and a size of the first intermediate feature map is related to the convolution step length; The first intermediate feature map is input into the second sub-layer to obtain the first feature map, and the second convolution layer in the second sub-layer is used to adjust the size of the first intermediate feature map to the same size as the spliced ​​image.

5. The method according to claim 4, characterized in that The first module further includes a padding layer and a truncation layer, and the spliced ​​image is input into the first sub-layer to obtain a first intermediate feature map, including: Inputting the stitched image into the padding layer, and determining through the padding layer whether the length and width of the stitched image are odd numbers; If the length and / or width of the stitched image is an odd number, performing a first expansion on pixels of the stitched image to obtain an updated stitched image, wherein the first expansion is used to make the length and width of the stitched image both even numbers; Inputting the updated spliced ​​image into the first sublayer to obtain the first intermediate feature map; The step of inputting the first intermediate feature map into the second sub-layer to obtain the first feature map includes: Inputting the first intermediate feature map into the second sub-layer to obtain a second intermediate feature map, wherein the second convolutional layer of the second sub-layer is used to adjust the size of the first intermediate feature map to the same size as the updated spliced ​​image; The second intermediate feature map is input into the truncation layer to obtain the first feature map, and the truncation layer is used to adjust the size of the second intermediate feature map to the same size as the stitched image.

6. The method according to claim 5, characterized in that The first image and the second image are two adjacent frames of images in the target video, or the first image and the second image are two frames of images corresponding to each other at an interval of N frames in the target video, where N is a positive integer greater than or equal to 1.

7. The method according to claim 5, characterized in that The generation process of the interpolation model includes: Acquire a plurality of training image groups, each of the plurality of training image groups comprising three images that are continuous in time series, and the second image in each of the training image groups is a labeled intermediate image corresponding to each of the training image groups; Inputting the first image and the third image in each training image group into a first sample module of a neural network to obtain a sample feature map, wherein the first sample module is used to extract pixel features of the first image and the third image through a convolutional layer included in the first sample module, and the sample feature map includes pixel features of the first image and pixel features of the third image; Input the sample feature map into the second sample module of the neural network to obtain an estimated intermediate image corresponding to each training image group, wherein the second sample module includes a first sample segmentation module and a first sample calculation module, wherein the first sample segmentation module is used to segment the sample feature map into a first sub-sample feature map and a second sub-sample feature map, and the first sample calculation module is used to determine sample difference data between the first sub-sample feature map and the second sub-sample feature map, and determine an estimated intermediate image according to the sample difference data, wherein the sample difference data is used to represent the movement of pixels from the first image to the third image in a time series; The neural network is trained according to the difference between the labeled intermediate image and the estimated intermediate image to obtain the interpolation model.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors call the computer instructions so that the electronic device executes the method as described in any one of claims 1 to 7.

9. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises instructions, which, when executed on an electronic device, cause the electronic device to perform the method as claimed in any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video frame insertion method and device, electronic equipment and storage medium

    CN111654746A

  • Video frame insertion method and device, training method and device, electronic equipment and storage medium

    CN115002379A