Video generation method and device, electronic equipment, storage medium and chip

By acquiring high-resolution reference images and performing frame interpolation and image quality enhancement, the frame rate limitation problem in the generation of high-speed moving object videos is solved, improving the clarity and smoothness of the videos.

CN120751282BActive Publication Date: 2025-11-07BEIJING X RING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511232012.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-07
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

When shooting fast-moving objects, the camera frame rate limitation makes it impossible to capture complete and continuous object motion trajectories, resulting in low video generation flexibility and requiring the use of multiple cameras to obtain accurate video.

Method used

By capturing video of the running object, a high-resolution reference image is obtained. Temporal filtering and optical flow information extraction are performed to generate interpolated video frames. Image quality enhancement is then performed based on the reference image to improve the video frame rate and clarity.

Benefits of technology

It achieves improved video clarity and smoothness while ensuring video generation efficiency and content integrity, generating high-resolution, detailed target videos with less noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751282B_ABST
    Figure CN120751282B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method and device, electronic equipment, storage medium and chip, and relates to the technical field of video processing. The video generation method comprises: video shooting on a running object, obtaining a reference image in response to a shooting end instruction; performing frame interpolation processing on the video to obtain an interpolated video; and performing image quality enhancement on the interpolated video based on the reference image to obtain a target video, thereby improving the video definition and fluency of the target video while ensuring the video generation efficiency and content integrity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of video processing, in particular to a video generation method and device, electronic equipment, storage medium and chip. BACKGROUND

[0002] When shooting a high-speed moving object, due to the limitation of the camera frame rate, the complete and continuous object motion trajectory cannot be shot, resulting in content loss and the need to use multiple cameras to obtain an accurate video, and the video generation flexibility is low. SUMMARY

[0003] To overcome the problems in the related art, the present disclosure provides a video generation method and device, electronic equipment, storage medium and chip.

[0004] According to a first aspect of an embodiment of the present disclosure, a video generation method is provided, comprising:

[0005] shooting a video of a running object, and in response to a shooting end instruction, acquiring a reference image, wherein the resolution of the reference image is higher than that of a video frame in the video;

[0006] performing time domain filtering on the video, and extracting optical flow information of adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames;

[0007] generating an interpolated video frame corresponding to the adjacent video frames according to a preset magnification and the first optical flow information, and inserting the interpolated video frame between the adjacent video frames to obtain an interpolated video;

[0008] performing image quality enhancement on the interpolated video based on the reference image to obtain a target video.

[0009] According to a second aspect of an embodiment of the present disclosure, a video generation device is provided, comprising:

[0010] an acquisition module configured to shoot a video of a running object, and in response to a shooting end instruction, acquire a reference image, wherein the resolution of the reference image is higher than that of a video frame in the video;

[0011] a processing module configured to perform time domain filtering on the video, and extract optical flow information of adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames, generate an interpolated video frame corresponding to the adjacent video frames according to a preset magnification and the first optical flow information, and insert the interpolated video frame between the adjacent video frames to obtain an interpolated video;

[0012] a generation module configured to perform image quality enhancement on the interpolated video based on the reference image to obtain a target video.

[0013] According to a third aspect of embodiments of the present disclosure, an electronic device is provided, comprising:

[0014] a processor;

[0015] a memory for storing processor-executable instructions;

[0016] wherein the processor is configured to implement the video generation method according to the first aspect.

[0017] According to a fourth aspect of embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to perform the video generation method according to the first aspect.

[0018] According to a fifth aspect of embodiments of the present disclosure, a computer program product is provided, comprising a computer program, when the computer program is executed by a processor, the video generation method according to the first aspect is implemented.

[0019] According to a sixth aspect of embodiments of the present disclosure, a chip is provided, comprising a processing unit and an interface circuit, the processing unit acquires program instructions through the interface circuit, the program instructions are executed by the processing unit, and the processing unit is configured to execute the video generation method according to the first aspect.

[0020] The technical solutions provided by the embodiments of the present disclosure can include the following beneficial effects:

[0021] The embodiments of the present disclosure capture a video of a running object in motion, acquire a high-resolution reference image when a video capturing end instruction is received, perform frame interpolation processing on the video, increase the frame rate of the video, obtain the video after frame interpolation, and use the high-resolution reference image containing rich details and less noise as a reference basis to assist the video after frame interpolation in video quality enhancement, so as to obtain the target video after detail enhancement. In this way, the video clarity and smoothness of the target video are improved while the video generation efficiency and content integrity are ensured.

[0022] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0024] Figure 1 is a flowchart of a video generation method according to some embodiments of the present disclosure.

[0025] Figure 2 is a flowchart of a video interpolation processing method according to some embodiments of the present disclosure.

[0026] Figure 3 is a first U-shaped convolutional network schematic diagram according to some embodiments of the present disclosure.

[0027] Figure 4 is a flowchart of obtaining a target video according to some embodiments of the present disclosure.

[0028] Figure 5 is a flowchart of obtaining a specified video frame according to some embodiments of the present disclosure.

[0029] Figure 6 is a second U-shaped convolutional network schematic diagram according to some embodiments of the present disclosure.

[0030] Figure 7 is a flowchart of another video generation method according to some embodiments of the present disclosure.

[0031] Figure 8 is a block diagram of a video generation apparatus according to some embodiments of the present disclosure.

[0032] Figure 9 is a schematic diagram of an electronic device according to some embodiments of the present disclosure.

[0033] Figure 10 is a schematic diagram of a chip according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0034] Some embodiments of the present disclosure will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements in which: the sequence of operations described herein is merely an example and is not intended to be limiting, except where otherwise indicated herein, as modifications that are obvious in light of this disclosure can be made by those of ordinary skill in the art, and the sequence of operations can be performed in other sequences than those described herein or even in parallel; and the description of features known in the art can be omitted in the interest of brevity and conciseness.

[0035] The implementations described below in some embodiments of the present disclosure are not meant to be representative of all implementations consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0036] Figure 1is a flowchart of a video generation method according to some embodiments of the present disclosure, as Figure 1 as shown, comprising the following steps:

[0037] S101, video shooting on a running object, in response to a shooting end instruction, a reference image is obtained.

[0038] Optionally, the running object can be a moving person, animal or other object, such as a walking person, a flying bird or a moving vehicle, etc.

[0039] Optionally, the image processor (Image Signal Processor, ISP) can be called by the camera to shoot the running object based on the camera, and the video shooting is performed at the highest frame rate. Since high frame rate can cause low exposure and high noise, this embodiment reduces the sensor resolution to improve the signal-to-noise ratio through binning, thereby obtaining a high frame rate and low resolution video.

[0040] When the shooting end instruction is received, for example, the user performs the operation of stopping shooting, the last frame image at the time of stopping shooting is taken as the reference image, wherein the resolution of the reference image is higher than that of the video frame, and in this embodiment, the reference image can be an image taken by the camera calling the same ISP at the highest resolution (normal exposure, non-binning).

[0041] S102, frame interpolation processing is performed on the video to obtain an interpolated video.

[0042] Video interpolation is a computer vision algorithm used to insert additional frames into a video to improve the smoothness and viewing experience of the video; the core principle is to insert additional frames between existing video frames to increase the frame rate of the video. Common interpolation algorithms include methods based on optical flow and methods based on deep learning.

[0043] In some embodiments, the video interpolation process can be determined according to the magnification, which refers to the proportion of the number of new video frames inserted between the original video frames. For example, when the magnification is 2, 1 frame is inserted between every two frames to obtain an interpolated video; when the magnification is 4, 3 frames are inserted between every two frames to obtain an interpolated video; after determining the required magnification, video interpolation is performed based on the magnification.

[0044] S103, image quality enhancement is performed on the interpolated video based on the reference image to obtain a target video.

[0045] It can be understood that the reference image is the highest resolution image in the video, and thus the reference image is taken as a reference basis to enhance the image quality of the video after the frame interpolation.

[0046] In some embodiments, the image quality enhancement can be performed based on a pre-trained deep learning model. For example, the reference image and the video after the frame interpolation are input into the pre-trained model, the model performs noise removal and detail enhancement on the video after the frame interpolation according to the features of the reference image, and outputs the enhanced target video, thereby improving the efficiency of generating the target video.

[0047] In this embodiment, the running object in motion is videoed, when a video shooting end instruction is received, a high-resolution reference image is obtained, the video is subjected to frame interpolation processing, the frame rate of the video is increased, and the video after the frame interpolation is obtained. The high-resolution reference image containing rich details and less noise is taken as a reference basis to assist the video after the frame interpolation to perform video quality enhancement, so as to obtain a target video, improve the video clarity and smoothness of the target video, and generate a video with higher efficiency and quality.

[0048] Figure 2 is a flowchart of a video frame interpolation processing method according to some embodiments of the present disclosure. As shown in Figure 2 , the method comprises the following steps:

[0049] S201, performing time domain filtering on the video, and extracting optical flow information of adjacent video frames in the video after noise removal to obtain first optical flow information corresponding to the adjacent video frames.

[0050] The time domain filtering on the video refers to filtering processing on the video in the time dimension, which can preliminarily remove the noise information contained in the video and improve the clarity of the generated video.

[0051] Optionally, the adjacent video frames can be input into a first U-shaped convolutional network, and the first U-shaped convolutional network comprises a plurality of cascaded convolutional blocks, as shown in Figure 3 a first U-shaped convolutional network diagram.

[0052] The first U-shaped convolutional network is used to extract the optical flow information of the adjacent video frames to obtain the first optical flow information corresponding to the adjacent video frames, that is, the input of the first U-shaped convolutional network is the low-resolution video frames of the adjacent two frames, and the output is the optical flow information of the adjacent two low-resolution video frames. The optical flow refers to the instantaneous velocity field of the spatial moving object on the imaging plane, and the optical flow information refers to the motion vector of each pixel between the adjacent two frames, which is a bottom feature of motion analysis and determines the reliability of subsequent video generation.

[0053] S202, generate an inserted video frame corresponding to the adjacent video frame pairs according to the preset magnification and the first optical flow information, and insert the inserted video frame between the adjacent video frames to obtain the video after the insertion.

[0054] Optionally, the number of inserted video frames required between the adjacent video frames can be determined according to the preset magnification. For example, when the magnification is 2, the number of inserted video frames required between the adjacent video frames is 1 frame, and when the magnification is 4, the number of inserted video frames required between the adjacent video frames is 3 frames.

[0055] In some embodiments, the first optical flow information can be divided according to the preset magnification. For example, when the magnification is 2, the first optical flow information is divided according to a ratio of 0.5, and when the magnification is 4, the first optical flow information is divided according to a ratio of 0.25, that is, the optical flow is divided according to ratios of 0.25, 0.5 and 0.75, so as to determine the insertion optical flow of each inserted video frame.

[0056] Further, the adjacent video frames can be deformed based on the insertion optical flow to obtain an intermediate frame, and the intermediate frame can be subjected to noise removal and ghost removal and other processing operations to obtain the inserted video frame. The inserted video frame is sequentially inserted between the adjacent video frames to obtain the video after the insertion, accurately captures the pixel motion information, and improves the smoothness of the video after the insertion.

[0057] In this embodiment, the video is subjected to time domain filtering for preliminary noise removal to improve the definition of the video. The first optical flow information between each adjacent two video frames in the video is extracted based on the first U-shaped convolutional network to fully reflect the motion information of the pixels and improve the reliability of the video generation. The insertion optical flow of each inserted video frame is determined based on the preset magnification and the first optical flow information. Then, the corresponding intermediate frame is generated according to the insertion optical flow, the inserted video frame is obtained after the intermediate frame is preprocessed, and the inserted video frame is sequentially inserted between the adjacent video frames to obtain the video after the insertion. The optical flow information is used as the reference basis for the insertion, the pixel motion information is accurately captured and the insertion processing is performed, and the smoothness of the video after the insertion is improved.

[0058] Figure 4 is a flowchart for obtaining a target video according to some embodiments of the present disclosure. As shown in Figure 4 , the following steps are included:

[0059] S401, performing resolution enhancement on the video after the insertion to obtain an enhanced video.

[0060] In some embodiments, a first video frame sequence can be extracted from the video after the insertion by sliding. The first video frame sequence includes consecutive N video frames, and N is an integer greater than or equal to 1. For example, in this embodiment, N is 5, and consecutive 5 video frames are extracted from the video after the insertion as the first video frame sequence.

[0061] Indeed the first video frame sequence corresponds to a set of optical flow information, the set of optical flow information includes first optical flow information corresponding to adjacent video frames in the N video frames.

[0062] The first video frame sequence and the set of optical flow information are convolved and up-sampled to obtain a specified video frame in the first video frame sequence after resolution enhancement; this embodiment gradually magnifies and restores high-frequency information of the image through the super-resolution network structure, as shown in Figure 5 The first video frame sequence and the set of optical flow information are input into the super-resolution network structure, and feature extraction, interpolation and up-sampling operations are performed through convolution blocks. Through convolution processing of multiple convolution blocks, the first video frame sequence is converted from a low-resolution video frame to a high-resolution video frame, and a specified video frame of the first video frame sequence is output, which better restores the details of the video frame.

[0063] In response to the end of the sliding of the interpolated video, based on the specified video frame after resolution enhancement obtained by each sliding, an enhanced video is obtained; that is, the first video frame sequence is extracted by sliding from the first video frame of the interpolated video, all first video frame sequences in the interpolated video are determined by sliding traversal, the specified video frame corresponding to each first video frame sequence after resolution enhancement is obtained, and all specified video frames are spliced in order according to the positions of the original video frames to obtain an enhanced video. The enhanced video has higher resolution and better video presentation effect than the original interpolated video.

[0064] S402, based on the reference image, details of the enhanced video are enhanced to obtain a target video.

[0065] In some embodiments, second optical flow information of the video frames in the reference image and the enhanced video can be determined; optionally, for a video frame i in the enhanced video, a second video frame sequence is determined based on the first video frame in the enhanced video to the video frame i, i is an integer greater than or equal to 1; the first optical flow information of adjacent video frames in the second video frame sequence is accumulated to obtain the second optical flow information of the reference frame and the video frame i; for example, for the video frame 5, based on the first video frame in the enhanced video to the video frame 5 as the second video frame sequence, that is, the second video frame sequence includes the video frame 1, the video frame 2, the video frame 3, the video frame 4 and the video frame 5, the first optical flow information between the video frame 1 and the video frame 2, the first optical flow information between the video frame 2 and the video frame 3, the first optical flow information between the video frame 3 and the video frame 4, and the first optical flow information between the video frame 4 and the video frame 5 in the second video frame sequence are accumulated to obtain the second optical flow information of the reference video and the video frame 5.

[0066] According to the second optical flow information, the video frames in the enhanced video and the reference image are preprocessed to obtain preprocessed video frames; optionally, the preprocessed video frames can be video frames determined based on the reference image for noise removal and detail enhancement of the video frames in the enhanced video, so as to improve the definition of the generated target video.

[0067] Further, the preprocessed video frames can be feature extracted and fused to obtain target video frames subjected to detail enhancement; optionally, the preprocessed video frames can be input into a second U-shaped convolutional network, which includes multiple cascaded convolutional blocks, such as Figure 6 As shown in FIG. 6, the preprocessed video frames are feature extracted and fused by the second U-shaped convolutional network to obtain target video frames subjected to detail enhancement, and then the target video frames are sequentially combined to obtain the target video, so as to ensure the generation quality and definition of the target video.

[0068] In this embodiment, the first video frame sequence is obtained by sliding sampling the interpolated video, each first video frame sequence is subjected to multiple convolution, difference and up-sampling processing to obtain a specified video frame with high resolution, the specified video frames of all first video frame sequences are spliced to obtain an enhanced video with high resolution, the enhanced video is further subjected to detail enhancement based on the reference image, the preprocessed video frames are obtained by preprocessing the video frames in the reference image and the enhanced video based on the second optical flow information, the preprocessed video frames are feature extracted and fused by the second U-shaped convolutional network to obtain target video frames subjected to detail enhancement, the target video frames have a small change range, so there is no flickering phenomenon, the target video is determined based on all target video frames, and the generated target video has higher definition and quality.

[0069] Figure 7 is a flowchart of another video generation method according to some embodiments of the present disclosure. As shown in FIG. 6, the method includes the following steps: Figure 7

[0070] S701, video shooting of a running object is performed, and a reference image is obtained in response to a shooting end instruction.

[0071] In some embodiments, a first camera module can be called to perform video shooting of the running object at a highest frame rate and a first resolution; in this embodiment, the first resolution is a lower resolution, so the shot video is a high-frame-rate, low-resolution video.

[0072] ​In response to the shooting end instruction, the first camera module is adjusted from the first resolution to the highest resolution, and a tail video frame is shot to obtain a video, wherein the tail video frame is the reference image, that is, the tail video frame of the video is shot at the highest resolution, and other video frames are shot at the first resolution, the same first camera module is called for shooting the scene, so that the problem of missing shooting content is avoided, and the video imaging consistency is ensured.

[0073] In some other embodiments, the first camera module can also be called to shoot a video of the running object based on the highest frame rate and the first resolution of the first camera module; in response to receiving a video shooting end instruction, the shooting is ended to obtain a candidate video; the second camera module is called, and the running object is shot based on the highest resolution to obtain a candidate tail video frame as the reference image; that is, different camera modules are used for video shooting, the tail video frame is shot by the second camera module based on the highest resolution, and the candidate video is shot by the first camera module based on the first resolution.

[0074] Based on the reference image, the candidate video is corrected to obtain a video; optionally, the reference image and the tail video frame in the candidate video can be pixel-level aligned, that is, the reference image is pixel-accurately matched with the tail video frame in the candidate video, so that the same scene or object is completely overlapped in the spatial position, thereby obtaining the video, and the smoothness and integrity of the video are ensured.

[0075] Optionally, the reference image can also be used to replace the tail video frame in the candidate video, that is, the last video frame in the candidate video is replaced by the reference image, thereby obtaining the video.

[0076] S702, time domain filtering is performed on the video, and optical flow information extraction is performed on adjacent video frames in the de-noised video to obtain first optical flow information corresponding to the adjacent video frames.

[0077] In the embodiments of the application, the implementation method of step S702 can be implemented by any one of the embodiments of the present disclosure, and here it is not limited, and will not be repeated.

[0078] S703, according to the preset magnification and the first optical flow information, a frame video frame corresponding to adjacent video frames is generated and inserted between the adjacent video frames to obtain a frame-inserted video.

[0079] In the embodiments of the application, the implementation method of step S703 can be implemented by any one of the embodiments of the present disclosure, and here it is not limited, and will not be repeated.

[0080] S704, the resolution of the frame-inserted video is enhanced to obtain an enhanced video.

[0081] In the embodiments of the present application, the implementation method of step S704 can be implemented by any one of the embodiments of the present disclosure, which will not be limited here and will not be described again.

[0082] S705, performing detail enhancement on the enhanced video based on the reference image to obtain a target video.

[0083] In the embodiments of the present application, the implementation method of step S705 can be implemented by any one of the embodiments of the present disclosure, which will not be limited here and will not be described again.

[0084] In the embodiments, the moving running object is videoed, the same camera module or different camera modules are used for videoing to obtain the video and the reference image, the resolution of the reference image is higher than that of other video frames, so as to improve the definition of the generated video; the video is temporally filtered to remove preliminary noise and improve the definition of the video; the first optical flow information between each two adjacent video frames in the video is extracted based on the first U-shaped convolutional network; based on the preset magnification and the first optical flow information, the inserted video frame is determined and inserted between the adjacent video frames in sequence to obtain the video after the insertion, the optical flow information is used as the reference basis for the insertion, the pixel motion information is accurately captured and the insertion processing is performed, and the fluency of the video after the insertion is improved; the resolution of the video after the insertion is processed to obtain the enhanced video with high resolution, and the enhanced video is further enhanced based on the reference image to obtain the target video frame, and the target video is determined according to all the target video frames, and the definition and quality of the generated target video are higher.

[0085] Figure 8 is a block diagram of a video generation apparatus according to some embodiments of the present disclosure. Referring to Figure 8 The video generation apparatus 800 includes:

[0086] The acquisition module 801 is configured to video the running object, and acquire the reference image in response to a videoing end instruction, wherein the resolution of the reference image is higher than that of the video frames in the video.

[0087] The processing module 802 is configured to perform insertion processing on the video to obtain the video after the insertion.

[0088] The generation module 803 is configured to perform image quality enhancement on the video after the insertion based on the reference image to obtain the target video.

[0089] In some implementations, the processing module 802 includes:

[0090] The video is temporally filtered, and the optical flow information of the adjacent video frames in the denoised video is extracted to obtain the first optical flow information corresponding to the adjacent video frames.

[0091] According to the preset magnification and the first optical flow information, an interpolated video frame corresponding to the adjacent video frames is generated and inserted between the adjacent video frames, to obtain an interpolated video.

[0092] In some implementations, the processing module 802 includes:

[0093] The adjacent video frames are input into a first U-shaped convolutional network, the first U-shaped convolutional network including a plurality of cascaded convolutional blocks, and the first optical flow information corresponding to the adjacent video frames is obtained by performing optical flow information extraction on the adjacent video frames through the first U-shaped convolutional network.

[0094] In some implementations, the generation module 803 includes:

[0095] The interpolated video is subjected to resolution enhancement to obtain an enhanced video.

[0096] The enhanced video is subjected to detail enhancement based on a reference image to obtain a target video.

[0097] In some implementations, the generation module 803 includes:

[0098] A first video frame sequence is extracted from the interpolated video by sliding, the first video frame sequence including N consecutive video frames, N being an integer greater than or equal to 1;

[0099] A set of optical flow information corresponding to the first video frame sequence is determined, the set of optical flow information including the first optical flow information corresponding to the adjacent video frames in the N video frames;

[0100] The first video frame sequence and the set of optical flow information are subjected to convolution and up-sampling processing to obtain a specified video frame in the first video frame sequence that has been subjected to resolution enhancement.

[0101] In response to the end of the sliding of the interpolated video, the specified video frame that has been subjected to resolution enhancement obtained each time the sliding is performed is used to obtain an enhanced video.

[0102] In some implementations, the generation module 803 includes:

[0103] Second optical flow information of the video frames in the reference image and the enhanced video is determined.

[0104] The video frames in the enhanced video and the reference image are preprocessed according to the second optical flow information to obtain preprocessed video frames.

[0105] The preprocessed video frames are subjected to feature extraction and fusion to obtain target video frames that have been subjected to detail enhancement.

[0106] The target video is obtained based on the target video frames.

[0107] In some implementations, the generation module 803 includes:

[0108] The pre-processed video frame is input into a second U-shaped convolutional network, the second U-shaped convolutional network comprising a plurality of cascaded convolutional blocks, and the pre-processed video frame is subjected to feature extraction and fusion through the second U-shaped convolutional network to obtain a target video frame subjected to detail enhancement.

[0109] In some implementations, the generating module 803 comprises:

[0110] For a video frame i in the enhanced video, a second video frame sequence is determined based on the first video frame in the enhanced video to the video frame i, i being an integer greater than or equal to 1;

[0111] The first optical flow information of adjacent video frames in the second video frame sequence is accumulated to obtain the second optical flow information of the reference frame and the video frame i.

[0112] In some implementations, the obtaining module 801 comprises:

[0113] The first camera module is called to use the highest frame rate and the first resolution to perform video shooting on the running object;

[0114] In response to a shooting end instruction, the first camera module is adjusted from the first resolution to the highest resolution, and a tail video frame is shot to obtain a video, wherein the tail video frame is the reference image.

[0115] In some implementations, the obtaining module 801 comprises:

[0116] The first camera module is called to perform video shooting on the running object based on the highest frame rate and the first resolution of the first camera module;

[0117] In response to receiving a video shooting end instruction, the shooting is ended to obtain a candidate video;

[0118] The second camera module is called and performs shooting on the running object based on the highest resolution to obtain a candidate tail video frame as the reference image;

[0119] The candidate video is corrected based on the reference image to obtain the video.

[0120] In some implementations, the obtaining module 801 comprises:

[0121] The reference image and the tail video frame in the candidate video are pixel-level aligned to obtain the video; or,

[0122] The reference image is used to replace the tail video frame in the candidate video to obtain the video.

[0123] As to the video generation apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the video generation method, and will not be described in detail here.

[0124] In this embodiment, the moving running object is videoed, and the video and the reference image are obtained by the same camera module or different camera modules, the resolution of the reference image is higher than that of other video frames, so as to improve the definition of the generated video; the video is time domain filtered to remove preliminary noise and improve the definition of the video; the first optical flow information between each two adjacent video frames in the video is extracted based on the first U-shaped convolution network, the interpolated video frame is determined based on the preset magnification and the first optical flow information, the interpolated video frame is sequentially inserted between the adjacent video frames, and the interpolated video is obtained; the optical flow information is used as the reference basis for interpolation, the pixel motion information is accurately captured and interpolation processing is performed, and the smoothness of the interpolated video is improved; the resolution of the interpolated video is processed to obtain an enhanced video with high resolution, the enhanced video is further detail enhanced based on the reference image, the target video frame is obtained, and the target video is determined according to all the target video frames, and the definition and quality of the generated target video are higher.

[0125] To achieve the above embodiments, the present disclosure further provides an electronic device, as shown in Figure 9 The electronic device 900 includes a processor 910, one or more memories 920 for storing instructions executable by the processor 910, wherein the processor 910 is configured to perform the video generation method described in the above embodiments, and the processor 910 and the memory 920 are connected through a communication bus.

[0126] To achieve the above embodiments, the present disclosure further provides a computer readable storage medium including instructions, such as a memory 920 including instructions, which can be executed by a processor 910 to complete the above method; optionally, the computer readable storage medium can be a ROM, a random access memory RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0127] To achieve the above embodiments, the present disclosure further provides a computer program product, including a computer program, characterized in that the computer program is executed by a processor to implement the video generation method described in the above embodiments.

[0128] To achieve the above embodiments, the present disclosure further provides a chip, as shown in Figure 10As shown, the chip includes at least one processor 1010 and at least one interface circuit 1020. The processor 1010 and the interface circuit 1020 can be interconnected via a line. For example, the interface circuit 1020 can be configured to receive a signal from another device (e.g., the processor 1010), and for example, the interface circuit 1020 can be configured to send a signal to another device (e.g., the processor 1010); illustratively, the interface circuit 1020 can read an instruction stored in a memory and send the instruction to the processor 1010.

[0129] When the instruction is executed by the processor 1010, each step of the video generation method described in the above embodiments can be performed by the port selection device. Of course, the chip can also include other discrete devices, and some embodiments of the present disclosure do not specifically limit this.

[0130] In some embodiments of the present disclosure, the interface circuit 1020 can obtain data, program instructions and / or information, etc. in an internal storage area of the chip; or can obtain data, program instructions and / or information, etc. outside the chip system.

[0131] Optionally, the chip further includes a memory for storing necessary computer programs and data.

[0132] Those skilled in the art can also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether the function is implemented by hardware or software depends on the specific application and design requirements of the overall system. Those skilled in the art can implement the described functions in various ways for each specific application, but such implementation should not be construed as beyond the scope of the embodiments of the present application.

[0133] Other embodiments of the present disclosure will be apparent to those skilled in the art with the consideration of the specification and practice of the disclosed application. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including modifications and equivalents that are obvious to those skilled in the art. The specification and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0134] It should be understood that the present disclosure is not limited to the precise structures described and shown in the above and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the claims appended hereto.

Claims

1. A method of video generation, the method comprising: The method comprises: Video shooting is performed on a running object, and a reference image is acquired in response to a shooting end instruction, wherein the resolution of the reference image is higher than that of a video frame in a video; Temporal filtering is performed on the video, and optical flow information extraction is performed on adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames; According to a preset magnification and the first optical flow information, an interpolated video frame corresponding to the adjacent video frames is generated and inserted between the adjacent video frames to obtain an interpolated video; Image quality enhancement is performed on the interpolated video based on the reference image to obtain a target video.

2. The method of claim 1, wherein, The optical flow information extraction on the adjacent video frames in the denoised video to obtain the first optical flow information corresponding to the adjacent video frames comprises: The adjacent video frames are input into a first U-shaped convolutional network, the first U-shaped convolutional network comprises a plurality of convolutional blocks connected in cascade, and the first optical flow information corresponding to the adjacent video frames is obtained by performing optical flow information extraction on the adjacent video frames through the first U-shaped convolutional network.

3. The method of claim 1, wherein, The image quality enhancement on the interpolated video based on the reference image to obtain the target video comprises: Resolution enhancement is performed on the interpolated video to obtain an enhanced video; Detail enhancement is performed on the enhanced video based on the reference image to obtain the target video.

4. The method of claim 3, wherein, The resolution enhancement on the interpolated video to obtain the enhanced video comprises: A first video frame sequence is extracted from the interpolated video by sliding, the first video frame sequence comprises N continuous video frames, and N is an integer greater than or equal to 1; A set of optical flow information corresponding to the first video frame sequence is determined, and the set of optical flow information comprises first optical flow information corresponding to adjacent video frames in the N video frames; Convolution and up-sampling processing are performed on the first video frame sequence and the set of optical flow information to obtain a specified video frame in the first video frame sequence that has been subjected to resolution enhancement; In response to the end of the sliding of the interpolated video, the enhanced video is obtained based on the specified video frame that has been subjected to resolution enhancement obtained each time the sliding is performed.

5. The method of claim 3, wherein, The detail enhancement on the enhanced video based on the reference image to obtain the target video comprises: Second optical flow information of video frames in the reference image and the enhanced video is determined; According to the second optical flow information, pre-processing is performed on the video frames in the enhanced video and the reference image to obtain pre-processed video frames; Feature extraction and fusion are performed on the pre-processed video frames to obtain target video frames that have been subjected to detail enhancement; The target video is obtained based on the target video frames.

6. The method of claim 5, wherein, The feature extraction and fusion on the pre-processed video frames to obtain the target video frames that have been subjected to detail enhancement comprises: The pre-processed video frames are input into a second U-shaped convolutional network, the second U-shaped convolutional network comprises a plurality of convolutional blocks connected in cascade, and the target video frames that have been subjected to detail enhancement are obtained by performing feature extraction and fusion on the pre-processed video frames through the second U-shaped convolutional network.

7. The method of claim 6, wherein, The determination of the second optical flow information of the video frames in the reference image and the enhanced video comprises: For the video frame i in the enhanced video, a second video frame sequence is determined based on the first video frame in the enhanced video to the video frame i, where i is an integer greater than or equal to 1; The first optical flow information of adjacent video frames in the second video frame sequence is accumulated to obtain the second optical flow information of the reference image and the video frame i.

8. The method according to any one of claims 1-7, characterized in that, The video of the running object is captured, and a reference image is obtained in response to a capture end instruction, including: A first camera module is called to capture a video of the running object at a highest frame rate and a first resolution; In response to a capture end instruction, the first camera module is adjusted from the first resolution to a highest resolution, and a final video frame is captured to obtain the video, where the final video frame is the reference image.

9. The method according to any one of claims 1-6, characterized in that, The video of the running object is captured, and a reference image is obtained in response to a capture end instruction, including: A first camera module is called to capture a video of the running object at a highest frame rate and a first resolution; In response to receiving a video capture end instruction, the capture is ended to obtain a candidate video; A second camera module is called, and a final video frame of the running object is captured at a highest resolution as the reference image; The candidate video is corrected based on the reference image to obtain the video.

10. The method of claim 9, wherein, The candidate video is corrected based on the reference image to obtain the video, including: The reference image and the final video frame in the candidate video are pixel-level aligned to obtain the video; or The final video frame in the candidate video is replaced with the reference image to obtain the video.

11. A video generating apparatus characterized by comprising: Including: An acquisition module is configured to capture a video of a running object, and obtain a reference image in response to a capture end instruction, where the resolution of the reference image is higher than that of a video frame in the video; A processing module is configured to perform time domain filtering on the video, extract optical flow information of adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames, generate an interpolated video frame corresponding to the adjacent video frames according to a preset magnification and the first optical flow information, and insert the interpolated video frame between the adjacent video frames to obtain an interpolated video; A generation module is configured to perform image quality enhancement on the interpolated video based on the reference image to obtain a target video.

12. An electronic device, comprising: Including: A processor; A memory for storing processor-executable instructions; The processor is configured to implement the video generation method of any one of claims 1-10.

13. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to perform the video generation method of any one of claims 1-10.

14. A computer program product, characterised in that, Including a computer program, when the computer program is executed by a processor, the video generation method of any one of claims 1-10 is implemented.

15. A chip, characterized by The chip comprises a processing unit and an interface circuit, the processing unit acquires program instructions through the interface circuit, the program instructions are executed by the processing unit, and the processing unit is used for executing the video generation method in any one of claims 1-10.

Citation Information

Patent Citations

  • Super-resolution processing method based on optical flow interpolation

    CN112488922A

  • Video processing method and device and storage medium

    CN114025202A