Video generation method and device, electronic equipment, storage medium and chip
By acquiring high-resolution reference images and performing frame interpolation and image quality enhancement, the problem of frame rate limitation in video shooting of high-speed moving objects is solved, and the clarity and smoothness of the video are improved.
Patent Information
- Application Number
- CN202511232012.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
When shooting high-speed moving objects, the camera frame rate limitation makes it impossible to capture the complete and continuous object movement trajectory, resulting in content loss. In addition, multiple cameras are required to obtain accurate video, and the video generation flexibility is low.
By shooting a video of the running object, a high-resolution reference image is obtained, time domain filtering and optical flow information extraction are performed, interpolated video frames are generated, and image quality enhancement is performed based on the reference image to improve the video frame rate and clarity.
It achieves the goal of improving the clarity and smoothness of the video while ensuring the efficiency of video generation and the integrity of the content, and generating a target video with high resolution, rich details and less noise.
Smart Images

Figure CN120751282A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video processing technology, and in particular to a video generation method, device, electronic device, storage medium, and chip. Background Art
[0002] When shooting high-speed moving objects, the camera frame rate limitation makes it impossible to capture the complete and continuous object movement trajectory, resulting in content loss and the need to use multiple cameras to obtain accurate video, resulting in low video generation flexibility. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides a video generation method, device, electronic device, storage medium and chip.
[0004] According to a first aspect of an embodiment of the present disclosure, a video generation method is provided, including: Performing video capture of the running object, and acquiring a reference image in response to a capture end instruction, wherein the reference image has a higher resolution than a video frame in the video; Performing time-domain filtering on the video, and extracting optical flow information from adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames; Generate interpolated video frames corresponding to the adjacent video frames according to a preset magnification and the first optical flow information, and insert the interpolated video frames between the adjacent video frames to obtain an interpolated video; Image quality enhancement is performed on the inserted video based on the reference image to obtain a target video.
[0005] According to a second aspect of an embodiment of the present disclosure, there is provided a video generating apparatus, including: an acquisition module, configured to capture a video of the running object and, in response to a capture end instruction, acquire a reference image, wherein the resolution of the reference image is higher than the video frame in the video; a processing module, configured to perform time-domain filtering on the video, extract optical flow information from adjacent video frames in the denoised video, obtain first optical flow information corresponding to the adjacent video frames, generate interpolated video frames corresponding to the adjacent video frames based on a preset magnification and the first optical flow information, and insert the interpolated video frames between the adjacent video frames to obtain an interpolated video; A generation module is used to enhance the image quality of the inserted video based on the reference image to obtain a target video.
[0006] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including: processor; a memory for storing processor-executable instructions; The processor is configured to implement the video generation method described in the embodiment of the first aspect.
[0007] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided. When instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to execute the video generation method described in the embodiment of the first aspect.
[0008] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the video generation method described in the embodiment of the first aspect.
[0009] According to the sixth aspect of the embodiment of the present disclosure, a chip is provided, including a processing unit and an interface circuit, wherein the processing unit obtains program instructions through the interface circuit, and the program instructions are executed by the processing unit, and the processing unit is used to execute the video generation method described in the embodiment of the first aspect.
[0010] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects: The disclosed embodiment captures a video of a moving object, obtains a high-resolution reference image upon receiving a video capture end instruction, performs frame interpolation on the video, increases the frame rate of the video, and obtains an interpolated video. The high-resolution reference image containing rich details and less noise is used as a reference basis to assist in enhancing the video quality of the interpolated video, thereby obtaining a target video with enhanced details, thereby improving the video clarity and smoothness of the target video while ensuring video generation efficiency and content integrity.
[0011] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0013] Figure 1 is a flowchart of a video generation method according to some embodiments of the present disclosure.
[0014] Figure 2 This is a flowchart of a video frame insertion processing method according to some embodiments of the present disclosure.
[0015] Figure 3 This is a schematic diagram of a first U-shaped convolutional network according to some embodiments of the present disclosure.
[0016] Figure 4The present invention is a flowchart illustrating a method of acquiring a target video according to some embodiments of the present disclosure.
[0017] Figure 5 The present invention is a flowchart illustrating a method of obtaining a specified video frame according to some embodiments of the present disclosure.
[0018] Figure 6 is a schematic diagram of a second U-shaped convolutional network according to some embodiments of the present disclosure.
[0019] Figure 7 is a flowchart illustrating another video generation method according to some embodiments of the present disclosure.
[0020] Figure 8 It is a block diagram of a video generating device according to some embodiments of the present disclosure.
[0021] Figure 9 is a schematic diagram of an electronic device according to some embodiments of the present disclosure.
[0022] Figure 10 is a schematic diagram of a chip according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0023] Some embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, for the sake of clarity and brevity, descriptions of features known in the art may be omitted.
[0024] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0025] Figure 1 is a flow chart of a video generation method according to some embodiments of the present disclosure. Figure 1 As shown, the following steps are included: S101 , capturing a video of a running object, and acquiring a reference image in response to a capture end instruction.
[0026] Optionally, the running object may be a moving person, animal, or other object, such as a walking person, a flying bird, or a moving vehicle.
[0027] Optionally, the camera can call an image signal processor (ISP) to shoot a video of the running object based on the camera. The video shooting is performed at the highest frame rate. Since a high frame rate may result in low exposure and high background noise, this embodiment reduces the sensor resolution through binning to improve the signal-to-noise ratio, thereby obtaining a high frame rate, low-resolution video.
[0028] When a shooting end instruction is received, for example, when the user stops shooting, the last frame image when shooting stops is used as a reference image, where the resolution of the reference image is higher than the video frame in the video. In this embodiment, the reference image can be an image captured by the camera using the same ISP and the highest resolution (normal exposure, non-binning).
[0029] S102: Perform frame insertion processing on the video to obtain a video after frame insertion.
[0030] Video interpolation is a computer vision algorithm used to insert additional frames into a video to improve video smoothness and viewing experience. The core principle is to increase the frame rate of the video by inserting additional frames between existing video frames. Common interpolation algorithms include optical flow-based methods and deep learning-based methods.
[0031] In some embodiments, the video interpolation process can be determined based on the magnification. The magnification refers to the ratio of the number of new video frames inserted between the original video frames. For example, when the magnification is 2, 1 frame is inserted between every two frames to obtain the interpolated video; when the magnification is 4, 3 frames are inserted between every two frames to obtain the interpolated video; after determining the required magnification, video interpolation is performed based on the magnification.
[0032] S103: Perform image quality enhancement on the inserted video based on the reference image to obtain a target video.
[0033] It can be understood that the reference image is the image with the highest resolution in the video, so the reference image is used as a reference basis to enhance the image quality of the interpolated video; optionally, the image quality enhancement includes but is not limited to noise removal and detail enhancement, thereby improving the video clarity and smoothness of the generated target video.
[0034] In some embodiments, image quality enhancement can be performed based on a pre-trained deep learning model. For example, a reference image and an interpolated video are input into a pre-trained model. The model removes noise and enhances details of the interpolated video according to the features of the reference image, and outputs an enhanced target video, thereby improving the efficiency of target video generation.
[0035] In this embodiment, a video of a moving object is captured. When a video capture end instruction is received, a high-resolution reference image is obtained, and the video is interpolated to increase the frame rate of the video to obtain an interpolated video. The high-resolution reference image containing rich details and less noise is used as a reference basis to assist in enhancing the video quality of the interpolated video, thereby obtaining a target video, improving the video clarity and smoothness of the target video, and generating a video with higher efficiency and quality.
[0036] Figure 2 FIG. 1 is a flow chart of a video frame insertion processing method according to some embodiments of the present disclosure. Figure 2 As shown, the following steps are included: S201 , performing time domain filtering on the video, and extracting optical flow information from adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames.
[0037] Temporal filtering of a video refers to filtering the video in the time dimension, which can preliminarily remove the noise information contained in the video and improve the clarity of the generated video.
[0038] Alternatively, adjacent video frames may be input into a first U-shaped convolutional network comprising multiple layers of cascaded convolutional blocks, such as Figure 3 Schematic diagram of the first U-shaped convolutional network shown.
[0039] The optical flow information of adjacent video frames is extracted through the first U-shaped convolutional network to obtain the first optical flow information corresponding to the adjacent video frames. That is, the input of the first U-shaped convolutional network is two adjacent low-resolution video frames, and the output is the optical flow information of two adjacent low-resolution video frames. Optical flow refers to the instantaneous velocity field of pixel motion of a spatial moving object on the imaging plane. Optical flow information is a motion vector that describes each pixel between two adjacent frames. As the underlying feature of motion analysis, it determines the reliability of subsequent video generation.
[0040] S202 : Generate interpolated video frames corresponding to adjacent video frames according to a preset magnification and first optical flow information, and insert the interpolated video frames between the adjacent video frames to obtain an interpolated video.
[0041] Optionally, the number of video frames required to be interpolated between adjacent video frames can be determined based on a preset magnification. For example, when the magnification is 2, the number of video frames required to be interpolated between adjacent video frames is 1 frame; when the magnification is 4, the number of video frames required to be interpolated between adjacent video frames is 3 frames.
[0042] In some embodiments, the first optical flow information can be divided according to a preset magnification. For example, when the magnification is 2, the first optical flow information is divided according to a ratio of 0.5. When the magnification is 4, the first optical flow information is divided according to a ratio of 0.25. That is, the optical flow is divided according to the ratios of 0.25, 0.5, and 0.75, thereby determining the interpolated optical flow of each interpolated video frame.
[0043] Furthermore, adjacent video frames can be deformed based on the interpolated optical flow to obtain intermediate frames, and noise removal and ghosting removal and other processing operations can be performed on the intermediate frames to obtain interpolated video frames. The interpolated video frames can be inserted between adjacent video frames in sequence to obtain the interpolated video, accurately capture pixel motion information, and improve the smoothness of the interpolated video.
[0044] In this embodiment, time-domain filtering is performed on the video to perform preliminary noise removal to improve the clarity of the video. The first optical flow information between each two adjacent video frames in the video is extracted based on the first U-shaped convolutional network to fully reflect the motion information of the pixels and improve the reliability of video generation. The interpolated optical flow of each interpolated video frame is determined based on the preset magnification and the first optical flow information, and then the corresponding intermediate frames are generated according to the interpolated optical flow. The intermediate frames are pre-processed to obtain interpolated video frames, and the interpolated video frames are sequentially inserted between adjacent video frames to obtain the interpolated video. The optical flow information is used as a reference basis for interpolation, and the pixel motion information is accurately captured and interpolation processing is performed to improve the smoothness of the interpolated video.
[0045] Figure 4 FIG. 1 is a flow chart showing a method of obtaining a target video according to some embodiments of the present disclosure. Figure 4 As shown, the following steps are included: S401: Perform resolution enhancement on the inserted frame video to obtain an enhanced video.
[0046] In some embodiments, a first video frame sequence can be slidingly extracted from the video after interpolation, and the first video frame sequence includes N consecutive video frames, where N is an integer greater than or equal to 1; for example, in this embodiment, N is 5, then 5 consecutive video frames are slidingly extracted from the video after interpolation as the first video frame sequence.
[0047] An optical flow information set corresponding to the first video frame sequence is determined, where the optical flow information set includes first optical flow information corresponding to adjacent video frames in the N video frames.
[0048] The first video frame sequence and the optical flow information set are convolved and up-sampled to obtain a specified video frame in the first video frame sequence with enhanced resolution. This embodiment uses a super-resolution network structure to gradually amplify and restore the high-frequency information of the image. Specifically, Figure 5 As shown, the first video frame sequence and the optical flow information set are input into the super-resolution network structure, and feature extraction, interpolation and upsampling operations are performed through the convolution block. Through the convolution processing of multiple convolution blocks, the first video frame sequence is converted from low-resolution video frames to high-resolution video frames, and the specified video frames of the first video frame sequence are output to better restore the details of the video frames.
[0049] In response to the end of sliding of the inserted video, an enhanced video is obtained based on the designated video frames with enhanced resolution obtained by each sliding; that is, the first video frame sequence is extracted by sliding starting from the first video frame of the inserted video, all the first video frame sequences in the inserted video are determined by sliding traversal, the designated video frames corresponding to each first video frame sequence after resolution enhancement are obtained, and all the designated video frames are spliced in sequence according to the positions of the original video frames to obtain an enhanced video. Compared with the original inserted video, the enhanced video has a higher resolution and a better video presentation effect.
[0050] S402: Perform detail enhancement on the enhanced video based on the reference image to obtain a target video.
[0051] In some embodiments, second optical flow information of video frames in a reference image and an enhanced video can be determined; optionally, for video frame i in the enhanced video, a second video frame sequence is determined based on the first video frame in the enhanced video to video frame i, where i is an integer greater than or equal to 1; the first optical flow information of adjacent video frames in the second video frame sequence is accumulated to obtain the second optical flow information of the reference frame and video frame i; for example, for video frame 5, based on the first video frame in the enhanced video to video frame 5, as the second video frame sequence, that is, the second video frame sequence includes video frame 1, video frame 2, video frame 3, video frame 4 and video frame 5, the first optical flow information between video frame 1 and video frame 2, the first optical flow information between video frame 2 and video frame 3, the first optical flow information between video frame 3 and video frame 4, and the first optical flow information between video frame 4 and video frame 5 in the second video frame sequence are accumulated to obtain the second optical flow information of the reference video and video frame 5.
[0052] According to the second optical flow information, the video frames in the enhanced video and the reference image are preprocessed to obtain preprocessed video frames; optionally, the preprocessed video frames can be video frames determined by removing noise and enhancing details of the video frames in the enhanced video based on the reference image to improve the clarity of the generated target video.
[0053] Furthermore, the pre-processed video frame can be subjected to feature extraction and fusion to obtain a target video frame with enhanced details; optionally, the pre-processed video frame can be input into a second U-shaped convolutional network, which includes a multi-layer cascade of convolutional blocks, such as Figure 6 As shown, the preprocessed video frames are feature extracted and fused through the second U-shaped convolutional network to obtain the target video frames with enhanced details, and then the target video frames are combined in sequence to obtain the target video, ensuring the generated image quality and clarity of the target video.
[0054] In this embodiment, sliding sampling is performed on the interpolated video to obtain a first video frame sequence, and multiple convolution, difference and upsampling processes are performed on each first video frame sequence to obtain a high-resolution designated video frame. The designated video frames of all the first video frame sequences are spliced to obtain a high-resolution enhanced video, and the enhanced video is further enhanced in detail based on the reference image. The reference image and the second optical flow information of the video frames in the enhanced video are preprocessed to obtain a preprocessed video frame with noise removed and detail enhanced. The preprocessed video frame is subjected to feature extraction and fusion through a second U-shaped convolutional network to obtain a target video frame with enhanced detail. The target video frames have a small variation range, so there will be no flickering. The target video is determined based on all the target video frames, and the generated target video has higher clarity and image quality.
[0055] Figure 7 FIG. 1 is a flow chart of another video generation method according to some embodiments of the present disclosure. Figure 7 As shown, the following steps are included: S701 , capturing a video of a running object, and acquiring a reference image in response to a capture end instruction.
[0056] In some embodiments, the first camera module can be called to use the highest frame rate and the first resolution to shoot video of the running object; in this embodiment, the first resolution is a lower resolution, so the captured video is a high frame rate, low resolution video.
[0057] In response to the shooting end instruction, the first camera module is adjusted from the first resolution to the highest resolution, and the last video frame is shot to obtain a video, wherein the last video frame is a reference image, that is, the last video frame of the video is shot at the highest resolution, and the other video frames are shot at the lower first resolution. The shooting scene calls the same first camera module to avoid the problem of loss of shooting content and ensure the consistency of video imaging.
[0058] In other embodiments, the first camera module can also be called to shoot video of the running object based on the highest frame rate and first resolution of the first camera module; in response to receiving the video shooting end instruction, the shooting is ended to obtain a candidate video; the second camera module is called, and the running object is shot based on the highest resolution to obtain a candidate end video frame as a reference image; that is, different camera modules are used to shoot videos, the end video frame is shot by the second camera module based on the highest resolution, and the candidate video is shot by the first camera module based on the first resolution.
[0059] Based on the reference image, the candidate video is corrected to obtain a video; optionally, the reference image and the last video frame in the candidate video can be aligned at the pixel level, that is, the reference image and the last video frame in the candidate video are accurately matched pixel by pixel to ensure that the same scene or object completely overlaps in spatial position, thereby obtaining a video and ensuring the smoothness and integrity of the video.
[0060] Optionally, the reference image may be used to replace the last video frame in the candidate video, that is, the last video frame in the candidate video may be replaced with the reference image, thereby obtaining the video.
[0061] S702 , performing time domain filtering on the video, and extracting optical flow information from adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames.
[0062] In the embodiment of the present application, the implementation method of step S702 can be implemented by any of the methods in the embodiments of the present disclosure, which is not limited here and will not be repeated.
[0063] S703 : Generate interpolated video frames corresponding to adjacent video frames according to a preset magnification and the first optical flow information, and insert the interpolated video frames between the adjacent video frames to obtain an interpolated video.
[0064] In the embodiment of the present application, the implementation method of step S703 can be implemented by any of the methods in the embodiments of the present disclosure, which is not limited here and will not be repeated.
[0065] S704: Perform resolution enhancement on the inserted frame video to obtain an enhanced video.
[0066] In the embodiment of the present application, the implementation method of step S704 can be implemented by any of the methods in the embodiments of the present disclosure, which is not limited here and will not be repeated.
[0067] S705 , performing detail enhancement on the enhanced video based on the reference image to obtain a target video.
[0068] In the embodiment of the present application, the implementation method of step S705 can be implemented by any of the methods in the embodiments of the present disclosure, which is not limited here and will not be repeated.
[0069] In this embodiment, a video of a moving object is captured using the same camera module or different camera modules to obtain a captured video and a reference image. The resolution of the reference image is higher than that of other video frames to improve the clarity of the generated video. The video is subjected to time-domain filtering for preliminary noise removal to improve the clarity of the video. The first optical flow information between each two adjacent video frames in the video is extracted based on a first U-shaped convolutional network. Based on a preset magnification and the first optical flow information, an interpolated video frame is determined. The interpolated video frames are sequentially inserted between adjacent video frames to obtain an interpolated video. The optical flow information is used as a reference basis for interpolation to accurately capture pixel motion information and perform interpolation processing to improve the smoothness of the interpolated video. Resolution processing is performed on the interpolated video to obtain a high-resolution enhanced video. The enhanced video is further enhanced in detail based on the reference image to obtain a target video frame. The target video is determined based on all the target video frames. The generated target video has higher clarity and image quality.
[0070] Figure 8 FIG is a block diagram of a video generation device according to some embodiments of the present disclosure. Figure 8 , the video generating device 800 includes: An acquisition module 801 is configured to capture a video of the running object and, in response to a capture end instruction, acquire a reference image, wherein the resolution of the reference image is higher than that of the video frame in the video; The processing module 802 is used to perform frame insertion processing on the video to obtain a video after frame insertion; The generating module 803 is configured to perform image quality enhancement on the inserted video based on the reference image to obtain a target video.
[0071] In some implementations, the processing module 802 includes: Performing time-domain filtering on the video and extracting optical flow information from adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames; According to the preset magnification and the first optical flow information, interpolated video frames corresponding to adjacent video frames are generated and inserted between the adjacent video frames to obtain an interpolated video.
[0072] In some implementations, the processing module 802 includes: Adjacent video frames are input into a first U-shaped convolutional network, which includes multiple layers of cascaded convolutional blocks. Optical flow information of adjacent video frames is extracted through the first U-shaped convolutional network to obtain first optical flow information corresponding to the adjacent video frames.
[0073] In some implementations, the generating module 803 includes: Performing resolution enhancement on the interpolated video to obtain an enhanced video; The details of the enhanced video are enhanced based on the reference image to obtain the target video.
[0074] In some implementations, the generating module 803 includes: Slidingly extracting a first video frame sequence from the interpolated video, the first video frame sequence including N consecutive video frames, where N is an integer greater than or equal to 1; Determine an optical flow information set corresponding to the first video frame sequence, where the optical flow information set includes first optical flow information corresponding to adjacent video frames in the N video frames; Performing convolution and upsampling processing on the first video frame sequence and the optical flow information set to obtain a specified video frame in the first video frame sequence with enhanced resolution; In response to the end of the video sliding after the interpolation frame, an enhanced video is obtained based on the designated video frame with enhanced resolution obtained in each sliding.
[0075] In some implementations, the generating module 803 includes: determining second optical flow information of the reference image and the video frames in the enhanced video; Preprocessing the video frame and the reference image in the enhanced video according to the second optical flow information to obtain a preprocessed video frame; Perform feature extraction and fusion on the preprocessed video frames to obtain the target video frames with enhanced details; A target video is obtained based on the target video frame.
[0076] In some implementations, the generating module 803 includes: The preprocessed video frame is input into the second U-shaped convolutional network, which includes multiple layers of cascaded convolutional blocks. The second U-shaped convolutional network is used to extract and fuse features of the preprocessed video frame to obtain a target video frame with enhanced details.
[0077] In some implementations, the generating module 803 includes: For video frame i in the enhanced video, determine a second video frame sequence starting from the first video frame to video frame i in the enhanced video, where i is an integer greater than or equal to 1; The first optical flow information of adjacent video frames in the second video frame sequence is accumulated to obtain the second optical flow information of the reference frame and the video frame i.
[0078] In some implementations, the acquisition module 801 includes: Calling the first camera module to use the highest frame rate and the first resolution to shoot a video of the running object; In response to the shooting end instruction, the first camera module is adjusted from the first resolution to the highest resolution, and the last video frame is shot to obtain a video, wherein the last video frame is a reference image.
[0079] In some implementations, the acquisition module 801 includes: Invoking the first camera module to capture a video of the running object based on a maximum frame rate and a first resolution of the first camera module; In response to receiving a video shooting end instruction, ending shooting and obtaining a candidate video; Calling the second camera module and shooting the running object based on the highest resolution to obtain a candidate end video frame as a reference image; Based on the reference image, the candidate video is corrected to obtain a video.
[0080] In some implementations, the acquisition module 801 includes: Perform pixel-level alignment on the reference image and the last video frame in the candidate video to obtain the video; or, The reference image is used to replace the last video frame in the candidate video to obtain the video.
[0081] Regarding the video generating device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the video generating method, and will not be elaborated here.
[0082] In this embodiment, a video of a moving object is captured using the same camera module or different camera modules to obtain a captured video and a reference image. The resolution of the reference image is higher than that of other video frames to improve the clarity of the generated video. The video is subjected to time-domain filtering for preliminary noise removal to improve the clarity of the video. The first optical flow information between each two adjacent video frames in the video is extracted based on a first U-shaped convolutional network. Based on a preset magnification and the first optical flow information, an interpolated video frame is determined. The interpolated video frames are sequentially inserted between adjacent video frames to obtain an interpolated video. The optical flow information is used as a reference basis for interpolation to accurately capture pixel motion information and perform interpolation processing to improve the smoothness of the interpolated video. Resolution processing is performed on the interpolated video to obtain a high-resolution enhanced video. The enhanced video is further enhanced in detail based on the reference image to obtain a target video frame. The target video is determined based on all the target video frames. The generated target video has higher clarity and image quality.
[0083] In order to implement the above embodiments, the present disclosure also provides an electronic device, such as Figure 9As shown, the electronic device 900 includes: a processor 910; one or more memories 920 for storing executable instructions of the processor 910; wherein the processor 910 is configured to execute the video generation method described in the above embodiment, wherein the processor 910 and the memory 920 are connected via a communication bus.
[0084] In order to implement the above embodiments, the present disclosure also provides a computer-readable storage medium including instructions, such as a memory 920 including instructions, and the above instructions can be executed by the processor 910 to complete the above method; optionally, the computer-readable storage medium can be a ROM, a random access memory RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0085] In order to implement the above embodiments, the present disclosure further provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the video generation method described in the above embodiments is implemented.
[0086] In order to implement the above embodiment, the present disclosure also provides a chip, such as Figure 10 As shown, the chip includes at least one processor 1010 and at least one interface circuit 1020. The processor 1010 and the interface circuit 1020 can be interconnected via lines; for example, the interface circuit 1020 can be used to receive signals from other devices (such as the processor 1010), or for another example, the interface circuit 1020 can be used to send signals to other devices (such as the processor 1010); illustratively, the interface circuit 1020 can read instructions stored in the memory and send the instructions to the processor 1010.
[0087] When the instructions are executed by the processor 1010, the port selection device can execute the steps of the video generation method described in the above embodiment. Of course, the chip can also include other discrete devices, which are not specifically limited in some embodiments of the present disclosure.
[0088] In some embodiments of the present disclosure, the interface circuit 1020 can obtain data, program instructions and / or information from the internal storage area of the chip; it can also obtain data, program instructions and / or information from outside the chip system.
[0089] Optionally, the chip also includes a memory for storing necessary computer programs and data.
[0090] Those skilled in the art will also appreciate that the various illustrative logic blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of the two. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present application.
[0091] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0092] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A video generation method, characterized in that: The method comprises: Performing video capture of the running object, and acquiring a reference image in response to a capture end instruction, wherein the reference image has a higher resolution than a video frame in the video; Performing time-domain filtering on the video, and extracting optical flow information from adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames; Generate interpolated video frames corresponding to the adjacent video frames according to a preset magnification and the first optical flow information, and insert the interpolated video frames between the adjacent video frames to obtain an interpolated video; Image quality enhancement is performed on the inserted video based on the reference image to obtain a target video.
2. The method according to claim 1, characterized in that The extracting optical flow information from adjacent video frames in the denoised video to obtain first optical flow information corresponding to the adjacent video frames includes: The adjacent video frames are input into a first U-shaped convolutional network, which includes multiple layers of cascaded convolutional blocks. The optical flow information of the adjacent video frames is extracted through the first U-shaped convolutional network to obtain first optical flow information corresponding to the adjacent video frames.
3. The method according to claim 1, characterized in that The performing image quality enhancement on the interpolated video based on the reference image to obtain a target video includes: Performing resolution enhancement on the interpolated video to obtain an enhanced video; The enhanced video is enhanced in detail based on the reference image to obtain the target video.
4. The method according to claim 3, characterized in that The step of performing resolution enhancement on the interpolated video to obtain an enhanced video includes: Slidingly extracting a first video frame sequence from the interpolated video, the first video frame sequence comprising N consecutive video frames, where N is an integer greater than or equal to 1; Determine an optical flow information set corresponding to the first video frame sequence, wherein the optical flow information set includes first optical flow information corresponding to adjacent video frames in the N video frames; Performing convolution and upsampling processing on the first video frame sequence and the optical flow information set to obtain a specified video frame in the first video frame sequence with enhanced resolution; In response to the completion of the sliding of the inserted video, the enhanced video is obtained based on the designated video frames with enhanced resolution obtained during each sliding.
5. The method according to claim 3, characterized in that The performing detail enhancement on the enhanced video based on the reference image to obtain the target video includes: Determining second optical flow information of the reference image and the video frame in the enhanced video; Preprocessing the video frame in the enhanced video and the reference image according to the second optical flow information to obtain a preprocessed video frame; Performing feature extraction and fusion on the preprocessed video frames to obtain target video frames with enhanced details; The target video is obtained based on the target video frame.
6. The method according to claim 5, characterized in that The extracting and fusing features of the pre-processed video frame to obtain a target video frame with enhanced details includes: The preprocessed video frame is input into a second U-shaped convolutional network, which includes a multi-layer cascaded convolutional block. The preprocessed video frame is subjected to feature extraction and fusion through the second U-shaped convolutional network to obtain a target video frame with enhanced details.
7. The method according to claim 6, characterized in that The determining of second optical flow information of the reference image and the video frame in the enhanced video includes: For video frame i in the enhanced video, determine a second video frame sequence starting from the first video frame in the enhanced video to the video frame i, where i is an integer greater than or equal to 1; The first optical flow information of adjacent video frames in the second video frame sequence is accumulated to obtain the second optical flow information of the reference frame and the video frame i.
8. The method according to any one of claims 1 to 7, characterized in that The step of capturing a video of the running object and obtaining a reference image in response to a capture end instruction includes: Calling the first camera module to use the highest frame rate and the first resolution to shoot a video of the running object; In response to a shooting end instruction, the first camera module is adjusted from the first resolution to the highest resolution, and the last video frame is shot to obtain the video, wherein the last video frame is the reference image.
9. The method according to any one of claims 1 to 6, characterized in that The step of capturing a video of the running object and obtaining a reference image in response to a capture end instruction includes: Invoking a first camera module to capture a video of the running object based on a maximum frame rate and a first resolution of the first camera module; In response to receiving a video shooting end instruction, ending shooting and obtaining a candidate video; Invoking a second camera module and shooting the running object based on the highest resolution to obtain a candidate end video frame as the reference image; Based on the reference image, the candidate video is modified to obtain the video.
10. The method according to claim 9, characterized in that The step of correcting the candidate video based on the reference image to obtain the video includes: Perform pixel-level alignment on the reference image and the last video frame in the candidate video to obtain the video; or, The reference image is used to replace the last video frame in the candidate video to obtain the video.
11. A video generating device, characterized in that: include: an acquisition module, configured to capture a video of the running object and, in response to a capture end instruction, acquire a reference image, wherein the resolution of the reference image is higher than the video frame in the video; a processing module, configured to perform time-domain filtering on the video, extract optical flow information from adjacent video frames in the denoised video, obtain first optical flow information corresponding to the adjacent video frames, generate interpolated video frames corresponding to the adjacent video frames based on a preset magnification and the first optical flow information, and insert the interpolated video frames between the adjacent video frames to obtain an interpolated video; A generation module is used to enhance the image quality of the inserted video based on the reference image to obtain a target video.
12. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the video generation method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute the video generation method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The invention comprises a computer program, which implements the video generation method according to any one of claims 1 to 10 when the computer program is executed by a processor.
15. A chip, characterized in that: The chip includes a processing unit and an interface circuit. The processing unit obtains program instructions through the interface circuit. The program instructions are executed by the processing unit. The processing unit is used to execute the video generation method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Video frame interpolation method and system based on optical flow method
CN105517671A
Super-resolution processing method based on optical flow interpolation
CN112488922A
Video processing method and device and storage medium
CN114025202A
Space-time video stream enhancement method and device, terminal and medium
CN116957937A
Video frame insertion method and frame insertion model training method
CN118101874A