A video processing method, apparatus, device, and storage medium

By extracting keyframes, determining the action primitive set, and performing local details and interpolation processing, the problem of image quality improvement in video compression is solved, and the image quality improvement and fluency after video compression is achieved.

CN120075473BActive Publication Date: 2025-07-22MT TITLIS BEIJING CONTROL TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510550304.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-22
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to improve the image quality of compressed video while compressing videos, especially in dynamic scenes, where the video quality is prone to distortion or incoherence.

Method used

By extracting keyframes from the original video, forming a key video sequence, determining the action primitive set, local detail refinement and interframe dynamic interpolation processing are performed to generate optimized videos.

Benefits of technology

While achieving video compression, it improves picture quality, ensuring the richness and authenticity of the video in local actions or scene transformation, avoiding visual jumps or distortion, and improving the smoothness of video playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075473B_ABST
    Figure CN120075473B_ABST
Patent Text Reader

Abstract

The present invention discloses a video processing method, apparatus, device, and storage medium, belonging to the technical field of video processing. The method includes: extracting key frames from an original video to obtain a key video sequence composed of the key frames; determining an action primitive set according to the key video sequence and the total diffusion time step; refining local details of the key video sequence according to the action primitive set to obtain an optimized video; and performing inter-frame dynamic interpolation processing on the optimized video to obtain a target video. The present invention improves the image quality of the compressed original video while realizing the compression of the original video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a video processing method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of information technology, video content has become an indispensable part of people's daily life and work. From social media, online education, remote conferencing to professional film production and broadcasting, the demand for high-quality video content is increasing day by day. Consequently, there is a continuous growing demand for video compression technology and image quality improvement methods. Especially in dynamic scenes, how to improve the image quality of compressed videos while achieving video compression has become an urgent problem to be solved. Summary of the Invention

[0003] The present invention provides a video processing method, apparatus, device, and storage medium to improve the image quality of compressed videos while achieving video compression.

[0004] According to one aspect of the present invention, a video processing method is provided, which includes:

[0005] Extracting key frames from the original video to obtain a key video sequence composed of key frames;

[0006] Determining an action primitive set according to the key video sequence and the total diffusion time step length;

[0007] Performing local detail refinement on the key video sequence according to the action primitive set to obtain an optimized video;

[0008] Performing inter-frame dynamic interpolation processing on the optimized video to obtain a target video.

[0009] According to another aspect of the present invention, a video processing apparatus is provided, which includes:

[0010] A key video sequence determination module for extracting key frames from the original video to obtain a key video sequence composed of key frames;

[0011] An action primitive set determination module for determining an action primitive set according to the key video sequence and the total diffusion time step length;

[0012] An optimized video determination module for performing local detail refinement on the key video sequence according to the action primitive set to obtain an optimized video;

[0013] A target video determination module for performing inter-frame dynamic interpolation processing on the optimized video to obtain a target video.

[0014] According to another aspect of the present invention, an electronic device is provided, and the electronic device includes:

[0015] At least one processor;

[0016] And a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the video processing method of any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the video processing method of any embodiment of the present invention when executed.

[0019] According to another aspect of the present invention, there is provided a computer program product including a computer program which implements the video processing method of any embodiment of the present invention when executed by a processor.

[0020] The technical solution of the embodiment of the present invention extracts key frames from the original video to obtain a key video sequence composed of the key frames; determines an action primitive set according to the key video sequence and the total diffusion time step length; refines local details of the key video sequence according to the action primitive set to obtain an optimized video; and performs inter-frame dynamic interpolation processing on the optimized video to obtain a target video. The above technical solution extracts key frames from the original video to obtain a key video sequence, realizes video compression of the original video, reduces the size of the original video, thereby reducing the computational cost of processing the original video and improving the processing efficiency of the original video; then, determines an action primitive set according to the key video sequence and the total diffusion time step length, providing necessary global video dynamic information for obtaining the subsequent target video; then, refines local details of the key video sequence according to the action primitive set to obtain an optimized video, ensuring the richness, accuracy and authenticity of video details in local actions or local scene transformations of the optimized video; then, obtains a target video by performing inter-frame dynamic interpolation processing on the optimized video, smoothing the difference between two adjacent video frames in the optimized video, making the visual content of the obtained target video more coherent and smoother, and avoiding visual jumps or distortions during the playback of the target video, thereby improving the video quality and picture quality of the target video and the smoothness of the target video playback. In short, the above technical solution improves the picture quality of the compressed original video while realizing the compression of the original video.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of a video processing method provided in Embodiment 1 of the present invention;

[0024] Figure 2 It is a flowchart of a video processing method provided in Embodiment 2 of the present invention;

[0025] Figure 3 It is a schematic structural diagram of a video processing device provided in Embodiment 3 of the present invention;

[0026] Figure 4 It is a schematic structural diagram of an electronic device for implementing the video processing method of the embodiments of the present invention. Detailed Embodiments

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be noted that the terms "original", "target", "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] In addition, it should be noted that in the technical solution of the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of the total diffusion time step length, detailed refinement table, etc. comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0030] Embodiment 1

[0031] Figure 1 As shown in the flowchart of a video processing method provided in Embodiment 1 of the present invention, this embodiment is applicable to the situation of processing videos with rich details in dynamic scenes (such as sports event videos). This method can be executed by a video processing device, and the device can be implemented in the form of hardware and / or software and can be configured in an electronic device. As Figure 1 shown, the method includes:

[0032] S101. Extract key frames from the original video to obtain a key video sequence composed of key frames.

[0033] Among them, the original video refers to the video to be processed. A key frame refers to a video frame in the original video that can reflect the change of visual information. For example, a key frame can be a video frame in the original video that contains a plot turning point. It should be noted that the number of key frames is multiple. A key video sequence refers to a video sequence composed of multiple key frames.

[0034] Specifically, a preset inter-frame difference method can be used to calculate the difference value between each video frame in the original video and its adjacent previous video frame; if the difference value is greater than the difference threshold, then this video frame is used as a key frame, so as to extract multiple key frames from the original video; according to the time arrangement order of the extracted key frames in the original video, the key frames are arranged in an orderly manner to obtain a key video sequence composed of multiple key frames.

[0035] Among them, the preset inter-frame difference method can be a pixel difference method, a histogram difference method, or an image feature difference method. Among them, the pixel difference method refers to directly comparing the pixel gray value or brightness value difference of adjacent video frames in the video, and usually uses indicators such as mean-square error (MSE) or Structural Similarity Index (SSIM); the histogram difference method refers to comparing the color histogram difference of adjacent video frames in the video; the image feature difference method refers to comparing the feature difference of adjacent video frames in the video. For example, comparing the edge difference of adjacent video frames in the video. The difference threshold can be preset according to the experience of those skilled in the art, and the embodiments of the present invention do not make specific limitations on it.

[0036] It can be understood that extracting key frames from the original video to obtain a key video sequence composed of key frames realizes the compression of the original video, reduces the size of the original video, decreases the storage space required for storing the original video, lowers the computational cost for processing the original video, improves the processing efficiency of the original video, and at the same time maintains the integrity and observability of the video content in the original video.

[0037] Among them, extracting key frames from the original video can also be: inputting the original video into a key frame extraction model, and after being processed by the key frame extraction model, obtaining multiple key frames. Among them, the key frame extraction model is obtained by training a deep learning model.

[0038] It can be understood that when the key frame extraction model extracts key frames from the original video, it not only considers the changes between adjacent video frames in the original video, but also considers the content complexity and information richness of each video frame in the original video frames, so that the key frames extracted based on the key frame extraction model are more accurate and reliable.

[0039] Optionally, in order to further improve the accuracy and efficiency of key frame extraction, preprocessing such as denoising, color correction, and image enhancement can also be performed on the original video before extracting key frames from the original video to obtain the preprocessed original video.

[0040] S102. Determine an action primitive set according to the key video sequence and the total diffusion time steps.

[0041] Among them, the diffusion time step refers to the time step used to control the number of diffusion iterations. The total diffusion time steps refer to the total duration of the diffusion time steps; optionally, the total diffusion time steps can be determined according to the video length and frame rate of the original video. For example, if the video length of the original video is 10 seconds and the frame rate is 30 frames per second, the total diffusion time steps can be 300, that is, 10 steps per frame; it can also be preset according to the content complexity of the original video. An action primitive refers to the smallest unit of an action that can represent the key content in the original video. The action primitive set refers to a set composed of multiple action primitives, which is used to characterize the global dynamic information of the original video.

[0042] Specifically, feature extraction can be performed on the key video sequence to obtain a key feature set; determine the initial action primitive; determine the action primitive set according to the key feature set, the initial action primitive, and the total diffusion time steps. Among them, the key feature refers to the feature extracted from the key frames in the key video sequence. It should be noted that one key frame corresponds to one key feature. The key feature set refers to a set composed of multiple key features. The initial action primitive refers to the action primitive when the diffusion time step is zero.

[0043] More specifically, a deep convolutional neural network can be adopted to extract features from each key frame in the key video sequence through the following feature extraction formula, so as to obtain the key features corresponding to each key frame in the key video sequence, and thus obtain a key feature set composed of multiple key features.

[0044] ;

[0045] Among them, represents the key feature corresponding to the k-th key frame in the key video sequence; N represents the total number of convolutional layers in the deep convolutional neural network; represents the weight corresponding to the i-th convolutional layer in the deep convolutional neural network, which is used to characterize the importance of this convolutional layer in the feature extraction process; represents the k-th key frame in the key video sequence; represents the convolution operation of the i-th convolutional layer in the deep convolutional neural network on the k-th key frame in the key video sequence.

[0046] After that, the initial action primitive can be determined according to the first key feature in the key feature set through the following initial primitive determination formula:

[0047] ;

[0048] Among them, represents the initial action primitive; represents the first key feature in the key feature set; G() represents a function for generating global features, which can be preset according to actual business requirements, and the embodiments of the present invention do not make specific limitations on it. It should be noted that the first key feature refers to the key feature corresponding to the first key frame in the key video sequence.

[0049] After that, according to the key feature set, the initial action primitive and the total diffusion time steps, the action primitive corresponding to each diffusion time step is determined through the following action primitive determination formula, so as to obtain an action primitive set composed of the action primitives corresponding to each diffusion time step.

[0050] ;

[0051] Among them, t represents the t-th diffusion time step, and ; T represents the total diffusion time steps; represents the action primitive corresponding to the t-th diffusion time step; represents the action primitive corresponding to the (t + 1)-th diffusion time step; represents the diffusion rate; D() represents the diffusion function; F represents the key feature set; G() represents a function for generating global features. It should be noted that the initial action primitive (i.e., ) It is not determined by the above action primitive determination formula.

[0052] Optionally, in order to determine the initial action primitive simply and quickly, the initial action primitive can also be randomly determined according to the experience of those skilled in the art, or the initial action primitive can be set to a fixed value.

[0053] It can be understood that based on the initial action primitive, the key features in the key feature set are transformed into a set of action primitives by means of a diffusion function, so that the obtained action primitive set can effectively summarize the global dynamics and key content of the original video, providing data support for obtaining the optimized video subsequently, thereby making the details of each video frame in the optimized video more accurate.

[0054] S103. According to the action primitive set, perform local detail refinement on the key video sequence to obtain an optimized video.

[0055] Among them, the optimized video refers to the video obtained after local detail refinement of the key video sequence. Specifically, based on a preset detail fusion algorithm, the key video sequence can be locally detailed and refined according to the action primitive set to obtain an optimized video.

[0056] S104. Perform inter-frame dynamic interpolation processing on the optimized video to obtain a target video.

[0057] Among them, the target video refers to the video obtained after inter-frame dynamic interpolation processing of the optimized video. Specifically, for each pair of video frames in the optimized video, according to the previous video frame, the subsequent video frame, and the total number of time steps between the previous video frame and the subsequent video frame in the pair of video frames, an intermediate frame between the previous video frame and the subsequent video frame is generated; where the pair of video frames consists of two adjacent video frames in the optimized video. Thus, the difference between two adjacent video frames in the optimized video is smoothed, making the visual content of the obtained target video more coherent and smoother, avoiding the occurrence of visual jumps or distortions when playing the target video, and further improving the image quality of the target video and the smoothness of playing the target video.

[0058] More specifically, for each pair of video frames in the optimized video, according to the previous video frame, the subsequent video frame, and the total number of time steps between the previous video frame and the subsequent video frame in the pair of video frames, through the following intermediate frame generation formula, an intermediate frame between the previous video frame and the subsequent video frame is generated, thereby filling the blank between two adjacent video frames in the optimized video to obtain a target video.

[0059] ;

[0060] Among them, represents the previous video frame in the pair of video frames; denote the latter video frame in the video frame pair; T denotes the total number of time steps between the former video frame and the latter video frame in the video frame pair; the value of t is ; denote the time weighting function, which is used to dynamically adjust the weight of interpolation according to the value of t, so as to control the generation of intermediate frames.

[0061] In the technical solution of the embodiment of the present invention, key frames are extracted from the original video to obtain a key video sequence composed of key frames; an action primitive set is determined according to the key video sequence and the total diffusion time steps; according to the action primitive set, local details of the key video sequence are refined to obtain an optimized video; and inter-frame dynamic interpolation processing is performed on the optimized video to obtain the target video. The above technical solution extracts key frames from the original video to obtain a key video sequence, realizing video compression of the original video, reducing the size of the original video, thereby reducing the computational cost of processing the original video and improving the processing efficiency of the original video; then, according to the key video sequence and the total diffusion time steps, an action primitive set is determined, providing necessary global video dynamic information for obtaining the target video subsequently; then, according to the action primitive set, local details of the key video sequence are refined to obtain an optimized video, ensuring the richness, accuracy, and authenticity of video details in local actions or local scene transformations of the optimized video; then, by performing inter-frame dynamic interpolation processing on the optimized video to obtain the target video, the difference between two adjacent video frames before and after in the optimized video is smoothed, making the visual content of the obtained target video more coherent and smoother, avoiding visual jumps or distortions during the playback of the target video, thereby improving the video quality and picture quality of the target video and enhancing the smoothness of the target video playback. All in all, the above technical solution improves the picture quality of the compressed original video while realizing the compression of the original video.

[0062] Embodiment 2

[0063] Figure 2 is a flowchart of a video processing method provided by Embodiment 2 of the present invention. On the basis of the above embodiment, this embodiment further optimizes "according to the action primitive set, local details of the key video sequence are refined to obtain an optimized video", and provides an optional implementation solution. It should be noted that for parts not detailed in the embodiments of the present invention, reference may be made to the relevant descriptions of other embodiments. As Figure 2 shown, the method includes:

[0064] S201. Extract key frames from the original video to obtain a key video sequence composed of key frames.

[0065] S202. Determine an action primitive set according to the key video sequence and the total diffusion time steps.

[0066] S203. Determine a base frame from the key video sequence according to the content importance degree of the key frames in the key video sequence.

[0067] Among them, for each key frame in the key video sequence, the content importance degree refers to the importance degree of the visual content of this key frame in the entire key video sequence. Optionally, the content importance degree of the key frame can be determined according to the visual prominence degree, motion intensity and entropy value of the pixel gray distribution in the key frame, and combined with the actual business requirements. Among them, the visual prominence degree of the key frame can be calculated by a pre-trained visual prominence degree calculation model; the motion intensity of the key frame can be obtained by calculating the displacement amplitude of each pixel in the key frame by the optical flow method. The base frame refers to the video frame used as a reference during the video reconstruction process.

[0068] Specifically, the key frame with the greatest content importance degree can be selected from all the key frames included in the key video sequence as the base frame. By doing so, the quality of the reconstructed frames generated based on the base frame subsequently can be improved, making the visual content of the reconstructed frames clearer and more reliable.

[0069] Optionally, in order to determine the base frame conveniently and quickly, the first key frame in the key video sequence can also be used as the base frame; or, a key frame can be randomly selected from the key video sequence as the base frame.

[0070] S204. Perform frame derivation processing on the base frame to obtain at least one variant frame corresponding to the base frame.

[0071] Among them, the variant frame refers to the video frame changed from the base frame. Specifically, affine transformations such as translation, rotation and scaling can be performed on the base frame to obtain at least one variant frame corresponding to the base frame.

[0072] Optionally, in order to generate variant frames simply and quickly, the color, contrast, brightness, etc. of the base frame can also be manually adjusted to obtain at least one variant frame corresponding to the base frame.

[0073] Optionally, in order to improve the quality and diversity of the variant frames, frame derivation processing can also be performed on the base frame based on a frame derivation model to obtain at least one variant frame corresponding to the base frame. Among them, the frame derivation model can be obtained by training a deep learning model (such as a generative adversarial network model) with a large number of sample video frames.

[0074] Optionally, in order to obtain more diverse variant frames, affine transformations such as translation, rotation, and scaling can also be performed on the base frame to obtain the first set of variant frames corresponding to the base frame; by manually adjusting the color, contrast, brightness, etc. of the base frame, the second set of variant frames corresponding to the base frame is obtained; based on the frame derivation model, frame derivation processing is performed on the base frame to obtain the third set of variant frames corresponding to the base frame, thereby obtaining three sets of variant frames corresponding to the base frame.

[0075] S205. Generate a reconstructed frame according to the action primitive set, the base frame, and at least one variant frame corresponding to the base frame.

[0076] Among them, the reconstructed frame refers to a video frame used for locally refining the details of the key video sequence. Specifically, the credibility of the variant frame can be determined according to the structural similarity index, mean square error, and feature difference value between the variant frame and the base frame; according to the credibility of each variant frame, the weight coefficient corresponding to each variant frame is determined; according to the action primitive set, the base frame, at least one variant frame corresponding to the base frame, and the weight coefficient corresponding to each variant frame, a reconstructed frame is generated.

[0077] Among them, the structural similarity index is an index for evaluating the quality of the variant frame by quantifying the similarity of the variant frame and the base frame in terms of brightness, contrast, and structure, that is:

[0078] ;

[0079] Among them, represents the structural similarity index between the variant frame and the base frame; represents the average brightness of the variant frame; represents the average brightness of the base frame; represents the standard deviation of the variant frame; represents the standard deviation of the base frame; represents the covariance between the variant frame and the base frame; and are constants. It should be noted that the theoretical value range of the structural similarity index is , but in practical applications, its value range is generally . The closer the value of the structural similarity index is to 1, the more similar the variant frame is to the base frame.

[0080] Among them, the mean square error is an index reflecting the pixel-level difference degree between the variant frame and the base frame, and is obtained by calculating the average value of the squared differences between the corresponding pixel points of the variant frame and the base frame, that is:

[0081] ;

[0082] Among them, represents the image size corresponding to the variant frame; Represents the image size corresponding to the base frame; Represents the variant frame at the position The pixel value at; Represents the base frame at the position The pixel value at.

[0083] Among them, the feature difference value is an index reflecting the distance between the variant frame and the base frame at the feature level, and usually uses cosine similarity or Euclidean distance.

[0084] More specifically, for each variant frame corresponding to the base frame, according to the mean square error between the variant frame and the base frame, through the following score conversion formula, determine the error score corresponding to the variant frame:

[0085] ;

[0086] Among them, Represents the error score corresponding to the variant frame; MSE represents the mean square error between the variant frame and the base frame. Then, perform a weighted sum processing or an average processing on the structural similarity index and the feature difference value between the variant frame and the base frame, and the error score corresponding to the variant frame, to obtain the credibility of the variant frame, so that the credibility of each variant frame corresponding to the base frame can be obtained. Then, according to the credibility of each variant frame, through the following weight coefficient determination formula, determine the weight coefficient corresponding to each variant frame corresponding to the base frame:

[0087] ;

[0088] Among them, Represents the weight coefficient corresponding to the j-th variant frame corresponding to the base frame; n represents the total number of variant frames corresponding to the base frame; Represents the credibility of the j-th variant frame corresponding to the base frame; Represents the sum of the credibilities of each variant frame corresponding to the base frame.

[0089] Then, according to the action primitive set, the base frame, at least one variant frame corresponding to the base frame, and the weight coefficient corresponding to each variant frame, through the following reconstruction frame generation formula, generate a reconstruction frame:

[0090] ;

[0091] Among them, Represents the reconstruction frame; Represents the base frame; Represents the j-th variant frame corresponding to the base frame; Represents the weight coefficient corresponding to the j-th variant frame corresponding to the base frame; n represents the total number of variant frames corresponding to the base frame; B represents the action primitive set; R() represents the reconstruction function.

[0092] It is understandable that by determining the credibility of each variant frame corresponding to the base frame, different weight coefficients are assigned to different variant frames, and then a reconstructed frame is generated based on the action primitive set, the base frame, at least one variant frame corresponding to the base frame, and the weight coefficients corresponding to each variant frame, thereby improving the quality of the reconstructed frame.

[0093] Optionally, in order to reduce the number of variant frames corresponding to the base frame and also to improve the quality of the variant frames, before generating the reconstructed frame based on the action primitive set, the base frame, and at least one variant frame corresponding to the base frame, it is also possible to: screen at least one variant frame corresponding to the base frame according to the visual quality evaluation result, dynamic consistency evaluation result, temporal consistency evaluation result, and content compatibility evaluation result of each variant frame corresponding to the base frame, so as to obtain the screened variant frames corresponding to the base frame.

[0094] Among them, the visual quality evaluation result refers to the result obtained after evaluating the visual quality of the variant frame. The dynamic consistency evaluation result is used to reflect the degree of consistency between the movement of the object in the variant frame and the movement of the object in the base frame. The temporal consistency evaluation result is used to reflect the smoothness of the action change between the variant frame and the front and back video frames after inserting the variant frame into the key video sequence. The content compatibility evaluation result is used to reflect the degree of compatibility between the variant frame and the video content in the key video sequence.

[0095] More specifically, for each variant frame corresponding to the base frame, a weighted sum processing is performed on the visual quality evaluation result, dynamic consistency evaluation result, temporal consistency evaluation result, and content compatibility evaluation result of the variant frame to obtain the influence degree of the variant frame on the video content reconstruction, so that the influence degrees of each variant frame corresponding to the base frame on the video content reconstruction can be obtained; the influence degrees of each variant frame on the video content reconstruction are respectively compared with the influence degree threshold, and the variant frames with the influence degree greater than or equal to the influence degree threshold are screened out as the screened variant frames corresponding to the base frame. Among them, the influence degree threshold can be preset according to the experience of those skilled in the art, and the embodiments of the present invention do not make specific limitations on it.

[0096] S206. Refine the local details of the key video sequence according to the reconstructed frame to obtain an optimized video.

[0097] Specifically, the frame pairs to be processed in the key video sequence can be determined according to the frame difference value and the difference threshold between adjacent video frame pairs in the key video sequence; according to the frame difference value between the frame pairs to be processed, based on the corresponding relationship between the frame difference value and the number of inserted frames in the detail refinement table, the number of frame insertions of the reconstructed frame is determined; according to the number of frame insertions, the reconstructed frame is inserted between the frame pairs to be processed to obtain an optimized video.

[0098] Among them, the frame difference value refers to the difference value between adjacent video frame pairs in the key video sequence; optionally, the frame difference value can be calculated by a preset inter-frame difference method. Among them, the preset inter-frame difference method can be a pixel difference method, a histogram difference method, or an image feature difference method. The frame pair to be processed refers to the video frame pair into which a reconstructed frame needs to be inserted in the middle. The detail refinement table refers to a data table used to formulate the number of reconstructed frame insertions in the video reconstruction process. The number of frame insertions refers to the total number of reconstructed frames that need to be inserted between the frame pairs to be processed.

[0099] More specifically, adjacent video frame pairs in the key video sequence with a frame difference value greater than the difference threshold can be used as the frame pairs to be processed in the key video sequence; the frame difference value between the frame pairs to be processed is matched with the frame difference value in the detail refinement table to obtain a matching frame difference value; among them, the matching frame difference value refers to the frame difference value in the detail refinement table that matches the frame difference value between the frame pairs to be processed; based on the correspondence between the frame difference value and the number of inserted frames in the detail refinement table, the number of inserted frames corresponding to the matching frame difference value is extracted from the detail refinement table as the number of frame insertions for the reconstructed frame; then, in the time steps between the two video frames in the frame pair to be processed, the number of frame insertions of reconstructed frames is inserted at a preset time interval to obtain an optimized video, realizing the detail reconstruction of the visually significant change area in the key video sequence, making the local video details of the obtained optimized video richer, more realistic and credible.

[0100] S207. Perform inter-frame dynamic interpolation processing on the optimized video to obtain the target video.

[0101] The technical solution of the embodiment of the present invention extracts key frames from the original video to obtain a key video sequence composed of key frames; determines an action primitive set according to the key video sequence and the total diffusion time step; determines a base frame from the key video sequence according to the content importance degree of the key frames in the key video sequence; performs frame derivation processing on the base frame to obtain at least one variant frame corresponding to the base frame; generates a reconstructed frame according to the action primitive set, the base frame, and at least one variant frame corresponding to the base frame; performs local detail refinement on the key video sequence according to the reconstructed frame to obtain an optimized video; and performs inter-frame dynamic interpolation processing on the optimized video to obtain a target video. The above technical solution extracts key frames from the original video to obtain a key video sequence, realizing video compression of the original video, reducing the size of the original video, thereby reducing the computational cost of processing the original video and improving the processing efficiency of the original video; then, determines an action primitive set according to the key video sequence and the total diffusion time step, providing necessary global video dynamic information for obtaining the target video in the future; then, based on the reconstructed frame determined according to the action primitive set and the key video sequence, performs local detail refinement on the key video sequence to obtain an optimized video, ensuring the richness, accuracy, and authenticity of video details in local actions or local scene transformations of the optimized video; then, obtains the target video by performing inter-frame dynamic interpolation processing on the optimized video, smoothing the difference between two adjacent video frames in the optimized video, making the visual content of the obtained target video more coherent and smoother, avoiding visual jumps or distortions during the playback of the target video, thereby improving the video quality and picture quality of the target video, enhancing the smoothness of the target video playback, and enhancing the viewing experience of the target video.

[0102] Embodiment III

[0103] Figure 3 FIG. is a schematic structural diagram of a video processing device provided in Embodiment III of the present invention. This embodiment is applicable to the case of processing videos with rich details in dynamic scenes (such as sports event videos). The device can be implemented in the form of hardware and / or software and can be configured in an electronic device. As Figure 3 shown, the device includes:

[0104] A key video sequence determination module 301, configured to extract key frames from the original video to obtain a key video sequence composed of key frames;

[0105] An action primitive set determination module 302, configured to determine an action primitive set according to the key video sequence and the total diffusion time step;

[0106] An optimized video determination module 303, configured to perform local detail refinement on the key video sequence according to the action primitive set to obtain an optimized video;

[0107] A target video determination module 304, configured to perform inter-frame dynamic interpolation processing on the optimized video to obtain a target video.

[0108] In the technical solution of the embodiment of the present invention, key frames are extracted from the original video to obtain a key video sequence composed of the key frames; an action primitive set is determined according to the key video sequence and the total diffusion time step; according to the action primitive set, local details of the key video sequence are refined to obtain an optimized video; and inter-frame dynamic interpolation processing is performed on the optimized video to obtain a target video. The above technical solution extracts key frames from the original video to obtain a key video sequence, realizes video compression of the original video, reduces the size of the original video, thereby reducing the computational cost of processing the original video and improving the processing efficiency of the original video; then, according to the key video sequence and the total diffusion time step, an action primitive set is determined, providing necessary global video dynamic information for obtaining the subsequent target video; then, according to the action primitive set, local details of the key video sequence are refined to obtain an optimized video, ensuring the richness, accuracy, and authenticity of video details in local actions or local scene transformations of the optimized video; then, by performing inter-frame dynamic interpolation processing on the optimized video to obtain a target video, the difference between two adjacent video frames before and after in the optimized video is smoothed, making the visual content of the obtained target video more coherent and smoother, avoiding visual jumps or distortions during the playback of the target video, thereby improving the video quality and picture quality of the target video and enhancing the smoothness of the playback of the target video. All in all, the above technical solution improves the picture quality of the compressed original video while realizing the compression of the original video.

[0109] Optionally, the action primitive set determination module 302 is specifically configured to:

[0110] Extract features from the key video sequence to obtain a key feature set;

[0111] Determine an initial action primitive;

[0112] Determine an action primitive set according to the key feature set, the initial action primitive, and the total diffusion time step.

[0113] Optionally, the optimized video determination module 303 includes:

[0114] A basic frame determination unit, configured to determine a basic frame from the key video sequence according to the content importance degree of the key frames in the key video sequence;

[0115] A variant frame determination unit, configured to perform frame derivation processing on the basic frame to obtain at least one variant frame corresponding to the basic frame;

[0116] A reconstructed frame generation unit, configured to generate a reconstructed frame according to an action primitive set, a base frame, and at least one variant frame corresponding to the base frame;

[0117] An optimized video determination unit, configured to perform local detail refinement on a key video sequence according to the reconstructed frame to obtain an optimized video.

[0118] Optionally, the reconstructed frame generation unit is specifically configured to:

[0119] Determine the credibility of the variant frame according to the structural similarity index, mean square error, and feature difference value between the variant frame and the base frame;

[0120] Determine the weight coefficient corresponding to each variant frame according to the credibility of each variant frame;

[0121] Generate a reconstructed frame according to the action primitive set, the base frame, at least one variant frame corresponding to the base frame, and the weight coefficient corresponding to each variant frame.

[0122] Optionally, the optimized video determination unit is specifically configured to:

[0123] Determine the frame pair to be processed in the key video sequence according to the frame difference value and the difference threshold between adjacent video frame pairs in the key video sequence;

[0124] Determine the number of frame insertions of the reconstructed frame based on the corresponding relationship between the frame difference value and the number of inserted frames in the detail refinement table according to the frame difference value between the frame pairs to be processed;

[0125] Insert the reconstructed frame between the frame pairs to be processed according to the number of frame insertions to obtain an optimized video.

[0126] Optionally, the target video determination module 304 is specifically configured to:

[0127] For each video frame pair in the optimized video, generate an intermediate frame between the previous video frame and the subsequent video frame according to the previous video frame, the subsequent video frame, and the total number of time steps between the previous video frame and the subsequent video frame in the video frame pair; wherein, the video frame pair is composed of two adjacent video frames in the optimized video.

[0128] The video processing device provided by the embodiments of the present invention can execute the video processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each video processing method.

[0129] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium, and a computer program product.

[0130] Embodiment 4

[0131] Figure 4FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0132] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0133] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0134] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the video processing method.

[0135] In some embodiments, the video processing method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the video processing method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the video processing method by any other suitable means (e.g., by means of firmware).

[0136] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0137] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0138] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0140] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0141] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0142] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0143] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video processing method, characterized in that, Including: Extracting key frames from the original video to obtain a key video sequence composed of key frames; Performing feature extraction on the key video sequence to obtain a key feature set; Determining an initial action primitive; Determining an action primitive set according to the key feature set, the initial action primitive, and the total diffusion time step length; wherein, the total diffusion time step length refers to the total duration of diffusion time steps; the diffusion time step refers to the time step used to control the number of diffusion iterations; the action primitive set refers to a set composed of multiple action primitives, which is used to represent the global dynamic information of the original video; the action primitive refers to the smallest unit of an action that can represent the key content in the original video; Performing local detail refinement on the key video sequence according to the action primitive set to obtain an optimized video; Performing inter-frame dynamic interpolation processing on the optimized video to obtain a target video.

2. The method according to claim 1, characterized in that The performing local detail refinement on the key video sequence according to the action primitive set to obtain an optimized video includes: Determining a base frame from the key video sequence according to the content importance degree of the key frames in the key video sequence; Performing frame derivation processing on the base frame to obtain at least one variant frame corresponding to the base frame; Generating a reconstructed frame according to the action primitive set, the base frame, and at least one variant frame corresponding to the base frame; Performing local detail refinement on the key video sequence according to the reconstructed frame to obtain an optimized video.

3. The method according to claim 2, wherein The generating a reconstructed frame according to the action primitive set, the base frame, and at least one variant frame corresponding to the base frame includes: Determining the credibility of the variant frame according to the structural similarity index, mean square error, and feature difference value between the variant frame and the base frame; Determining the weight coefficient corresponding to each variant frame according to the credibility of each variant frame; Generating a reconstructed frame according to the action primitive set, the base frame, at least one variant frame corresponding to the base frame, and the weight coefficient corresponding to each variant frame.

4. The method according to claim 2, wherein The performing local detail refinement on the key video sequence according to the reconstructed frame to obtain an optimized video includes: Determining a pair of frames to be processed in the key video sequence according to the frame difference value and difference threshold between adjacent video frame pairs in the key video sequence; Determining the number of frame insertions of the reconstructed frame based on the corresponding relationship between the frame difference value and the number of inserted frames in the detail refinement table according to the frame difference value between the pair of frames to be processed; Inserting the reconstructed frame between the pair of frames to be processed according to the number of frame insertions to obtain an optimized video.

5. The method according to claim 1, characterized in that, The performing inter-frame dynamic interpolation processing on the optimized video to obtain a target video includes: For each pair of video frames in the optimized video, generating an intermediate frame between the previous video frame and the subsequent video frame according to the previous video frame, the subsequent video frame, and the total number of time steps between the previous video frame and the subsequent video frame in the pair of video frames; wherein, the pair of video frames is composed of two adjacent video frames in the front and back of the optimized video.

6. A video processing device, characterized in that, Including: A key video sequence determination module, configured to extract key frames from the original video to obtain a key video sequence composed of key frames; An action primitive set determination module, configured to extract features from the key video sequence to obtain a key feature set; determine an initial action primitive; determine an action primitive set according to the key feature set, the initial action primitive, and the total diffusion time step length; wherein, the total diffusion time step length refers to the total duration of the diffusion time step; the diffusion time step refers to the time step used to control the number of diffusion iterations; the action primitive set refers to a set composed of multiple action primitives, which is used to represent the global dynamic information of the original video; the action primitive refers to the smallest unit of an action that can represent the key content in the original video; An optimized video determination module, configured to perform local detail refinement on the key video sequence according to the action primitive set to obtain an optimized video; A target video determination module, configured to perform inter-frame dynamic interpolation processing on the optimized video to obtain a target video.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video processing method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the video processing method according to any one of claims 1-5 when executed.

9. A computer program product, including a computer program, and the computer program implements the video processing method according to any one of claims 1-5 when executed by a processor.

Citation Information

Patent Citations

  • Video coding method and device, equipment and storage medium

    CN118573870A