Video coding method and system, electronic equipment and storage medium

By using forward and backward accumulated optical flows in video encoding to calculate the inter-frame motion intensity, the problem of inaccurate long-term motion estimation in bidirectional video encoding is solved, and the encoding performance and quality are improved.

CN120075457AActive Publication Date: 2025-05-30PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510205102.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing intelligent two-way video encoding method has inaccurate problems in dealing with long-distance motion estimation between long-distance frames with large motion scenes, resulting in the impact of encoding performance.

Method used

By dividing the original video sequence into multiple image groups, the forward accumulated optical flow between the video frame and the forward reference frame, and the back accumulated optical flow between the video frame and the backward reference frame, the inter-frame motion intensity is calculated based on these optical flow information, and motion estimation and encoding are performed.

Benefits of technology

Improves the accuracy of motion estimation during bidirectional video encoding, thereby improving encoding efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075457A_ABST
    Figure CN120075457A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding method and system, electronic equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: dividing an original video sequence into a plurality of image groups, and determining a forward accumulated optical flow and a backward accumulated optical flow for any video frame of any image group in the image groups; determining a forward inter-frame motion intensity based on the forward cumulative optical flow, and determining a backward inter-frame motion intensity based on the backward cumulative optical flow; determining a forward motion estimation result based on the forward inter-frame motion intensity, and determining a backward motion estimation result based on the backward inter-frame motion intensity; and coding the video frame based on the forward motion estimation result and the backward motion estimation result to obtain video coding data corresponding to the video frame. According to the invention, by introducing the forward and backward accumulated optical streams and calculating the inter-frame motion intensity based on the optical streams, the accuracy of motion estimation during bidirectional video coding can be improved, so that the coding efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a video encoding method, system, electronic device, and storage medium. Background Art

[0002] Intelligent video encoding refers to analyzing and compressing video content to improve the compression efficiency while maintaining the video quality. According to different usage scenarios, there are mainly two common configurations for video encoding: Low Delay (LD) and Random Access (RA). The LD configuration emphasizes minimizing the encoding delay and is suitable for real-time applications such as video calls and live broadcasts. In the LD configuration, Predictive (P) frames are encoded unidirectionally, that is, only the previously encoded frames are referred to during encoding. The RA configuration allows viewers to freely browse the video and is suitable for applications that require random access such as video on demand. The key frame type in the RA configuration is the Bi-directional (B) frame, and the B frame uses past and future frames as references.

[0003] However, existing intelligent bi-directional video encoding methods have inaccurate problems in processing long-term motion estimation between distant frames with large motion scenes. For example, some methods are trained on small Groups-of-Picture (GoP) with small motion intensities but tested on large GoP with various motion intensities. The domain shift problem between the training and testing scenarios causes the motion estimation module to be unable to effectively process long-term motion, resulting in the performance of bi-directional video encoding being affected.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a video encoding method, system, electronic device, and storage medium, aiming to solve the technical problem that long-term motion estimation is inaccurate during bi-directional video encoding, resulting in affected encoding performance.

[0006] To achieve the above purpose, this application proposes a video encoding method, and the method includes:

[0007] Divide the original video sequence into multiple groups of pictures. For any video frame in any group of pictures among the groups of pictures, determine the forward cumulative optical flow between the video frame and the forward reference frame, and determine the backward cumulative optical flow between the video frame and the backward reference frame;

[0008] Determine the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow, and determine the backward inter-frame motion intensity between the video frame and the backward reference frame based on the backward cumulative optical flow;

[0009] Determine the forward motion estimation result between the video frame and the reference frame based on the forward inter-frame motion intensity, and determine the backward motion estimation result between the video frame and the reference frame based on the backward inter-frame motion intensity;

[0010] Encode the video frame based on the forward motion estimation result and the backward motion estimation result to obtain the video coding data corresponding to the video frame.

[0011] In one embodiment, the step of determining the forward cumulative optical flow between the video frame and the forward reference frame includes:

[0012] Based on the first local optical flow between the video frame and the first adjacent frame, perform a warping operation on the second local optical flow between the first adjacent frame and the second adjacent frame to obtain a warped optical flow, and superimpose the warped optical flow and the second local optical flow to obtain the cumulative optical flow between the video frame and the second adjacent frame, where the first adjacent frame is a forward video frame adjacent to the video frame, and the second adjacent frame is a forward video frame adjacent to the first adjacent frame;

[0013] Update the cumulative optical flow between the video frame and the second adjacent frame to the first local optical flow, and update the third local optical flow between the second adjacent frame and the third adjacent frame to the second local optical flow, and perform the step of performing a warping operation on the second local optical flow between the first adjacent frame and the second adjacent frame based on the first local optical flow between the video frame and the first adjacent frame to obtain a warped optical flow and subsequent steps until the forward cumulative optical flow between the video frame and the forward reference frame is obtained, where the third adjacent frame is a forward video frame adjacent to the second adjacent frame.

[0014] In one embodiment, after the step of determining the forward cumulative optical flow between the video frame and the forward reference frame, the method further includes:

[0015] Perform a warping operation based on the forward reference frame to obtain a predicted reference frame;

[0016] Input the predicted reference frame, the forward reference frame, and the forward cumulative optical flow into a preset convolutional layer to obtain a first processing result;

[0017] Superimpose the first processing result and the forward cumulative optical flow to obtain a second processing result, update the forward cumulative optical flow with the second processing result, and perform the step of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow.

[0018] In one embodiment, the step of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow includes:

[0019] Map the forward cumulative optical flow to a motion intensity factor;

[0020] If the motion intensity factor is greater than or equal to a preset intensity threshold, determine that the forward inter-frame motion intensity between the video frame and the forward reference frame is high intensity;

[0021] If the motion intensity factor is less than the preset intensity threshold, determine that the forward inter-frame motion intensity is low intensity.

[0022] In one embodiment, the step of mapping the forward cumulative optical flow to a motion intensity factor includes:

[0023] Based on the forward cumulative optical flow, determine the cumulative horizontal offset and the cumulative vertical offset of each pixel point in the video frame;

[0024] Traverse each pixel point, calculate the Euclidean norm of the cumulative horizontal offset and the cumulative vertical offset to obtain the cumulative offset length between the pixel point in the video frame and the forward reference frame;

[0025] Obtain the total offset length based on the cumulative offset lengths of each pixel point, divide the total offset length by the number of pixels in the video frame to obtain the average optical flow intensity of each pixel point, and determine the average optical flow intensity as the motion intensity factor.

[0026] In one embodiment, the step of determining the forward motion estimation result between the video frame and the forward reference frame based on the forward inter-frame motion intensity includes:

[0027] If the forward inter-frame motion intensity is low intensity, input the video frame and the forward reference frame into a preset optical flow network to obtain the forward direct optical flow, and determine the forward direct optical flow as the forward motion estimation result between the video frame and the forward reference frame;

[0028] If the forward inter-frame motion intensity is high intensity, determine the forward cumulative optical flow as the forward motion estimation result.

[0029] In one embodiment, the step of encoding the video frame based on the forward motion estimation result and the backward motion estimation result to obtain the video coding data corresponding to the video frame includes:

[0030] Encoding based on the forward motion estimation result and the backward motion estimation result to obtain encoded optical flow data, and decoding the encoded optical flow data to obtain decoded optical flow data;

[0031] Generating a multi-scale temporal context based on the decoded optical flow data, the forward reference features of the forward reference frame, and the backward reference features of the backward reference frame;

[0032] Calculating the difference between the multi-scale temporal context as a condition and the video frame to obtain residual information;

[0033] Determining the forward motion estimation, the backward motion estimation, and the residual information as the video coding data corresponding to the video frame.

[0034] In addition, to achieve the above object, the present application also proposes a video coding system, which includes:

[0035] An adaptive motion estimation module, configured to divide an original video sequence into multiple groups of pictures (GOPs). For any video frame in any GOP among the GOPs, determining the forward cumulative optical flow between the video frame and a forward reference frame, and determining the backward cumulative optical flow between the video frame and a backward reference frame;

[0036] The adaptive motion estimation module is further configured to determine the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow, and determine the backward inter-frame motion intensity between the video frame and the backward reference frame based on the backward cumulative optical flow;

[0037] The adaptive motion estimation module is further configured to determine the forward motion estimation result between the video frame and the reference frame based on the forward inter-frame motion intensity, and determine the backward motion estimation result between the video frame and the reference frame based on the backward inter-frame motion intensity;

[0038] A video coding module, configured to encode the video frame based on the forward motion estimation result and the backward motion estimation result to obtain the video coding data corresponding to the video frame.

[0039] In addition, to achieve the above object, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the video coding method as described above.

[0040] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the video encoding method described above are implemented.

[0041] In the present application, an original video sequence is divided into multiple groups of pictures. For any video frame in any group of pictures among the groups of pictures, a forward cumulative optical flow between the video frame and a forward reference frame is determined, and a backward cumulative optical flow between the video frame and a backward reference frame is determined; a forward inter-frame motion intensity between the video frame and the forward reference frame is determined based on the forward cumulative optical flow, and a backward inter-frame motion intensity between the video frame and the backward reference frame is determined based on the backward cumulative optical flow. By calculating the forward cumulative optical flow between the video frame and the forward reference frame and the backward cumulative optical flow between the video frame and the backward reference frame, motion information of the video frame in different directions can be captured, so that the motion relationship between the video frame and the reference frame can be estimated, providing a basis for subsequent motion estimation and reducing the problem of inaccurate long-term motion estimation caused by domain offset.

[0042] A forward motion estimation result between the video frame and the reference frame is determined based on the forward inter-frame motion intensity, and a backward motion estimation result between the video frame and the reference frame is determined based on the backward inter-frame motion intensity; the video frame is encoded based on the forward motion estimation result and the backward motion estimation result to obtain video encoding data corresponding to the video frame. Through more accurate long-term motion estimation, the correlation between video frames can be utilized to improve the accuracy of subsequent video encoding, thereby improving the video encoding quality.

[0043] In summary, by introducing forward and backward cumulative optical flows and calculating the inter-frame motion intensity based on the optical flow, the present application can improve the accuracy of motion estimation during bidirectional video encoding, thereby improving the encoding efficiency and quality. Description of the Drawings

[0044] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0045] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0046] Figure 1Schematic flowchart provided for the first embodiment of the video encoding method of this application;

[0047] Figure 2 Schematic diagram of the bidirectional video encoding framework provided for one embodiment of this application;

[0048] Figure 3 Schematic diagram of the framework for direct optical flow estimation provided for one embodiment of this application;

[0049] Figure 4 Schematic diagram of the process for cumulative optical flow estimation provided for one embodiment of this application;

[0050] Figure 5 Schematic diagram of the framework for cumulative optical flow estimation provided for one embodiment of this application;

[0051] Figure 6 Schematic diagram of the framework for optical flow fine-tuning provided for one embodiment of this application;

[0052] Figure 7 Schematic flowchart of the adaptive motion estimation provided for one embodiment of this application;

[0053] Figure 8 Schematic diagram of the effect comparison of optical flow estimation under small motion provided for one embodiment of this application;

[0054] Figure 9 Schematic diagram of the effect comparison of optical flow estimation under large motion provided for one embodiment of this application;

[0055] Figure 10 Schematic diagram of the module structure of the video encoding system according to the embodiment of this application;

[0056] Figure 11 Schematic diagram of the device structure of the hardware operating environment involved in the video encoding method according to the embodiment of this application.

[0057] The realization of the purpose, functional features and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0058] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0059] For a better understanding of the technical solutions of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and the specific implementation manners.

[0060] The main solution of the embodiment of the present application is as follows: The original video sequence is divided into multiple groups of pictures (GOPs). For any video frame in any GOP among the GOPs, the forward cumulative optical flow between the video frame and the forward reference frame is determined, and the backward cumulative optical flow between the video frame and the backward reference frame is determined; the forward inter-frame motion intensity between the video frame and the forward reference frame is determined based on the forward cumulative optical flow, and the backward inter-frame motion intensity between the video frame and the backward reference frame is determined based on the backward cumulative optical flow; the forward motion estimation result between the video frame and the reference frame is determined based on the forward inter-frame motion intensity, and the backward motion estimation result between the video frame and the reference frame is determined based on the backward inter-frame motion intensity; the video frame is encoded based on the forward motion estimation result and the backward motion estimation result to obtain the video coding data corresponding to the video frame.

[0061] In this embodiment, for the convenience of description, the following takes an electronic device as the execution subject for elaboration.

[0062] Currently, the performance of intelligent bidirectional video coding methods lags far behind the traditional video coding standard H.266 / VVC in the RA configuration. A main reason for the poor performance of intelligent bidirectional video coding methods is the inaccurate long-term motion estimation between long-distance frames, especially in large motion scenarios. Existing intelligent bidirectional video coding methods are trained on small GOPs with small motion intensities, but tested on large GOPs with various motion intensities. The domain shift problem between the training and testing scenarios causes the motion estimation module to be unable to handle long-term motion. Therefore, existing intelligent video coding methods perform well in the LD configuration but poorly in the RA configuration. Currently, a coding scheme combining B frames and P frames is proposed to solve this problem. First, P frames are encoded with a time step of 2, and then the intermediate B frames are encoded using a video frame interpolation method. Although this method keeps the distance from the reference frame unchanged and reduces the motion variability, this method does not fully utilize the temporal correlation of the video, resulting in a decrease in compression performance. Another method is to select the most suitable frame type for each GOP, using B frame coding for GOPs with small motion and P frame coding for GOPs with large motion. The disadvantage is that making a decision for each GOP will introduce a time delay and the computational complexity is relatively high. In addition, neither of these two methods fundamentally improves the accuracy of long-term motion estimation.

[0063] The present application provides a solution. By calculating the forward cumulative optical flow between a video frame and a forward reference frame, as well as the backward cumulative optical flow between the video frame and a backward reference frame, the motion information of the video frame can be captured more comprehensively. Based on the optical flow information, the forward and backward inter-frame motion intensities are further determined, so as to more accurately estimate the long-term motion relationship between the video frame and the reference frame, thereby improving the accuracy of long-term motion estimation.

[0064] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an electronic device, etc. that can implement the above functions. Hereinafter, an electronic device will be taken as an example to illustrate this embodiment and the following embodiments.

[0065] Based on this, an embodiment of the present application provides a video coding method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the video coding method of the present application.

[0066] In this embodiment, the video coding method includes steps S10 to S40:

[0067] Step S10: Divide the original video sequence into multiple groups of pictures. For any video frame in any group of pictures in each of the groups of pictures, determine the forward cumulative optical flow between the video frame and a forward reference frame, and determine the backward cumulative optical flow between the video frame and a backward reference frame;

[0068] The original video sequence refers to a set of consecutive video frames that are uncompressed or uncoded. A group of pictures (GOP) refers to a set of several consecutive video frames into which the original video sequence is divided. In video coding, a GOP serves as the basic unit of coding. Each GOP contains one or more key frames (I-frames) and non-key frames (such as P-frames and B-frames). Key frames can be decoded independently, while non-key frames need to rely on other frames (such as the previous I-frame or P-frame) for decoding. In this embodiment, any video frame in any GOP is a non-key frame in the GOP. It should be noted that the length and composition of a GOP can be adjusted according to the coding strategy and video content, and are not limited herein. Exemplarily, for a group of GOPs including 32 frames of images, the process of frame coding is as follows: using the first frame in the GOP as the forward reference frame and the first frame in the next adjacent GOP (i.e., the 33rd frame in the video sequence) as the backward reference frame, encode the 17th frame, where the first frame and the 33rd frame are encoded as key frames; using the first frame in the GOP as the forward reference frame and the 17th frame as the backward reference frame, encode the 9th frame, using the 17th frame as the forward reference frame and the 33rd frame as the backward reference frame, encode the 25th frame; and so on until all 32 video frames are encoded. It should be noted that for the last GOP, since there is no next adjacent GOP, after unidirectionally encoding the last frame in this GOP with the first frame in this GOP as the reference frame, use the encoded last frame as the reference frame for the above encoding process.

[0069] A forward reference frame refers to a frame used for forward prediction of a video frame during video coding. The forward cumulative optical flow can reflect the overall movement trend of pixels in the image from the forward reference frame to the video frame; a backward reference frame refers to a frame used for backward prediction of a video frame during video coding; the backward cumulative optical flow can reflect the overall movement trend of pixels in the image from the backward reference frame to the video frame.

[0070] Step S20, determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow, and determining the backward inter-frame motion intensity between the video frame and the backward reference frame based on the backward cumulative optical flow;

[0071] The forward inter-frame motion intensity refers to the motion intensity or degree of motion change between a video frame and a forward reference frame. The forward inter-frame motion intensity includes high intensity and low intensity. The backward inter-frame motion intensity refers to the motion intensity or degree of motion change between a video frame and a backward reference frame. The backward inter-frame motion intensity includes high intensity and low intensity.

[0072] Calculate the forward inter-frame motion intensity between the video frame and the forward reference frame according to parameters such as the amplitude or velocity of the forward cumulative optical flow. Similarly, calculate the backward inter-frame motion intensity between the video frame and the backward reference frame according to parameters such as the amplitude or velocity of the backward cumulative optical flow. By calculating the inter-frame motion intensity, a basis can be provided for motion estimation and encoding in subsequent steps.

[0073] Step S30: Determine the forward motion estimation result between the video frame and the reference frame based on the forward inter-frame motion intensity, and determine the backward motion estimation result between the video frame and the reference frame based on the backward inter-frame motion intensity;

[0074] The forward motion estimation result refers to the optical flow output by using an optical flow estimation network (such as SpyNet) based on the forward inter-frame motion intensity, and the backward motion estimation result refers to the optical flow output by using the optical flow estimation network based on the backward inter-frame motion intensity.

[0075] Step S40: Encode the video frame based on the forward motion estimation result and the backward motion estimation result to obtain the video coding data corresponding to the video frame.

[0076] Encode the video frame according to the forward and backward motion estimation results and information such as the forward and backward reference frames. The specific encoding process is not elaborated here. After encoding, generate the video coding data corresponding to the video frame for subsequent storage and transmission. By encoding based on the motion estimation result, redundant information between video frames can be more effectively removed, thereby improving the encoding efficiency and compression ratio.

[0077] In a feasible embodiment, the step S40 of encoding the video frame based on the forward motion estimation result and the backward motion estimation result to obtain the video coding data corresponding to the video frame includes:

[0078] Step S401: Encode based on the forward motion estimation result and the backward motion estimation result to obtain encoded optical flow data, and decode the encoded optical flow data to obtain decoded optical flow data;

[0079] Generate optical flow field data according to the motion vectors obtained from the forward and backward motion estimations, compress the optical flow field data to obtain encoded optical flow data. Then decode the encoded optical flow data to restore the original or approximately original optical flow field information to obtain decoded optical flow data. By encoding and decoding the optical flow data, the motion information between video frames can be efficiently transmitted during the transmission process to meet the requirements of video data transmission and processing.

[0080] In this embodiment, given the original video sequence X = {x 0 ,x 1 ,…,xt , …}, where x t represents the frame at time t. X is divided into several groups of pictures (GoP) and compressed separately. Each GoP contains 32 video frames. All the video frames of the K-th GoP unit X K are used as input, and the optical flow between the video frame x t and two original reference frames x t-i , x t+i is estimated The estimated original motion is jointly encoded and transmitted to obtain the reconstructed motion after decoding Among them, the optical flow codec is implemented based on an autoencoder

[0081] Step S402: Generate multi-scale temporal context based on the decoded optical flow data, the forward reference features of the forward reference frame, and the backward reference features of the backward reference frame

[0082] Feature information such as edges, textures, color histograms, etc. is extracted from the forward and backward reference frames. Combining the decoded optical flow data and the reference frame features, a multi-scale analysis method (such as a pyramid structure, wavelet transform, etc.) is used to construct context information at different temporal scales. Specifically, based on the decoded motion the reference features are warped to generate the multi-scale temporal context

[0083]

[0084] Step S403: Calculate the difference between the video frame and the multi-scale temporal context as the conditional residual information

[0085] Conditioned on the multi-scale temporal context the video frame x t is compressed and decoded by the B-frame codec into the intermediate feature F t and the reconstructed frame where F t is referenced when encoding or decoding the next frame

[0086] Step S404: Determine the forward motion estimation, the backward motion estimation, and the residual information as the video coding data corresponding to the video frame

[0087] The forward motion estimation, the backward motion estimation, and the residual information are determined as the video coding data corresponding to the video frame. It can be understood that the residual information can reflect the information loss and reconstruction error during the encoding process. Taking the encoded video data and the residual information together as the video coding data can improve the decoded video quality and encoding efficiency

[0088] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be elaborated hereinafter. On this basis, the step S10: the step of determining the forward cumulative optical flow between the video frame and the forward reference frame includes:

[0089] Step S101, based on the first local optical flow between the video frame and the first adjacent frame, perform a warping operation on the second local optical flow between the first adjacent frame and the second adjacent frame to obtain a warped optical flow, and superimpose the warped optical flow and the second local optical flow to obtain the cumulative optical flow between the video frame and the second adjacent frame, where the first adjacent frame is a forward video frame adjacent to the video frame, and the second adjacent frame is a forward video frame adjacent to the first adjacent frame;

[0090] Based on the local optical flow information between a video frame and its directly adjacent forward video frame (the first adjacent frame), perform a warping operation on the local optical flow between the first adjacent frame and its previous video frame (the second adjacent frame) to simulate the propagation of the optical flow between consecutive frames. Then, superimpose the warped optical flow and the original local optical flow between the second adjacent frame to obtain the cumulative optical flow between the video frame and the second adjacent frame.

[0091] Step S102, update the cumulative optical flow between the video frame and the second adjacent frame to the first local optical flow, and update the third local optical flow between the second adjacent frame and the third adjacent frame to the second local optical flow, and execute the step of performing a warping operation on the second local optical flow between the first adjacent frame and the second adjacent frame based on the first local optical flow between the video frame and the first adjacent frame to obtain a warped optical flow and subsequent steps until the forward cumulative optical flow between the video frame and the forward reference frame is obtained, where the third adjacent frame is a forward video frame adjacent to the second adjacent frame.

[0092] At the beginning of each iteration, update the local optical flow between the current video frame and the adjacent frame to the last part of the cumulative optical flow calculated in the previous iteration (i.e., the local optical flow of the previous frame directly adjacent to the current video frame), and update the local optical flow between the next adjacent frame and the next adjacent frame to the local optical flow between the adjacent frames considered in the current iteration. Based on the updated local optical flow information, repeat the warping and superimposing operations to calculate the cumulative optical flow between the current video frame and the next adjacent frame until the forward reference frame is reached. At this time, the obtained cumulative optical flow is the forward cumulative optical flow between the video frame and the forward reference frame.

[0093] To accurately estimate long-term motion, this embodiment proposes an optical flow estimation method based on accumulation, that is, recursively accumulating the local optical flow between adjacent frames. Exemplarily, directly estimate the original adjacent frames {x 0, x 1 , x 2 , x 3 , x 4 The local optical flow {v 1→0 , v 2→1 , v 3→2 , v 4→3} between... In step 1, in order to obtain the optical flow v 2 between x 0 and x 2→0 , v 1→0 is warped to v 2→1 as follows:

[0094] v 2→0 = W(v 1→0 , v 2→1 ) + v 2→1

[0095] where W represents the warping operation. In steps 2 and 3, the target optical flow v 4→0 is recursively accumulated as follows:

[0096] v 3→0 = W(v 2→0 , v 3→2 ) + v 3→2

[0097] v 4→0 = W(v 3→0 , v 4→3 ) + v 4→3

[0098] In a feasible embodiment, after the step S10: determining the forward cumulative optical flow between the video frame and the forward reference frame, the following is further included:

[0099] Step S50, performing a warping operation based on the forward reference frame to obtain a predicted reference frame;

[0100] Due to the influence of occlusion, the optical flow accumulation process will bring cumulative errors, resulting in inaccurate optical flow estimation in non-occluded areas. To solve this problem, this embodiment designs an Optical Flow Refinement (OFR) module to correct the errors in each accumulation process.

[0101] Specifically, based on the forward cumulative optical flow information, a warping operation is performed on the forward reference frame to generate a predicted reference frame. The predicted reference frame is an attempt to simulate the possible state of the video frame on the forward time axis, providing an approximate representation of the video frame based on forward motion for subsequent motion estimation and coding.

[0102] Step S60: Input the predicted reference frame, the forward reference frame, and the forward cumulative optical flow into a preset convolutional layer to obtain a first processing result;

[0103] Use the predicted reference frame, the forward reference frame, and the forward cumulative optical flow as inputs, and process them through a preset convolutional layer to extract features and generate a first processing result. Feature fusion can capture the internal connections between the predicted reference frame, the forward reference frame, and the forward cumulative optical flow, improving the accuracy of subsequent processing.

[0104] Step S70: Superimpose the first processing result and the forward cumulative optical flow to obtain a second processing result, update the forward cumulative optical flow with the second processing result, and perform the step of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow.

[0105] Perform a pixel-level superimposition operation on the first processing result and the forward cumulative optical flow, output the superimposed result as the second processing result, update the forward cumulative optical flow according to the second processing result to reflect more accurate motion information. Use the updated forward cumulative optical flow to perform subsequent steps to improve the overall coding efficiency and video quality.

[0106] Specifically, based on the accumulated optical flow First, distort x t-i to obtain the predicted frame of video frame x t Then, design a convolutional network consisting of 3 3×3 convolutional layers to refine the initial cumulative optical flow: where Conv represents the convolutional layer.

[0107]

[0108] In a feasible embodiment, the step S20 of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow includes:

[0109] Step S201: Map the forward cumulative optical flow to a motion intensity factor;

[0110] Convert the forward cumulative optical flow (i.e., the accumulated motion information between the video frame and the forward reference frame) into a quantization index, that is, a motion intensity factor. The motion intensity factor is used to measure the severity of the motion of the video frame relative to the forward reference frame. The specific mapping rule can be set according to actual needs and is not limited here.

[0111]

[0112] ​Step S202, if the motion intensity factor is greater than or equal to a preset intensity threshold, determine that the forward inter-frame motion intensity between the video frame and the forward reference frame is high intensity;

[0113] Compare the motion intensity factor with a preset intensity threshold. If the motion intensity factor is large, it is considered that the forward inter-frame motion intensity between the video frame and the forward reference frame is high intensity. The intensity threshold can be determined according to the application scenario and coding requirements and is not limited here.

[0114] Step S203, if the motion intensity factor is less than the preset intensity threshold, determine that the forward inter-frame motion intensity is low intensity.

[0115] If the motion intensity factor is small, it is considered that the forward inter-frame motion intensity between the video frame and the forward reference frame is low intensity. It can be understood that by quantifying the motion intensity, comparing the preset threshold and making a determination, it provides an important basis for the selection of subsequent coding strategies.

[0116] In a feasible embodiment, the step S201: the step of mapping the forward cumulative optical flow to a motion intensity factor includes:

[0117] Step S2011, based on the forward cumulative optical flow, determine the respective cumulative horizontal offset and cumulative vertical offset of each pixel point in the video frame;

[0118] Extract the motion vector of each pixel point from the forward cumulative optical flow data, decompose the motion vector of each pixel point into a horizontal component and a vertical component, which respectively represent the offset amount of the pixel point in the horizontal direction and the vertical direction. For each pixel point, accumulate the offset amounts of all frames on its forward time axis to obtain the cumulative horizontal offset and the cumulative vertical offset.

[0119] Step S2012, traverse each of the pixel points, calculate the Euclidean norm of the cumulative horizontal offset and the cumulative vertical offset to obtain the cumulative offset length between the pixel point in the video frame and the forward reference frame;

[0120] Traverse each pixel point in the video frame in turn. For each pixel point, calculate the cumulative offset length according to its cumulative horizontal offset and cumulative vertical offset using the Euclidean norm formula. The cumulative offset length can reflect the overall motion distance of each pixel point relative to the forward reference frame and provide a basis for calculating the average optical flow intensity later.

[0121] Step S2013, obtain the total offset length based on the cumulative offset lengths of each of the pixel points, divide the total offset length by the number of pixels in the video frame to obtain the average optical flow intensity of each of the pixel points, and determine the average optical flow intensity as the motion intensity factor.

[0122] Add up the cumulative offset lengths of all pixels to obtain the total offset length. Divide the total offset length by the number of pixels in the video frame to get the average optical flow intensity, and output the calculated average optical flow intensity as the motion intensity factor.

[0123] That is, define the calculation method of the motion intensity factor as:

[0124]

[0125] where m x and m y represent the horizontal and vertical offsets of the cumulative optical flow, and H and W represent the height and width of the frame. If the average MIF of the bidirectional optical flow is less than the given threshold T, it is defined as small motion and the direct optical flow estimation method is used. Otherwise, it is defined as large motion and the cumulative-based optical flow estimation method is used. Exemplarily, the threshold T can be set to 10.

[0126] In a feasible embodiment, the step S30: the step of determining the forward motion estimation result between the video frame and the forward reference frame based on the forward inter-frame motion intensity includes:

[0127] Step S301, if the forward inter-frame motion intensity is low intensity, input the video frame and the forward reference frame into a preset optical flow network to obtain the forward direct optical flow, and determine the forward direct optical flow as the forward motion estimation result between the video frame and the forward reference frame;

[0128] Take the video frame and the forward reference frame as input data and send them into a preset optical flow network. The optical flow network extracts features and matches the two input frames through internal convolutional layers, activation functions and other structures, and finally outputs the optical flow information between the two frames, that is, the forward direct optical flow. As Figure 6 shown, the direct estimation method directly inputs two frames x t , x t-i into the optical flow estimation network to estimate the optical flow This method performs well in small motion scenarios.

[0129] Step S302, if the forward inter-frame motion intensity is high intensity, determine the forward cumulative optical flow as the forward motion estimation result.

[0130] Obtain the already calculated forward cumulative optical flow, and directly use the forward cumulative optical flow as the forward motion estimation result between the video frame and the forward reference frame. For high-intensity motion scenarios, using cumulative optical flow can capture motion information more comprehensively and avoid errors that may occur in the optical flow network under complex motions.

[0131] It is understandable that different forward motion estimation methods are selected according to the different intensities of forward inter-frame motion. In low-intensity scenarios, an optical flow network is used for accurate estimation; in high-intensity scenarios, the cumulative optical flow is directly utilized for estimation. By flexibly choosing appropriate motion estimation methods according to the characteristics of different scenarios, the performance and accuracy of video coding can be improved.

[0132] Exemplarily, to facilitate understanding of the implementation process of the video coding method obtained based on the above embodiments, please refer to Figure 2 , the framework diagram of the intelligent bidirectional video coding method based on this embodiment is as shown in Figure 2 . Given the original video sequence X = {x 0 , x 1 , …, x t , …}, where x t represents the frame at time t. X is divided into several groups of pictures (GoP) and compressed separately. A GoP contains 32 video frames. The intelligent bidirectional video coding method mainly consists of the following units:

[0133] Adaptive Motion Estimation (AME). The AME module takes all the video frames of the Kth GoP unit X K as input and estimates the bidirectional optical flow between the current frame x t and two original reference frames x t-i , x t+i . In this implementation, SpyNet is used as the optical flow estimation network.

[0134] Optical Flow Codec. The estimated original motion is jointly encoded and transmitted to decode and obtain the reconstructed motion . Among them, the optical flow codec is implemented based on an autoencoder.

[0135] Temporal Context Mining. Based on the decoded motion to warp and transform the reference features to generate multi-scale temporal context

[0136] B-frame Codec. Conditional on the multi-scale temporal context , the current frame x t is compressed and decoded by the B-frame codec into intermediate features F t and the reconstructed frame . Among them, F t is referenced when encoding or decoding the next frame.

[0137] The AME method proposed in this embodiment consists of two motion estimation methods: direct optical flow estimation and cumulative optical flow estimation. As Figure 3 shown, the direct estimation method directly inputs two frames x t ,x t-i into the optical flow estimation network to estimate the optical flow This method performs well in small motion scenarios. To accurately estimate long-term motion, this embodiment proposes a cumulative optical flow estimation method, that is, recursively accumulating the local optical flow between adjacent frames. Please refer to Figure 4 , Figure 4 which provides an example of the optical flow accumulation process for estimating the optical flow between x 4 and x 0 . First, directly estimate the local optical flow {v 0 ,v 1 ,v 2 ,v 3 ,v 4} between the original adjacent frames {x 1→0 ,x 2→1 ,x 3→2 ,x 4→3}. In step 1, to obtain the optical flow v 2 between x 0 and x 2→0 , warp v 1→0 to v 2→1 in the following way:

[0138] v 2→0 =W(v 1→0 ,v 2→1 )+v 2→1

[0139] where W represents the warping operation. In steps 2 and 3, recursively accumulate to generate the target optical flow v 4→0 in the following way:

[0140] v 3→0 =W(v 2→0 ,v 3→2 )+v 3→2

[0141] v 4→0 =W(v 3→0 ,v 4→3 )+v 4→3

[0142] The framework diagram of the cumulative optical flow estimation method is as shown in Figure 5As shown. Due to the influence of occlusion, the optical flow accumulation process will introduce cumulative errors, resulting in inaccurate optical flow estimation in non-occluded areas. To solve this problem, this embodiment also designs an Optical Flow Refinement (OFR) module to correct the errors in each accumulation process. The detailed network structure of the optical flow refinement module is as Figure 6 shown. Based on the accumulated optical flow First, distort x t-i to obtain the predicted frame of the current frame x t Then, design a convolutional network consisting of 3 3×3 convolutional layers to refine the initial accumulated optical flow:

[0143]

[0144] where Conv represents the convolutional layer.

[0145] Meanwhile, define the Motion Intensity Factor (MIF) to evaluate the motion intensity between two frames, and its calculation method is:

[0146]

[0147] where m x and m y represent the horizontal and vertical offsets of the accumulated optical flow, and H and W represent the height and width of the frame. If the average MIF of the bidirectional optical flow is less than the given threshold T, it is defined as small motion and the direct optical flow estimation method is used. Otherwise, it is defined as large motion and the optical flow estimation method based on accumulation is used. In this embodiment, the threshold T is set to 10, Figure 7 shows the flowchart of the adaptive motion estimation method.

[0148] However, as Figure 8 shown, as the time interval increases, due to the domain shift problem, the accuracy of optical flow estimation drops severely or even becomes unacceptable. As Figure 9 shown, the AME module proposed in this embodiment achieves accurate motion estimation and prediction results in both small motion and large motion scenarios.

[0149] ​In this embodiment, ablation experiments were conducted on the UVG, MCL-JCV, and HEVC datasets to prove the effectiveness of this embodiment. These datasets contain videos with different motion intensities and resolutions and are widely used to evaluate the performance of intelligent video coding methods. Table 1 calculates the BD-Rate (%) comparison with PSNR as the metric. Among them, Method A uses direct optical flow estimation and is set as the anchor point. In this embodiment, the optical flow estimation method of Method A is replaced with adaptive motion estimation. Through comparison, it can be found that the AME method proposed in this embodiment achieves an average bitrate saving of 7.9%, which proves the effectiveness of the AME method proposed in this embodiment. Since the AME method directly estimates small motions, no gain is obtained on the HEVC E dataset that only contains small motions.

[0150] Table 1: Ablation experiments on different datasets

[0151]

[0152] In addition, this embodiment also conducted ablation studies on video sequences of different motion types. Among them, the direct estimation method is set as the anchor point. As shown in Table 2, the cumulative optical flow estimation method brings performance improvement in large motion scenarios but suffers performance loss in small motion scenarios. Adding an optical flow refinement module to the cumulative-based method can bring significant performance improvement. This is because the OFR module can effectively correct the cumulative error. The comparison experiment results show that the AME method proposed in this embodiment can effectively handle motions of different intensities.

[0153] Table 2: Ablation experiments on datasets with different motion intensities

[0154]

[0155]

[0156] To find the key factors leading to the low compression performance of existing intelligent bidirectional video compression methods, the motion estimation and prediction results of the prior art B-CANF were visualized in Figure 1 . It can be seen from the figure that the motion estimation of B-CANF performs well in small motion scenarios. However, when the frame distance increases, B-CANF cannot effectively handle the large motion between distant frames, resulting in inaccurate prediction.

[0157] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the video coding method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0158] This application also provides a video coding system. Please refer to Figure 10 , the video coding system includes:

[0159] The adaptive motion estimation module 10 is configured to divide an original video sequence into multiple groups of pictures. For any video frame in any group of pictures, determine the forward cumulative optical flow between the video frame and a forward reference frame, and determine the backward cumulative optical flow between the video frame and a backward reference frame;

[0160] The adaptive motion estimation module 10 is further configured to determine the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow, and determine the backward inter-frame motion intensity between the video frame and the backward reference frame based on the backward cumulative optical flow;

[0161] The adaptive motion estimation module 10 is further configured to determine the forward motion estimation result between the video frame and the reference frame based on the forward inter-frame motion intensity, and determine the backward motion estimation result between the video frame and the reference frame based on the backward inter-frame motion intensity;

[0162] The video encoding module 20 is configured to encode the video frame based on the forward motion estimation result and the backward motion estimation result to obtain video encoding data corresponding to the video frame.

[0163] Optionally, the adaptive motion estimation module 10 is further configured to:

[0164] Perform a warping operation on a second local optical flow between the first adjacent frame and a second adjacent frame based on a first local optical flow between the video frame and the first adjacent frame to obtain a warped optical flow, and superimpose the warped optical flow and the second local optical flow to obtain the cumulative optical flow between the video frame and the second adjacent frame, where the first adjacent frame is a forward video frame adjacent to the video frame, and the second adjacent frame is a forward video frame adjacent to the first adjacent frame;

[0165] Update the cumulative optical flow between the video frame and the second adjacent frame to the first local optical flow, and update the third local optical flow between the second adjacent frame and a third adjacent frame to the second local optical flow, and perform the step of performing a warping operation on the second local optical flow between the first adjacent frame and the second adjacent frame based on the first local optical flow between the video frame and the first adjacent frame to obtain a warped optical flow and subsequent steps until the forward cumulative optical flow between the video frame and the forward reference frame is obtained, where the third adjacent frame is a forward video frame adjacent to the second adjacent frame.

[0166] Optionally, the adaptive motion estimation module 10 is further configured to:

[0167] Perform a warping operation based on the forward reference frame to obtain a predicted reference frame;

[0168] Input the prediction reference frame, the forward reference frame, and the forward cumulative optical flow into a preset convolutional layer to obtain a first processing result;

[0169] Superimpose the first processing result and the forward cumulative optical flow to obtain a second processing result, update the forward cumulative optical flow with the second processing result, and perform the step of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward cumulative optical flow.

[0170] The adaptive motion estimation module 10 is further configured to:

[0171] Determine the forward inter-frame motion intensity between the video frame and the forward reference frame based on the second processing result.

[0172] Optionally, the adaptive motion estimation module 10 is further configured to:

[0173] Map the forward cumulative optical flow to a motion intensity factor;

[0174] If the motion intensity factor is greater than or equal to a preset intensity threshold, determine that the forward inter-frame motion intensity between the video frame and the forward reference frame is high intensity;

[0175] If the motion intensity factor is less than the preset intensity threshold, determine that the forward inter-frame motion intensity is low intensity.

[0176] Optionally, the adaptive motion estimation module 10 is further configured to:

[0177] Based on the forward cumulative optical flow, determine the respective cumulative horizontal offset and cumulative vertical offset of each pixel point in the video frame;

[0178] Traverse each pixel point, calculate the Euclidean norm of the cumulative horizontal offset and the cumulative vertical offset to obtain the cumulative offset length between the pixel point in the video frame and the forward reference frame;

[0179] Obtain the total offset length based on the cumulative offset lengths of each pixel point, divide the total offset length by the number of pixels in the video frame to obtain the average optical flow intensity of each pixel point, and determine the average optical flow intensity as the motion intensity factor.

[0180] Optionally, the adaptive motion estimation module 10 is further configured to:

[0181] If the forward inter-frame motion intensity is low intensity, input the video frame and the forward reference frame into a preset optical flow network to obtain the forward direct optical flow, and determine the forward direct optical flow as the forward motion estimation result between the video frame and the forward reference frame;

[0182] If the forward inter-frame motion intensity is high intensity, determine the forward cumulative optical flow as the forward motion estimation result.

[0183] Optionally, the video encoding module 20 includes:

[0184] An optical flow encoding and decoding unit, configured to encode to obtain encoded optical flow data based on the forward motion estimation result and the backward motion estimation result, and decode the encoded optical flow data to obtain decoded optical flow data;

[0185] A temporal context mining unit, configured to generate a multi-scale temporal context based on the decoded optical flow data, the forward reference features of the forward reference frame, and the backward reference features of the backward reference frame;

[0186] A bidirectional frame encoding and decoding unit, configured to calculate the difference between the multi-scale temporal context as a condition and the video frames to obtain residual information;

[0187] A video frame encoding unit, configured to determine the forward motion estimation, the backward motion estimation, and the residual information as the video encoding data corresponding to the video frame.

[0188] The video encoding system provided by the present application adopts the video encoding method in the above embodiment, and can solve the technical problem that the long-term motion estimation is inaccurate during bidirectional video encoding, resulting in the influence on the encoding performance. Compared with the prior art, the beneficial effects of the video encoding system provided by the present application are the same as those of the video encoding method provided by the above embodiment, and other technical features in the video encoding system are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0189] The present application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video encoding method in the above embodiment.

[0190] Next, refer to Figure 11 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as laptop computers, digital broadcast receivers, PADs (Portable Application Description: tablet computers), PMPs (Portable Media Players: portable multimedia players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0191] As Figure 11 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.

[0192] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0193] The electronic device provided by the present application, adopting the video encoding method in the above embodiments, can solve the technical problem that the long-term motion estimation is inaccurate during two-way video encoding, resulting in the encoding performance being affected. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the video encoding method provided by the above embodiments, and other technical features in the electronic device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated herein.

[0194] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0195] As mentioned above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0196] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the video encoding method in the above embodiments.

[0197] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0198] The above computer-readable storage medium can be included in an electronic device; or it can exist separately without being assembled into the electronic device.

[0199] The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the embodiments of the above video encoding method.

[0200] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0201] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0202] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0203] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned video coding method, which can solve the technical problem that the long-term motion estimation is inaccurate during bidirectional video coding, resulting in the impact on the coding performance. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the video coding method provided by the above embodiments, and will not be elaborated here.

[0204] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the video encoding method as described above.

[0205] The computer program product provided by the present application can solve the technical problem that long-term motion estimation is inaccurate during bidirectional video encoding, resulting in the encoding performance being affected. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the video encoding method provided in the above embodiments, and will not be elaborated herein.

[0206] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A video encoding method, characterized in that: The video encoding method comprises: Dividing the original video sequence into a plurality of image groups, for any video frame of any image group in each of the image groups, determining a forward accumulated optical flow between the video frame and a forward reference frame, and determining a backward accumulated optical flow between the video frame and a backward reference frame; Determine a forward inter-frame motion strength between the video frame and the forward reference frame based on the forward accumulated optical flow, and determine a backward inter-frame motion strength between the video frame and the backward reference frame based on the backward accumulated optical flow; Determine a forward motion estimation result between the video frame and the reference frame based on the forward inter-frame motion strength, and determine a backward motion estimation result between the video frame and the reference frame based on the backward inter-frame motion strength; The video frame is encoded based on the forward motion estimation result and the backward motion estimation result to obtain video encoding data corresponding to the video frame.

2. The video encoding method according to claim 1, characterized in that: The step of determining the forward accumulated optical flow between the video frame and the forward reference frame comprises: Based on a first local optical flow between the video frame and a first adjacent frame, a second local optical flow between the first adjacent frame and a second adjacent frame is distorted to obtain a distorted optical flow, and the distorted optical flow and the second local optical flow are superimposed to obtain a cumulative optical flow between the video frame and the second adjacent frame, wherein the first adjacent frame is a forward video frame adjacent to the video frame, and the second adjacent frame is a forward video frame adjacent to the first adjacent frame; The accumulated optical flow between the video frame and the second adjacent frame is updated to the first local optical flow, and the third local optical flow between the second adjacent frame and the third adjacent frame is updated to the second local optical flow, and the step of performing a warping operation on the second local optical flow between the first adjacent frame and the second adjacent frame based on the first local optical flow between the video frame and the first adjacent frame to obtain a warped optical flow and subsequent steps are performed until a forward accumulated optical flow between the video frame and the forward reference frame is obtained, wherein the third adjacent frame is a forward video frame adjacent to the second adjacent frame.

3. The video encoding method according to claim 2, characterized in that: After the step of determining the forward accumulated optical flow between the video frame and the forward reference frame, the method further includes: Performing a warping operation based on the forward reference frame to obtain a predicted reference frame; Inputting the predicted reference frame, the forward reference frame and the forward accumulated optical flow into a preset convolution layer to obtain a first processing result; The first processing result and the forward accumulated optical flow are superimposed to obtain a second processing result, the forward accumulated optical flow is updated according to the second processing result, and the step of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward accumulated optical flow is performed.

4. The video encoding method according to claim 3, characterized in that: The step of determining the forward inter-frame motion intensity between the video frame and the forward reference frame based on the forward accumulated optical flow comprises: Mapping the forward accumulated optical flow into a motion intensity factor; If the motion intensity factor is greater than or equal to a preset intensity threshold, determining that the forward inter-frame motion intensity between the video frame and the forward reference frame is high intensity; If the motion intensity factor is less than the preset intensity threshold, the forward inter-frame motion intensity is determined to be low intensity.

5. The video encoding method according to claim 4, characterized in that: The step of mapping the forward accumulated optical flow into a motion intensity factor comprises: Determine, based on the forward accumulated optical flow, an accumulated horizontal offset and an accumulated vertical offset of each pixel point in the video frame; Traversing each of the pixel points, calculating the Euclidean norm of the cumulative horizontal offset and the cumulative vertical offset, and obtaining the cumulative offset length of the pixel point between the video frame and the forward reference frame; A total offset length is obtained based on the cumulative offset length of each pixel point, and an average optical flow intensity of each pixel point is obtained by dividing the total offset length by the number of pixels of the video frame. The average optical flow intensity is determined as the motion intensity factor.

6. The video encoding method according to claim 4, characterized in that: The step of determining the forward motion estimation result between the video frame and the forward reference frame based on the forward inter-frame motion strength comprises: If the forward inter-frame motion intensity is low, inputting the video frame and the forward reference frame into a preset optical flow network to obtain a forward direct optical flow, and determining the forward direct optical flow as a forward motion estimation result between the video frame and the forward reference frame; If the forward inter-frame motion intensity is high, the forward accumulated optical flow is determined as the forward motion estimation result.

7. The video encoding method according to any one of claims 1 to 6, characterized in that: The step of encoding the video frame based on the forward motion estimation result and the backward motion estimation result to obtain video encoding data corresponding to the video frame includes: Encoding based on the forward motion estimation result and the backward motion estimation result to obtain encoded optical flow data, and decoding the encoded optical flow data to obtain decoded optical flow data; generating a multi-scale temporal context based on the decoded optical flow data, the forward reference features of the forward reference frame, and the backward reference features of the backward reference frame; Calculating the difference between the multi-scale temporal context as a condition and the video frame to obtain residual information; The forward motion estimation, the backward motion estimation and the residual information are determined as video coding data corresponding to the video frame.

8. A video encoding system, characterized in that: The video encoding system comprises: An adaptive motion estimation module, configured to divide an original video sequence into a plurality of image groups, and for any video frame of any image group in each of the image groups, determine a forward accumulated optical flow between the video frame and a forward reference frame, and determine a backward accumulated optical flow between the video frame and a backward reference frame; The adaptive motion estimation module is further used to determine the forward inter-frame motion strength between the video frame and the forward reference frame based on the forward accumulated optical flow, and to determine the backward inter-frame motion strength between the video frame and the backward reference frame based on the backward accumulated optical flow; The adaptive motion estimation module is further used to determine a forward motion estimation result between the video frame and the reference frame based on the forward inter-frame motion strength, and to determine a backward motion estimation result between the video frame and the reference frame based on the backward inter-frame motion strength; A video encoding module is used to encode the video frame based on the forward motion estimation result and the backward motion estimation result to obtain video encoding data corresponding to the video frame.

9. An electronic device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the video encoding method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the video encoding method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Video coding method and device, electronic equipment and storage medium

    CN112954348A

  • Apparent motion joint weak and small motion target detection method in combination with inter-frame optical flow

    CN113936034A

  • Video image processing method and device, electronic equipment and storage medium

    CN117395423A

  • Neural Network-Based Video Compression with Spatial-Temporal Adaptation

    US20220394240A1

  • Block motion video coding and decoding

    WO2000019725A1