Video Encoding System LTR Frame Reference for Packet Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video call technologies face challenges in maintaining a balance between video quality and smoothness due to high burst packet loss rates and network congestion, leading to frame freezing and poor user experience, especially in scenarios with weak signals or high latency.
Innovation Solution
A method for encoding and transmitting video frames that includes indicating inter-frame reference relationships in the bitstream, allowing for adaptive selection of reference frames to shorten inter-frame reference distances and improve coding quality, even in poor network conditions, by dynamically determining the LTR frame marking interval based on network status and motion scene.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Duration of action of moving object
If a transmit end re-encodes an I frame to restore a smooth video picture when data is incomplete, then video smoothness is improved, but frame freezing occurs and video definition deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-marking certain video frames as Long-Term Reference (LTR) frames before transmission. These LTR frames are identified and marked in advance based on motion scene analysis and network status, so that when packet loss occurs, the receiver can immediately use the pre-designated LTR frames for recovery without waiting for re-encoding operations, thus avoiding frame freezing while maintaining video definition
Solution Approach 2:
The patent introduces LTR frames as intermediary elements that mediate between the transmitted video data and the reconstruction process at the receiver. These LTR frames serve as stable reference points that can be used to recover from packet loss without requiring re-encoding operations, thus resolving the contradiction between smoothness and definition by providing a intermediate recovery mechanism
2Stability of the object's composition
If video encoding and decoding and transmission control are performed by two separate subsystems with stable reference relationships, then system stability is improved, but video smoothness and video definition cannot be balanced
Solution Approach 1:
The patent merges the functions of video encoding, decoding, and transmission control into a unified system that jointly optimizes for both smoothness and definition. By combining these previously separate subsystems, the invention can dynamically coordinate reference frame selection, encoding parameters, and transmission strategies to achieve balanced video quality adaptation
Solution Approach 2:
The patent introduces dynamics by making the reference frame relationships adaptive rather than stable. The system dynamically adjusts which frames are marked as LTR frames based on real-time motion scene analysis and network conditions, allowing the reference relationships to change flexibly to balance smoothness and definition requirements under different operating conditions
Data Source
AI summary
Embodiments of this disclosure provide a method for transmitting a video picture, a device for sending a video picture, and a video call method and device, where the video picture includes a plurality of video frames. According to the method for transmitting a video picture, the plurality of video frames are encoded to obtain an encoded bitstream, where the bitstream includes at least information indicating inter-frame reference relationships. In this application, each of N frames preceding to a current frame references a forward long-term reference (LTR) frame that has a shortest temporal distance to each of the preceding N frames, and the forward LTR frame is an encoded video frame marked by a transmit end device as an LTR frame.


