Video processing method and related products

By generating synchronous packet code streams in the RTC communication system, the server only sends keyframe and adjacent frame information to the problem receiving end, solving the video quality reduction and network pressure problems caused by Infinite GOP encoding, and improving the smoothness of video playback and network efficiency.

CN119316628BActive Publication Date: 2025-07-04BEIJING HONGYUN RONGTONG TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411128512.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-07-04
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

In RTC communication system, Infinite GOP encoding technology leads to sparse keyframes, video quality declines and playback interruptions when the network is unstable. The existing solution increases the overall code stream data volume, causing network transmission pressure.

Method used

After receiving the request, the server determines the target video frame to generate keyframes, and generates connection forward prediction frames and fusion compensation information based on the keyframes and adjacent frames to generate synchronous packet code streams, which are only sent to the receiving end where problems occur, to avoid the increase in the overall code stream data volume.

Benefits of technology

It effectively reduces network transmission pressure, improves the decoding quality and video playback fluency of the receiver, and reduces the waiting time of the receiver.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316628B_ABST
    Figure CN119316628B_ABST
Patent Text Reader

Abstract

The present invention provides a video processing method and related products. The method is executed by a server, and the method includes: receiving a first request for requesting to regenerate a key frame; determining a first target video frame according to the first request, where the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request; generating a key frame according to the first target video frame; generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame; generating fusion compensation information according to the connection forward prediction frame and a third video frame adjacent to the second video frame; generating a synchronous packet bitstream according to the key frame, the connection forward prediction frame, and the fusion compensation information; and at least directionally sending the synchronous packet bitstream to a receiving end corresponding to the first request, so that the receiving end can decode the first encoded bitstream synchronously again. Embodiments of the present invention can avoid an increase in the overall bitstream data volume of the server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video processing, and particularly relates to a video processing method and related products. Background Art

[0002] The RTC (Real-Time Communication) communication system is a real-time communication technology that allows people to communicate instantly at different geographical locations. In the RTC communication system, video encoding is a key link because it directly affects the quality of the video and the load of the network.

[0003] To reduce traffic, the RTC communication system can use the Infinite Group of Pictures (Infinite GOP) encoding technology for video encoding. Traditionally, the length of the GOP is fixed, but in some advanced encoding technologies, to improve the encoding efficiency and video quality, the length of the GOP can be made infinitely long, so it is called Infinite GOP, and it is achieved by setting the "--keyint infinite" parameter.

[0004] For example, in many video conferences, live applications, or remote teaching, this encoding technology is used. While ensuring the continuity and clarity of the video, it can minimize the consumption of network bandwidth as much as possible. However, since there are very few key frames (i.e., I frames) in Infinite GOP, when an error occurs in one frame during transmission, it may affect multiple subsequent frames until the next key frame appears to be repaired, which may lead to a decline in video quality and playback interruption, and may also cause problems such as out-of-sync between audio and video.

[0005] To solve the above problems, there is an urgent need to propose a new video processing method. Summary of the Invention

[0006] To overcome the above-mentioned defects in the prior art, the present invention provides a video processing method, which is executed by a server. The method includes: receiving a first request for requesting to regenerate a key frame; determining a first target video frame according to the first request, where the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request; generating a key frame according to the first target video frame; generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame; generating fusion compensation information according to the connection forward prediction frame and a third video frame adjacent to the second video frame; generating a synchronization data packet according to the key frame, the connection forward prediction frame and the fusion compensation information; and at least directing the synchronization data packet to a receiving end corresponding to the first request, so that the receiving end decodes the first encoded bitstream synchronously again.

[0007] The present invention also provides a video processing method, which is executed by a receiving end. The method includes: sending a first request for requesting to regenerate a key frame; receiving a synchronization packet bitstream and a first encoded bitstream forwarded by the server, where the synchronization packet bitstream is generated according to the key frame, the connection forward prediction frame and the fusion compensation information and obtained through encoding processing, the key frame is generated according to a first target video frame, the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request, the connection forward prediction frame is generated according to the key frame and a second video frame adjacent to the first video frame, and the fusion compensation information is generated according to the connection forward prediction frame and a third video frame adjacent to the second video frame; and decoding according to the synchronization packet bitstream and the first encoded bitstream to obtain an original video frame matching the playback format of the receiving end.

[0008] The present invention also provides a video processing device, which is configured in a server. The device includes: a request receiving unit for receiving a first request for requesting to regenerate a key frame; a target frame determining unit for determining a first target video frame according to the first request, where the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request; a key frame generating unit for generating a key frame according to the first target video frame; a prediction frame generating unit for generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame; a fusion information generating unit for generating fusion compensation information according to the connection forward prediction frame and a third video frame adjacent to the second video frame; a synchronization packet bitstream generating unit for generating a synchronization packet bitstream according to the key frame, the connection forward prediction frame and the fusion compensation information; and a bitstream forwarding unit for at least directing the synchronization packet bitstream to a receiving end corresponding to the first request, so that the receiving end decodes the first encoded bitstream synchronously again.

[0009] The present invention also provides a video processing device, which is configured as a client. The device includes: a request sending unit, configured to send a first request to a server, where the first request is used to request the regeneration of a key frame; a bitstream receiving unit, configured to receive a synchronization packet bitstream and a first encoded bitstream forwarded by the server. The synchronization packet bitstream is generated based on a key frame, a connected forward prediction frame, and fusion compensation information, and is obtained through encoding processing. The key frame is generated based on a first target video frame, and the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request. The connected forward prediction frame is generated based on the key frame and a second video frame adjacent to the first video frame. The fusion compensation information is generated based on the connected forward prediction frame and a third video frame adjacent to the second video frame; a bitstream decoding unit, configured to decode according to the synchronization packet bitstream and the first encoded bitstream to obtain an original video frame matching the playback format of the receiving end.

[0010] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to perform the video processing method described above.

[0011] The present invention also provides a computer device, which includes a memory and a processor. Among them: the memory stores a computer program, and when the processor executes the computer program, it implements the video processing method described above.

[0012] The video processing method and related products provided by the present invention are executed by the server. The method first receives a first request, which is used to request the regeneration of a key frame; then determines a first target video frame according to the first request, and the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request; generates a key frame according to the first target video frame; generates a connected forward prediction frame according to the key frame and a second video frame adjacent to the first video frame; generates fusion compensation information according to the connected forward prediction frame and a third video frame adjacent to the second video frame; generates a synchronization data packet according to the key frame, the connected forward prediction frame, and the fusion compensation information; and at least directionally sends the synchronization data packet to a receiving end corresponding to the first request, so that the receiving end can decode the first encoded bitstream synchronously again. The embodiments provided by the present invention avoid sending unnecessary IDR frames to all receiving ends by generating synchronization data packets, thereby reducing the overall bitstream data volume, reducing the pressure on network transmission, and further improving the quality of the decoded images at the receiving end through the fusion compensation information. Description of the Drawings

[0013] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting specific embodiments with reference to the following drawings:

[0014] Figure 1 Shows a schematic diagram of the application scenario of the video processing method provided by an embodiment of the present invention;

[0015] Figure 2 Shows a schematic flowchart of the video processing method 200 provided by an embodiment of the present invention;

[0016] Figure 3 Shows a schematic flowchart of the video processing method 300 provided by an embodiment of the present invention;

[0017] Figure 4 Shows a schematic flowchart of the video processing method 400 provided by an embodiment of the present invention;

[0018] Figure 5 Shows a schematic diagram of the first encoded bitstream and the synchronization packet bitstream provided by an embodiment of the present invention;

[0019] Figure 6 Shows a schematic framework diagram of the video processing method provided by an embodiment of the present invention;

[0020] Figure 7 Shows a schematic structural diagram of the video processing apparatus 700 provided by an embodiment of the present invention;

[0021] Figure 8 Shows a schematic structural diagram of the video processing apparatus 800 provided by an embodiment of the present invention;

[0022] Figure 9 Shows a schematic structural diagram of the video processing system 900 provided by an embodiment of the present invention;

[0023] Figure 10 Shows a schematic structural diagram of the computer device 1000 provided by an embodiment of the present invention.

[0024] Identical or similar reference numerals in the drawings represent identical or similar components. Detailed Description of the Invention

[0025] In order to better understand and explain the present invention, the present invention will be further described in detail below with reference to the accompanying drawings. The present invention is not limited solely to these specific embodiments. On the contrary, modifications or equivalent replacements made to the present invention shall be covered within the scope of the claims of the present invention.

[0026] It should be noted that numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present invention can be implemented without these specific details. In the following multiple detailed descriptions, well-known principles, structures, and components are not described in detail in order to highlight the gist of the present invention.

[0027] As Figure 1 shown, the RTC service platform (Real-Time Communication Service Platform) is a platform that provides real-time audio and video communication services. It utilizes advanced real-time audio and video processing and transmission technologies, combined with powerful cloud computing capabilities, to provide users with stable and high-quality real-time audio and video services. Through the RTC service platform, users can quickly build real-time audio and video applications across multiple platforms and achieve cross-platform real-time audio and video communication.

[0028] The RTC service platform generally includes three parts: a streaming end 102, a server end 101, and a receiving end 103. The server end 101 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The server end 101 can also provide signaling services for media negotiation and channel establishment between the two communication parties.

[0029] The streaming end 102 is used to push media streams such as audio, video, or data to the server end 101 in real time. The streaming end 102 can be a web application, a mobile application, or a dedicated hardware device, etc.

[0030] The receiving ends 103, 104, 105, 106 can also be referred to as receiving clients or receiving ends, which are used to obtain media streams from the server end 101 and play or process them locally. The receiving end 103 can be a web application, a mobile application, or a dedicated hardware device, etc. The receiving ends 103, 104, 105, 106 can be installed on electronic devices, which can include but are not limited to mobile devices such as smartphones and tablets, and computer devices such as desktop computers. Among them, the receiving end 105 is the receiving end that discovers packet loss, and the receiving end 106 is the receiving end that newly joins the call.

[0031] A network connection can be established between the streaming end 102 and the server end 101, and also between the server end 101 and the pulling end 103. The network is used to support the connection between the streaming end 102 or the pulling end 103 and the server end 101 under any network conditions. It can include, but is not limited to, devices such as access proxies and intelligent gateways. Optionally, the above network can include a wireless network or a wired network, using standard communication technologies and / or protocols through the wireless network or the wired network. The network is usually the Internet, but can also be any network, including, but not limited to, any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network.

[0032] Throughout the process, the streaming end 102 and the pulling end 103 can interact through a browser or an application, while the server end 101 is responsible for processing and forwarding the media stream, thus achieving high-quality real-time communication.

[0033] In live streaming or real-time video communication, to ensure a smooth playback experience, Infinite GOP can be adopted to avoid buffering and latency caused by frequent insertion of I-frames. In this mode, the encoder will try not to insert key frames unless it automatically detects a scene change that requires the addition of key frames. By setting fewer key frames, the compression ratio can be maximized. At the same time, by reducing the number of key frames, the Infinite GOP mode also helps to improve video quality and encoding efficiency.

[0034] However, in live streaming or real-time video communication, users may encounter video stuttering caused by unstable network. Due to the unstable network, the receiving end detects consecutive lost frames and will send a request to the sending end to resynchronize the video. In related technologies, after the sending end receives the Picture Loss Indicator (PLI) request or the Full Instra Request (FIR request) information sent by the receiving end, in the main bitstream, a P-frame is replaced with an Instantaneous Decoder Refresh Frame (IDR) frame. Subsequent P-frames depend on this IDR frame, and then the main bitstream with the temporarily inserted IDR frame and subsequent P-frames is sent to all receiving ends. After such an operation, the receiving end with problems can resynchronize and use the same video bitstream. However, the problem caused by this solution is that for one receiving end with problems, all receiving ends need to receive an unnecessary IDR frame, which increases the overall bitstream data volume, puts pressure on network transmission, and may cause potential network transmission problems such as congestion and latency.

[0035] Therefore, the present invention expects to propose a video processing method to solve the above problems and improve the user experience.

[0036] Please refer to Figure 2 , Figure 2 which shows a schematic flowchart of a video processing method 200 provided by an embodiment of the present invention. This method can be implemented by a video processing device, which is configured on the server side. As Figure 2 shown, the method includes:

[0037] Step S201, receiving a first request for requesting to regenerate a key frame.

[0038] In the above step, the first request is used to request the server to regenerate a key frame. This key frame is an instantaneous decoding refresh frame, that is, an IDR frame. The IDR frame is a special I-frame. It not only has the characteristics of an I-frame but also is used to indicate the encoder and decoder to clear the decoded image cache. This means that all frames after the IDR frame cannot refer to the frames before the IDR frame, making the front and back video sequences completely independent.

[0039] The first request can be a full instra request or a picture loss indicator request.

[0040] A full internal frame request is a request mechanism used in real-time video transmission. It mainly requests the sender to regenerate and send a complete key frame (I-frame) so that the receiver can correctly decode the video stream. This mechanism is particularly important in video conferencing and video streaming transmission, especially when the network environment is unstable or a new participant joins the conference. In practical applications, such as video conferencing or live broadcast scenarios, when a new participant joins, they may only receive P-frames or B-frames and cannot perform correct decoding due to the lack of necessary I-frames as references. In this case, the new participant will send a FIR request asking other participants to resend a data packet containing an I-frame to ensure that they can smoothly decode and watch the video.

[0041] An image loss indication request is a mechanism used in real-time video transmission to indicate the frame loss situation detected by the receiver and request the sender to take measures to correct it. This mechanism is very important in video conferencing and online streaming transmission, especially in an unstable network environment or when packet loss occurs. When this happens, the receiver will send a PLI message to the sender, and the sender can respond by retransmitting these lost packets or generating a new key frame (I-frame).

[0042] In practical applications, such as video conferencing or live broadcast scenarios, when the network environment is unstable and sudden packet loss occurs, the receiver will send a PLI request to inform the sender which frames are affected. After receiving the PLI message, the sender can choose to retransmit the lost data packets or generate a brand-new key frame to send to the receiver to ensure that the video can be smoothly decoded and played.

[0043] The first request can be received by the pushing end from the pulling end and then forwarded by the pushing end to the server; the first request can also be directly received by the server from the pulling end.

[0044] If the receiving end with problems directly sends the first request to the sending end, the sending end can use its own encoder to generate the first video frame, then decode and reconstruct the first video frame, and then generate an IDR frame using the decoded and reconstructed result, and then the sending end sends the IDR frame to the server.

[0045] If the receiving end with problems directly sends the first request to the server, since the server (i.e., the media forwarding server) is closer to the receiving end, after receiving the first request, the server will respond to the first request faster.

[0046] Step S202, determine the first target video frame according to the first request. The first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request.

[0047] In the above steps, the server receives the real-time raw stream from the streaming end. Before transcoding or encoding, the server usually needs to decode the real-time raw stream. For example, use FFmpeg or other similar video processing libraries to parse and obtain the original video format. Then, perform transcoding or encoding processing on the decoded result so that the transcoding or encoding processing result adapts to specific playback devices or network conditions.

[0048] For example, re-encode to generate a new stream that meets the requirements of the target format. Use H.264, H.265 (HEVC), AVC or other coding standards to generate a new stream. For the sake of distinction, the new stream obtained by the server through re-encoding the real-time raw stream received from the streaming end is called the main stream or the first encoded stream here.

[0049] In some embodiments, the server forwards the first encoded stream to the corresponding pulling end. When the pulling end detects packet loss, it sends a first request to the streaming end or the server. The server receives and parses the first request, and then, according to the first request, determines the first video frame in the first encoded stream that has not been forwarded to the pulling end after packet loss is detected. As Figure 6 shown, the server forwards the first encoded stream {P1, P2, ……, P8} to the pulling end. When the pulling end finds problems such as one frame or consecutive multiple frames being lost or decoding errors in the received first encoded stream, assuming that the pulling end directly sends a PLI request to the server as the first request, the server responds to the first request and determines the video frame adjacent to the position of the lost consecutive multiple frames in the first encoded stream that has not been forwarded to the pulling end as the first video frame. For example, if P2 and P3 frames are continuously lost, then determine P4 frame as the first video frame. That is, determine the first video frame ranked first from the first encoded stream to be forwarded. Assuming the first encoded stream to be forwarded is {P4, P5, ……, P8}, the first video frame ranked first is P4 frame.

[0050] The first video frame can also be the currently latest completed encoded video frame that is waiting to be transmitted. For example, the server forwards the first encoded stream {P1, P2, ……, P8} to the pulling end. When the pulling end finds a lost frame, it immediately sends a first request to the server. The first request carries information and timestamps of the lost video frame. When the server receives the first request, it determines the moment of receiving the first request and determines the latest completed encoded video frame after this moment as the first video frame. Assuming the latest completed encoded video frame is P6, determine P6 as the first video frame.

[0051] Taking the most recently encoded video frame as the first video frame is beneficial for the receiving end to better solve the problem of video playback lag at the receiving end. After taking the most recently encoded video frame as the first video frame, the reconstructed frame of the most recently encoded video frame can be cached so that subsequent decoding results all depend on the reconstructed frame of the most recently encoded video frame.

[0052] After the reconstruction of the most recently encoded video frame, the previously cached reconstructed frame can be discarded.

[0053] Step S203, generate a key frame according to the first target video frame.

[0054] In the above steps, the server generates a key frame according to the first target video frame, that is, generates an IDR frame. Generating a key frame according to the first target video frame may include reconstructing the first target video frame to obtain first video frame data; re-encoding the first video frame data in an intra-coded manner to obtain a key frame.

[0055] Optionally, after reconstructing the first target video frame to obtain first video frame data, an IDR frame can be generated by notifying the encoder by setting the frame type to IDR frame. For example, it can be achieved by checking the properties of the encoder or directly setting the frame type. Then, use a specific function such as x264_reference_reset to clear the reference frame queue of the decoder to ensure that the decoder starts re-decoding from the IDR frame.

[0056] In some embodiments, at least a user identifier and frame-drop information are carried in the first request. The user identifier can be a username or any other identification information that can uniquely identify the user, such as a hardware identifier, etc.

[0057] In the embodiments of the present application, generating an IDR frame according to the first video frame can effectively refresh the decoder state and provide a reliable random access point, thereby improving the efficiency and quality of video transmission and playback.

[0058] As Figure 6 shown, decode the code stream corresponding to the P4 frame in the first encoded bitstream to obtain the P4 frame, and then input the P4 frame into the first reconstruction unit to obtain first video frame data. Then input the first video frame data into the IDR encoding unit, and the IDR encoding unit is used to re-encode the first video frame data in an intra-coded manner to obtain a generated IDR frame. The key frame generator may at least include a first reconstruction unit and an IDR encoding unit.

[0059] The encoding process of an IDR frame involves steps such as initialization, frame type selection, macroblock segmentation, prediction mode selection, transformation and quantization, entropy encoding, generating the IDR frame, and output. The IDR encoding unit is used to initialize and set parameters including resolution, bit rate, frame rate, etc. at the start of a video coding sequence. The IDR encoding unit is used to divide a video frame into macroblocks, which are the basic units of encoding. Each macroblock can further be divided into smaller blocks, for example, in the H.264 coding standard, usually 4x4 or 8x8 blocks. For an IDR frame, all macroblocks use the intra-frame prediction mode, that is, they do not rely on information from other frames. The IDR encoding unit is used to determine the optimal intra-frame prediction mode to minimize the prediction error and perform a transformation on the prediction residual, usually using the discrete cosine transform (DCT) or its variants, such as integer transform. The transformed coefficients are quantized to reduce the amount of data. The quantized coefficients are entropy encoded to further compress the data. The IDR encoding unit is also used to encode auxiliary information such as the intra-frame prediction mode and motion vectors, and pack the encoded data into an IDR frame, which can contain intra-frame encoded image data as well as relevant header information and auxiliary information.

[0060] Step S204: Generate a connected forward prediction frame based on the key frame and the second video frame adjacent to the first video frame.

[0061] In the above steps, the connected forward prediction frame is a video frame obtained by re-encoding the second video frame in the first coded bitstream. The connected forward prediction frame is used to connect the IDR frame and the third video frame in the first coded bitstream. When performing forward encoding with reference to the reconstructed data of the IDR frame, the encoding parameters used for forward encoding can inherit the encoding parameters for forward encoding the second video frame. On this basis, the encoding parameters are further adjusted to make the reconstructed data of the connected forward video frame as close as possible to the reconstructed data of the second video frame.

[0062] The second video frame adjacent to the first video frame refers to the video frame adjacent in the sorting relationship to the first video frame in the first coded bitstream. As Figure 5 shown, in the main bitstream, the P5 frame adjacent to the P4 frame, that is, the second video frame.

[0063] In some embodiments, generating a connected forward prediction frame based on the key frame and the second video frame adjacent to the first video frame includes: reconstructing the second standard video frame to obtain second video frame data; performing inter-frame predictive coding with reference to the reconstructed data of the key frame to obtain a forward prediction frame to be adjusted; reconstructing the forward prediction frame to be adjusted to obtain third video frame data; determining a first residual value between the second video frame data and the third video frame data; and adjusting the encoding parameters for inter-frame predictive coding according to the first residual value, so that when the residual value is less than a threshold, output the forward prediction frame to be adjusted as the connected forward prediction frame.

[0064] By adjusting the quantization parameters of the encoding, the residual difference between the second video frame data and the third video frame data can be made to approach a certain threshold, so that the second video frame data and the third video frame data are approximately equal.

[0065] In video encoding, inter-frame predictive coding is an efficient method for removing temporal redundancy. The pixel values of the encoded neighboring images (reference frames) are used to predict the pixel values of the current image, and the predicted residuals are encoded. For example, each pixel block in the second video frame data searches for the best matching block in the reconstructed data of the key frame. The reference block most similar to the current block can be found within a given search range through a specific search algorithm (such as global search, three-step search method, diamond search method, etc.). Then, matching criteria such as SAD (Sum of Absolute Differences), SATD (Sum of Absolute Transformed Differences), or MSE (Mean Square Error) are used to determine the best matching block to minimize the prediction residual. Finally, motion estimation with sub-pixel accuracy can improve the prediction accuracy. For example, H.264 supports motion estimation with 1 / 4 pixel accuracy, while H.265 supports higher sub-pixel accuracy, such as 1 / 8 pixel interpolation.

[0066] After the motion vectors obtained according to the motion estimation, the predicted block is obtained from the reconstructed data of the key frame, and the predicted value of the current block is constructed. Then, the predicted block is subtracted from the current block to generate a residual signal for further encoding. The motion vector of the current block can be predicted using the motion vectors of neighboring blocks within the same image. Or other similar methods are used for the prediction and encoding of motion vectors.

[0067] As Figure 6 shown, the bitstream corresponding to the P5 frame in the first encoded bitstream is decoded to obtain the P5 frame, the P5 frame is input into the fourth reconstruction unit to obtain the third video frame data, and then the third video frame data is input into the encoding parameter adjustment unit for adjustment; the IDR frame is input into the second reconstruction unit to obtain the IDR reconstructed data, the IDR reconstructed data is input into the LP encoding unit to obtain the encoded data, the encoded data is input into the third reconstruction unit to obtain the fourth video frame data, and then the fourth video frame data is input into the encoding parameter adjustment unit. According to the residual between the third video frame data and the fourth video frame data, the encoding parameters of the LP encoding unit are adjusted, and finally the linked forward prediction frame, that is, the LP frame (English full name LinkPredictive-Frame), is output, so that the residual between the LP frame and the reconstructed data of the P5 frame is minimized.

[0068] Step S205, generating fusion compensation information according to the linked forward prediction frame and the third video frame adjacent to the second video frame.

[0069] In the above steps, the fusion compensation information refers to information used to correct the decoding distortion result after switching to the first coded bitstream. The fusion compensation information is generated by the difference between the YUV data reconstructed according to the third video frame in the first coded bitstream and the YUV data reconstructed according to the connected forward prediction frame when encoding the third video frame. After decoding the third video frame in the first coded bitstream at the decoding end, the fusion compensation information can be used for compensation to obtain a better quality image result.

[0070] Step S206, generating a synchronous packet stream according to the key frame, the connected forward prediction frame and the fusion compensation information.

[0071] In the above steps, entropy coding is performed on the transformed and quantized data, such as using Huffman coding as entropy coding.

[0072] The server generates a synchronization data packet based on the IDR frame, the connected forward prediction frame and the fusion compensation information. Figure 5 As shown, the synchronous packet code stream is obtained by encoding the IDR frame, the connected forward prediction frame (LP frame) and the fusion compensation information respectively. In some embodiments, the fusion compensation information is stored in the side information (English full name Side Information, English abbreviation SI).

[0073] The coded code stream corresponding to the P6 frame to the P8 frame is the coded code stream corresponding to the P6 frame to the P8 frame in the first coded code stream to be forwarded. The server generates a synchronous packet code stream and only sends the synchronous packet code stream to the receiving end with the problem in a directional manner, thereby effectively avoiding an increase in the number of overall code streams on the server and reducing the pressure on network transmission.

[0074] Step S207: at least send the synchronous packet code stream directionally to the receiving end corresponding to the first request, so that the receiving end can decode the first encoded code stream synchronously again.

[0075] In the above steps, when the server receives the first request, it records the user identifier of the receiving end that sends the first request; after generating the synchronous packet code stream, it forwards the synchronous packet code stream directionally to the receiving end corresponding to the user identifier according to the user identifier, and forwards the first encoded code stream to all receiving ends.

[0076] The embodiment provided by the present application, by sending the synchronous packet code stream and the first encoding code stream to the receiving end where the problem occurs, the receiving end where the problem occurs can synchronously decode the first encoding code stream, thereby saving the time of the receiving end where the problem occurs waiting for the fault to be resolved, reducing the overall code stream data volume of the server, and reducing the pressure on network transmission.

[0077] A network connection is established between the server and the receiving end with problems or the newly established receiving end respectively. For video transmissions that require high reliability, such as live broadcasts or video conferences, TCP is usually selected. Before sending the encoded bitstream, the encoded bitstream will be encapsulated, which includes adding necessary header information, timestamps, sequence numbers, etc., to ensure that the client can correctly parse and play the received video data. The encapsulation process may also involve fragmenting the bitstream to adapt to the maximum transmission unit of the network and optimize the transmission efficiency. Once the encapsulation is completed, the encapsulated result is forwarded to the client. This is usually accomplished using the network protocol stack, and address addressing and routing are performed through lower-layer protocols (such as the IP protocol). During the forwarding process, the server may need to process feedback information from the network, such as congestion control and retransmission requests, to ensure the stability of data transmission.

[0078] After the data is received at the receiving end, it is first de-encapsulated to restore the original encoded bitstream. Then, the decoder at the receiving end will decode the received encoded bitstream and finally present the video content.

[0079] In the embodiment provided by the present invention, when a new receiving end joins the RTC service platform or a packet loss problem occurs at a certain receiving end, after the server receives the first request, it generates a synchronization packet bitstream and only sends the synchronization packet bitstream to the receiving end directionally, which can effectively avoid the increase in the overall bitstream data volume of the server and reduce the pressure on network transmission.

[0080] The following combines Figure 3 , and further describes the video processing method provided by the embodiments of the present application. Please refer to Figure 3 , Figure 3 shows a schematic flowchart of a video processing method 300 provided by an embodiment of the present invention. This method can be implemented by a video processing device, which is configured on the server side. As Figure 3 shown, this method includes:

[0081] Step S301, receiving a first request sent by the client, where the first request is used to request the regeneration of a key frame.

[0082] Step S302, determining a first target video frame according to the first request, where the first target video frame is the first video frame ranked first determined from the first encoded bitstream to be forwarded by the server after receiving the first request.

[0083] Step S303, generating a key frame according to the first target video frame.

[0084] Step S304, performing inter-frame predictive coding using the mode data of the second target video frame to obtain an encoding result.

[0085] Step S305, reconstructing the encoding result to obtain fourth video frame data.

[0086] Step S306, performing inter-frame prediction coding on the key frame to obtain a forward prediction frame to be adjusted; reconstructing the forward prediction frame to be adjusted to obtain third video frame data.

[0087] Step S307: determining a second residual value between the third video frame data and the fourth video frame data.

[0088] Step S308, adjusting the coding parameters for inter-frame prediction coding according to the second residual value, so that when the residual value is less than the threshold, the forward prediction frame to be adjusted is output as the connected forward prediction frame.

[0089] The mode data of a video frame refers to the data information used to describe the internal structure and properties of a video frame during the video encoding process. The mode data of a video frame includes, but is not limited to, partition, motion and other mode data. For example, partition mode data refers to the way in which a video frame is divided into different regions (or blocks) during video encoding. The purpose of partitioning is to process the video content more accurately so that the most appropriate encoding method can be applied to different regions. Motion mode data is mainly used to describe the motion trajectory of an object in a video frame.

[0090] Through accurate partitioning and motion estimation, as well as appropriate coding mode selection, the compression rate of video data can be greatly improved, while ensuring the smoothness and quality of video playback. In practical applications, the characteristics and advantages of various mode data should be comprehensively considered, and appropriate coding strategies and technologies should be selected in a targeted manner to achieve the best coding effect.

[0091] The following takes the H.264 coding standard as an example to illustrate the process of using the mode data of the second video frame to perform inter-frame prediction coding to obtain the coding result. First, select one or more encoded frames as reference frames for the current frame. For example, select the IDR frame as the reference frame and the second video frame as the current frame. For each macroblock in the current frame, a motion vector is calculated by comparing its difference with the macroblock at the corresponding position in the reference frame. The motion vector is used to describe the position offset of the macroblock in the current frame relative to the macroblock at the corresponding position in the IDR frame. According to the motion vector, the corresponding pixel value is obtained from the IDR frame and filled into the corresponding position of the current frame to obtain the predicted frame. The predicted frame is subtracted from the current frame to obtain the residual frame. The residual frame contains the difference information between the current frame and the predicted frame. Finally, the residual frame is subjected to discrete cosine transform and quantization operations to obtain the coding result.

[0092] Step S309: determine to connect the forward prediction frame as a reference frame, reconstruct the third video frame adjacent to the second video frame, and obtain fifth video frame data.

[0093] Step S310: Determine the second video frame as the reference frame, and reconstruct the third video frame adjacent to the second video frame to obtain the sixth video frame data.

[0094] Step S311: Obtain the fusion compensation information according to the third residual value between the fifth video frame data and the sixth video frame data.

[0095] In the above steps, the fusion compensation information is used for compensation when the second encoded bitstream is decoded at the client.

[0096] In the above steps, the second video frame is a video frame in the first encoded bitstream, and the connected forward prediction frame is the reconstructed decoded data of the reference IDR frame and the video frame obtained by reconstructing and decoding the second video frame. Using the decoded IDR frame as the reference frame to perform prediction, compensation, and reconstruction on the second video frame is to utilize the temporal redundancy information in the video frame sequence to improve the compression efficiency and maintain the video quality.

[0097] After determining the reference frame, use the prediction frame obtained from the reference frame to calculate the residual between the current frame and the prediction frame. The residual reflects the error after prediction.

[0098] To ensure that subsequent frames can be correctly decoded, the third video frame is reconstructed with reference to the connected forward prediction frame and the second video frame respectively to obtain the fifth video frame data and the sixth video frame data corresponding to the third video frame. Then, according to the third residual value between the fifth video frame data and the sixth video frame data, the fusion compensation information is obtained, and the fusion compensation information is combined with the motion vector and mode decision information of the third video frame through entropy coding to form the encoded bitstream of the P6 frame.

[0099] In the embodiments of the present application, by reconstructing new frames with reference to the decoded video frames, the amount of data required for encoding is effectively reduced, while ensuring the continuity and smoothness of video playback.

[0100] Step S312: Generate a synchronization packet bitstream according to the key frame, the connected forward prediction frame, and the fusion compensation information.

[0101] Step S313: Directly send at least the synchronization packet bitstream to the receiving end corresponding to the first request, so that the receiving end can decode the first encoded bitstream synchronously again.

[0102] In the embodiments provided by the present invention, after the server receives the first request, a new IDR frame is generated, and then the connected forward prediction frame is generated by using the mode data of the second video frame and the reconstructed data of the IDR frame, which can effectively improve the encoding speed; furthermore, by generating the fusion compensation information, it helps to improve the decoding quality of the first encoded bitstream at the receiving end with problems, thereby ensuring the smoothness of video playback at the receiving end with problems.

[0103] The following describes from the receiving - end side to further describe the video processing method provided by this application. Please refer to Figure 4 , Figure 4 FIG. shows a schematic flowchart of a video processing method 400 provided by an embodiment of the present invention. This method can be implemented by a video processing device, which is configured on the receiving - end side. As Figure 4 shown, the method includes:

[0104] Step S401: Send a first request to the server. The first request is used to request the regeneration of key frames.

[0105] Step S402: Receive the synchronized packet bitstream and the first encoded bitstream forwarded by the server.

[0106] In the above steps, the synchronized packet bitstream is generated based on key frames, connected forward - predicted frames, and fusion compensation information, and is obtained through encoding processing. The key frame is generated based on a first target video frame. The first target video frame is the first video frame ranked first in the first encoded bitstream to be forwarded determined by the server after receiving the first request. The connected forward - predicted frame is generated based on the key frame and a second video frame adjacent to the first video frame. The fusion compensation information is generated based on the connected forward - predicted frame and a third video frame adjacent to the second video frame.

[0107] Step S403: Decode according to the synchronized packet bitstream and the first encoded bitstream to obtain an original video frame that matches the playback format of the receiving end.

[0108] In the above steps, the decoder at the receiving end is used to parse the bitstream corresponding to the IDR frame. For example, extract metadata such as frame - header information, sequence parameter sets, and picture parameter sets. Since the IDR frame is a completely independent frame and does not depend on any previously received frames, it does not require prediction with reference to other frames. Then, the decoder at the receiving end is also used to perform entropy - decoding operations to convert the variable - length - coded codewords into quantization coefficients. The quantization coefficients obtained after entropy decoding need to be reconstructed into pixel values through inverse transformation (such as inverse discrete cosine transform) and inverse quantization processes. The inverse transformation converts frequency - domain data back into spatial - domain data, and inverse quantization restores an approximation of the original pixel values according to the quantization parameters.

[0109] To improve the image quality, the decoder at the receiving end is also used to perform loop - filtering processing on the decoded image. This includes de - blocking effect filtering and other image - enhancement techniques, which can reduce the distortion and noise introduced during the encoding process. Finally, the decoder of the client is used to output the decoded image data to the display device to complete the decoding process.

[0110] After the decoder at the receiving end decodes an IDR frame, it continues to decode the bitstream corresponding to the connected forward prediction frame received. The decoder at the receiving end parses out the motion vectors and mode decision information, and then uses this information to reconstruct the connected forward prediction frame. Next, the decoder at the client side is also used to perform inverse quantization and inverse DCT operations on the residual data to recover the residual frame. Finally, the decoder at the receiving end adds the residual frame back to the prediction frame to obtain the reconstructed frame corresponding to the connected forward prediction frame.

[0111] In some embodiments, the fusion compensation information is obtained based on the third residual value between the fifth video frame data and the sixth video frame data. The fifth video frame data is obtained by reconstructing the third video frame adjacent to the second video frame with reference to the connected forward prediction frame, and the sixth video frame data is obtained by reconstructing the third video frame adjacent to the second video frame with reference to the second video frame. Decoding according to the synchronization packet bitstream and the first encoded bitstream includes: decoding the synchronization packet bitstream to obtain key frames, connected forward prediction frames, and fusion compensation information; referring to the connected forward prediction frames, decoding the first encoded bitstream to obtain the reconstructed frame corresponding to the third video frame adjacent to the second video frame; using the fusion compensation information to compensate the reconstructed frame so that the compensated result is approximated to the reconstructed frame of the third video frame.

[0112] In the embodiments provided by the present invention, when the receiving end discovers packet loss or newly joins the RTC service platform, it actively sends a first request, expecting the server to respond quickly, reducing the waiting time of the problematic receiving end or the new receiving end. After the receiving end receives the synchronization packet bitstream, it can use the synchronization packet bitstream to connect the first encoded bitstream, reducing the time required for the problematic receiving end or the new receiving end to wait for decoding synchronization, and greatly improving the user experience.

[0113] Please refer to Figure 7 , Figure 7 FIG. shows a schematic structural diagram of a video processing apparatus 700 provided by an embodiment of the present invention. The apparatus is configured on the server side, and the apparatus includes:

[0114] A request receiving unit 701, configured to receive a first request sent by a client, where the first request is used to request to regenerate a key frame;

[0115] A target frame determining unit 702, configured to determine a first target video frame according to the first request, where the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request;

[0116] A key frame generating unit 703, configured to generate a key frame according to the first target video frame;

[0117] A prediction frame generation unit 704, configured to generate a connection forward prediction frame according to a key frame and a second video frame adjacent to the first video frame;

[0118] A fusion information generation unit 705, configured to generate fusion compensation information according to the connection forward prediction frame and a third video frame adjacent to the second video frame;

[0119] A synchronization packet bitstream generation unit 706, configured to generate a synchronization packet bitstream according to the key frame, the connection forward prediction frame, and the fusion compensation information;

[0120] A bitstream forwarding unit 707, configured to direct at least the synchronization packet bitstream to a receiving end corresponding to the first request, so that the receiving end decodes the first encoded bitstream synchronously again.

[0121] Optionally, the key frame generation unit 703 is further configured to reconstruct the first target video frame to obtain first video frame data; re-encode the first video frame data in an intra-coding manner to obtain a key frame.

[0122] Optionally, the prediction frame generation unit 704 is further configured to reconstruct the second target video frame to obtain second video frame data; perform inter-frame prediction coding with reference to the reconstruction data of the key frame to obtain a to-be-adjusted forward prediction frame; reconstruct the to-be-adjusted forward prediction frame to obtain third video frame data; determine a first residual value between the second video frame data and the third video frame data; adjust encoding parameters for inter-frame prediction coding according to the first residual value, so that when the first residual value is less than a threshold, output the to-be-adjusted forward prediction frame as the connection forward prediction frame.

[0123] Optionally, the prediction frame generation unit 704 is further configured to perform inter-frame prediction coding using the mode data of the second video frame to obtain an encoding result; reconstruct the encoding result to obtain fourth video frame data; perform inter-frame prediction coding with reference to the reconstruction data of the key frame to obtain a to-be-adjusted forward prediction frame; reconstruct the to-be-adjusted forward prediction frame to obtain third video frame data; determine a second residual value between the third video frame data and the fourth video frame data; adjust encoding parameters for inter-frame prediction coding according to the second residual value, so that when the second residual value is less than a threshold, output the to-be-adjusted forward prediction frame as the connection forward prediction frame.

[0124] Optionally, the fusion information generation unit 705 further includes: a fifth reconstruction subunit, configured to determine the connection forward prediction frame as a reference frame and reconstruct the third video frame adjacent to the second video frame to obtain fifth video frame data; a sixth reconstruction subunit, configured to determine the second video frame as a reference frame and reconstruct the third video frame adjacent to the second video frame to obtain sixth video frame data; a fusion compensation subunit, configured to obtain fusion compensation information according to a third residual value between the fifth video frame data and the sixth video frame data.

[0125] In the embodiment provided by the present application, the video processing device is configured on the server side. After the server receives the first request, it first generates a new IDR frame, and then generates a connection forward prediction frame by using the mode data of the second video frame and the reconstruction data of the IDR frame. This can effectively improve the encoding speed, thereby saving the waiting time of the receiving end with problems.

[0126] Furthermore, by generating fusion compensation information, it helps to improve the quality of decoding the first encoded bitstream at the receiving end with problems, thereby ensuring the smoothness of video playback at the receiving end with problems.

[0127] Please refer to Figure 8 , Figure 8 which shows a schematic structural diagram of a video processing device 800 provided by an embodiment of the present invention. The device is configured at the receiving end, and the device includes:

[0128] A request sending unit 801, configured to send a first request, where the first request is used to request the regeneration of a key frame;

[0129] A bitstream receiving unit 802, configured to receive a synchronization packet bitstream and a first encoded bitstream forwarded by the server. The synchronization packet bitstream is generated according to a key frame, a connection forward prediction frame, and fusion compensation information and obtained through encoding processing. The key frame is generated according to a first target video frame, and the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request. The connection forward prediction frame is generated according to the key frame and a second video frame adjacent to the first video frame, and the fusion compensation information is generated according to the connection forward prediction frame and a third video frame adjacent to the second video frame;

[0130] A bitstream decoding unit 803, configured to perform decoding according to the synchronization packet bitstream and the first encoded bitstream to obtain an original video frame matching the playback format of the receiving end.

[0131] Optionally, the fusion compensation information is obtained according to a third residual value between fifth video frame data and sixth video frame data. The fifth video frame data is obtained by reconstructing a third video frame adjacent to the second video frame with reference to the connection forward prediction frame, and the sixth video frame data is obtained by reconstructing a third video frame adjacent to the second video frame with reference to the second video frame. The bitstream decoding unit 803 is further configured to parse the synchronization packet bitstream to obtain the key frame, the connection forward prediction frame, and the fusion compensation information; decode the first encoded bitstream with reference to the connection forward prediction frame to obtain a reconstructed frame corresponding to the third video frame adjacent to the second video frame; and compensate the reconstructed frame by using the fusion compensation information so that the compensated result is approximated to the reconstructed frame of the third video frame.

[0132] In the embodiment provided by the present invention, the video processing device is configured on the receiving end side. When the receiving end finds packet loss or newly joins the RTC service platform, it actively sends a first request, expecting the service end to respond quickly, which can reduce the waiting time of the receiving end with problems or new receiving ends. After receiving the synchronous packet stream, the receiving end uses the synchronous packet stream to connect and decode the first encoding stream, which reduces the time required for the receiving end with problems or new receiving ends to wait for decoding synchronization, greatly improving the user experience.

[0133] Please refer to Figure 9 , Figure 9 FIG. 9 is a schematic diagram showing the structure of a video processing system 900 provided by an embodiment of the present invention. The system includes a server 901 and at least one receiving end 902. The server includes Figure 7 The video processing device shown in FIG. 1 includes a receiving end 902 including: Figure 8 The video processing device shown.

[0134] Please refer to Figure 10 , Figure 10 FIG. 1 shows a schematic diagram of the structure of a computer device 1000 provided in an embodiment of the present invention. The computer device 1000 may be used to forward a coded stream to a client. Figure 10 As shown, the computer device 1000 includes at least a memory 1001 and a processor 1002. For example, the computer device may also include Figure 10 The embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned Figures 2 - 4 Describe the method.

[0135] In particular, according to the embodiment provided by the present invention, the above reference flow chart Figures 2 - 4 The described process can be implemented as a computer software program. For example, the embodiments provided by the present invention include a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program contains program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the system of the present invention are executed.

[0136] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the methods and computer program products described in various embodiments provided by the present invention. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0138] The units or modules involved in the embodiments provided by the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a request receiving unit, a target frame determining unit, a key frame generating unit, a predicted frame generating unit, a fusion information generating unit, and a bitstream forwarding unit. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases. For example, the request receiving unit can also be described as "a unit for receiving a first request".

[0139] As another aspect, an embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium can be included in the computer device described in the above embodiments; it can also exist alone without being assembled into the computer device. The above computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors, they are used to perform the video processing method described in the present invention.

[0140] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present invention.

Claims

1. A video processing method, characterized in that, This method is executed by the server, and the method includes: Receiving a first request for requesting the regeneration of a key frame; Determining a first target video frame according to the first request, where the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request; Generating a key frame according to the first target video frame; Generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame; Generating fusion compensation information according to the connection forward prediction frame and a third video frame adjacent to the second video frame, where the fusion compensation information is generated based on the difference between the YUV data reconstructed according to the third video frame in the first encoded bitstream and the YUV data reconstructed according to the connection forward prediction frame when encoding the third video frame; Generating a synchronization packet bitstream according to the key frame, the connection forward prediction frame, and the fusion compensation information; Directing at least the synchronization packet bitstream to a receiving end corresponding to the first request, so that the receiving end decodes the first encoded bitstream synchronously again; Among them, generating a key frame according to the first target video frame includes: reconstructing the first target video frame to obtain first video frame data; re-encoding the first video frame data in an intra-coded manner to obtain a key frame; Generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame includes: reconstructing the second target video frame to obtain second video frame data; performing inter-frame prediction coding with reference to the reconstructed data of the key frame to obtain a forward prediction frame to be adjusted; reconstructing the forward prediction frame to be adjusted to obtain third video frame data; determining a first residual value between the second video frame data and the third video frame data; adjusting the coding parameters for inter-frame prediction coding according to the first residual value, so that when the residual value is less than a threshold, outputting the forward prediction frame to be adjusted as the connection forward prediction frame.

2. The method according to claim 1, wherein Generating a key frame according to the first target video frame includes: Reconstructing the first target video frame to obtain first video frame data; Re-encoding the first video frame data in an intra-coded manner to obtain the key frame.

3. The method according to claim 1, wherein Generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame includes: Reconstructing the second video frame to obtain second video frame data; Performing inter-frame prediction coding with reference to the reconstructed data of the key frame to obtain a forward prediction frame to be adjusted; Reconstructing the forward prediction frame to be adjusted to obtain third video frame data; Determining a first residual value between the second video frame data and the third video frame data; Adjusting the coding parameters for inter-frame prediction coding according to the first residual value, so that when the first residual value is less than a threshold, outputting the forward prediction frame to be adjusted as the connection forward prediction frame.

4. The method according to claim 1, wherein Generating a connection forward prediction frame according to the key frame and a second video frame adjacent to the first video frame includes: Performing inter-frame prediction coding using the mode data of the second video frame to obtain a coding result; Reconstruct the encoded result to obtain the fourth video frame data; Perform inter-frame predictive coding with reference to the reconstruction data of the key frame to obtain a forward prediction frame to be adjusted; Reconstruct the forward prediction frame to be adjusted to obtain the third video frame data; Determine a second residual value between the third video frame data and the fourth video frame data; Adjust the coding parameters for inter-frame predictive coding according to the second residual value, such that when the second residual value is less than a threshold, output the forward prediction frame to be adjusted as the connected forward prediction frame.

5. The method according to claim 1, characterized in that, The method further includes: When receiving the first request, record the user identifier corresponding to the receiving end that sends the first request; The at least directing the synchronization packet bitstream to the receiving end corresponding to the first request includes: Directly forwarding the synchronization packet bitstream to the receiving end corresponding to the user identifier; And continue to forward the first encoded bitstream to all receiving ends that have established a communication connection with the server.

6. The method according to any one of claims 1-5, generating fusion compensation information according to the connected forward prediction frame and the third video frame adjacent to the second video frame, includes: Determine the connected forward prediction frame as a reference frame, and reconstruct the third video frame adjacent to the second video frame to obtain the fifth video frame data; Determine the second video frame as a reference frame, and reconstruct the third video frame adjacent to the second video frame to obtain the sixth video frame data; Obtain the fusion compensation information according to a third residual value between the fifth video frame data and the sixth video frame data.

7. A video processing method, characterized in that, The method is executed by a receiving end, and the method includes: Send a first request, where the first request is used to request the regeneration of a key frame; Receive a synchronization packet bitstream and a first encoded bitstream forwarded by the server, where the synchronization packet bitstream is generated according to the key frame, the connected forward prediction frame, and the fusion compensation information, and is obtained after encoding processing, the key frame is generated according to a first target video frame, the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request, the connected forward prediction frame is generated according to the key frame and the second video frame adjacent to the first video frame, and the fusion compensation information is generated according to the connected forward prediction frame and the third video frame adjacent to the second video frame; wherein the fusion compensation information is generated based on the difference between the YUV data obtained by reconstructing the third video frame in the first encoded bitstream and the YUV data obtained by reconstructing the third video frame according to the connected forward prediction frame when encoding the third video frame; Decode according to the synchronization packet bitstream and the first encoded bitstream to obtain an original video frame matching the playback format of the receiving end; Wherein, the key frame is generated according to the first target video frame, including: reconstructing the first target video frame to obtain the first video frame data; re-encoding the first video frame data in an intra-coding manner to obtain the key frame; The connected forward prediction frame is generated based on the key frame and a second video frame adjacent to the first video frame, and includes: reconstructing the second standard video frame to obtain second video frame data; performing inter-frame predictive coding with reference to the reconstruction data of the key frame to obtain a forward prediction frame to be adjusted; reconstructing the forward prediction frame to be adjusted to obtain third video frame data; determining a first residual value between the second video frame data and the third video frame data; adjusting the coding parameters for inter-frame predictive coding according to the first residual value, so that when the residual value is less than a threshold, outputting the forward prediction frame to be adjusted as the connected forward prediction frame.

8. The method according to claim 7, wherein The fusion compensation information is obtained according to a third residual value between fifth video data and sixth video data. The fifth video data is obtained by reconstructing a third video frame adjacent to the second video frame with reference to the connected forward prediction frame. The sixth video data is obtained by reconstructing a third video frame adjacent to the connected forward prediction frame with reference to the second video frame. The decoding according to the synchronization packet bitstream and the first coding bitstream includes: Decoding the synchronization packet bitstream to obtain the key frame, the connected forward prediction frame, and the fusion compensation information; Decoding the first coding bitstream with reference to the connected forward prediction frame to obtain a reconstructed frame corresponding to the third video frame adjacent to the second video frame; Compensating the reconstructed frame with the fusion compensation information so that the compensated result approximates the reconstructed frame of the third video frame.

9. A video processing device, characterized in that, The device is configured in a server, and the device includes: A request receiving unit, configured to receive a first request for requesting to regenerate a key frame; A target frame determining unit, configured to determine a first target video frame according to the first request. The first target video frame is the first video frame ranked first determined by the server from the first coding bitstream to be forwarded after receiving the first request; A key frame generating unit, configured to generate a key frame according to the first target video frame; A predictive frame generating unit, configured to generate a connected forward prediction frame according to the key frame and a second video frame adjacent to the first video frame; A fusion information generating unit, configured to generate fusion compensation information according to the connected forward prediction frame and a third video frame adjacent to the second video frame. The fusion compensation information is generated based on the difference between the YUV data obtained by reconstructing the third video frame according to the third video frame in the first coding bitstream and the YUV data obtained by reconstructing the third video frame according to the connected forward prediction frame when encoding the third video frame; A synchronization packet bitstream generating unit, configured to generate a synchronization packet bitstream according to the key frame, the connected forward prediction frame, and the fusion compensation information; A bitstream forwarding unit, configured to at least direct the synchronization packet bitstream to a receiving end corresponding to the first request, so that the receiving end decodes the first coding bitstream synchronously again; Among them, generating a key frame according to the first target video frame includes: reconstructing the first target video frame to obtain first video frame data; re-encoding the first video frame data in an intra-coding manner to obtain a key frame; Generating a connected forward prediction frame based on a key frame and a second video frame adjacent to the first video frame, including: reconstructing the second standard video frame to obtain second video frame data; performing inter-frame prediction coding with reference to the reconstruction data of the key frame to obtain a forward prediction frame to be adjusted; reconstructing the forward prediction frame to be adjusted to obtain third video frame data; determining a first residual value between the second video frame data and the third video frame data; adjusting the coding parameters for inter-frame prediction coding according to the first residual value, such that when the residual value is less than a threshold, outputting the forward prediction frame to be adjusted as the connected forward prediction frame.

10. A video processing device, characterized in that, The apparatus is configured at a receiving end, and the apparatus includes: A request sending unit, configured to send a first request to a server, where the first request is used to request regeneration of a key frame; A bitstream receiving unit, configured to receive a synchronization packet bitstream and a first encoded bitstream forwarded by the server, where the synchronization packet bitstream is generated according to the key frame, the connected forward prediction frame, and fusion compensation information, and is obtained through encoding processing. The key frame is generated according to a first target video frame, where the first target video frame is the first video frame ranked first determined by the server from the first encoded bitstream to be forwarded after receiving the first request. The connected forward prediction frame is generated according to the key frame and a second video frame adjacent to the first video frame. The fusion compensation information is generated according to the connected forward prediction frame and a third video frame adjacent to the second video frame. The fusion compensation information is generated as a difference between the YUV data reconstructed according to the third video frame in the first encoded bitstream and the YUV data reconstructed according to the connected forward prediction frame when encoding the third video frame. A bitstream decoding unit, configured to decode according to the synchronization packet bitstream and the first encoded bitstream to obtain an original video frame matching the playback format of the receiving end; Wherein, the key frame is generated according to a first target video frame, including: reconstructing the first target video frame to obtain first video frame data; re-encoding the first video frame data in an intra-coding manner to obtain the key frame; The connected forward prediction frame is generated according to the key frame and a second video frame adjacent to the first video frame, including: reconstructing the second standard video frame to obtain second video frame data; performing inter-frame prediction coding with reference to the reconstruction data of the key frame to obtain a forward prediction frame to be adjusted; reconstructing the forward prediction frame to be adjusted to obtain third video frame data; determining a first residual value between the second video frame data and the third video frame data; adjusting the coding parameters for inter-frame prediction coding according to the first residual value, such that when the residual value is less than a threshold, outputting the forward prediction frame to be adjusted as the connected forward prediction frame.

11. A video processing system, characterized in that, The system includes a server and at least one receiving end. The server includes the video processing apparatus as described in claim 9, and the receiving end includes the video processing apparatus as described in claim 10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to perform the video processing method as described in any one of claims 1 to 8.

13. A computer device, the computer device comprising a memory and a processor, wherein: The memory stores a computer program, and when the processor executes the computer program, the video processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Video stream data transmission method and device, storage medium and electronic equipment

    CN114567799A

  • Video coding apparatus and method for inserting key frame adaptively

    US20050169371A1