Video error spread suppression method and device for high-definition video

By adjusting the image group structure of the video data, selecting frames with high cross-correlation as enhancement P frames, and changing their reference relationships, the problem of inability to effectively control the video code rate in the prior art is solved, and efficient transmission and high-quality reception of high-definition videos are achieved.

CN120034656APending Publication Date: 2025-05-23SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109030.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing video error spread suppression methods cannot effectively control the video code rate, resulting in wasted spectrum resources.

Method used

By adjusting the image group structure of the video data, selecting frames with high cross-correlation as enhancement P frames (EP frames), and changing their reference relationships, so that the I frame or the previous EP frame directly references the EP frame, thereby reducing the amount of inter prediction residual data and reducing the code rate.

Benefits of technology

It achieves the implementation of maintaining a low bit rate while suppressing the spread of video errors, avoiding the waste of spectrum resources, and improving the transmission efficiency and quality of high-definition videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034656A_ABST
    Figure CN120034656A_ABST
Patent Text Reader

Abstract

The invention discloses a video error spread suppression method and device for a high-definition video, and the method comprises the following steps: S1, obtaining original video information from a shooting end, and transmitting the original video information to an encoder; s2, adjusting an image group structure of video data in an encoder; s3, encoding the adjusted video data to generate a video code stream; and S4, sending the coded video code stream. The device comprises a data receiving module, an image group adjusting module, a video coding module and a data transmission module, wherein the data transmission module is used for transmitting the data coded by the video coding module. According to the method, the coding enhancement reference frame is selected by utilizing the characteristic that the inter-frame prediction residual data volume of the frames with higher cross correlation is smaller, so that the coding system has a smaller code rate, and meanwhile, the error spreading phenomenon is fully inhibited. Error spread in an image group is controlled within an acceptable range, and meanwhile, the code rate only fluctuates within a small range, so that the code rate cost is minimum while the quality of a high-definition video is improved in a transmission process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video coding and decoding, and in particular, to a method and apparatus for suppressing video error propagation for high-definition video. Background Art

[0002] Currently, multimedia transmission services have become a major part of fields such as communication and the Internet of Things, and high-definition video has become the core part of multimedia services. The transmission efficiency and reception quality of high-definition video highly depend on video coding and decoding technology. During the video coding process, important steps for processing video data include inter-frame prediction and group of pictures (GOP) structure division. Therefore, in order to make the reference structure of the GOP more efficient, it is necessary to optimize the GOP structure to contain the error propagation within the GOP during the transmission process and improve the image compression and transmission efficiency.

[0003] Currently, the main methods used to contain error propagation are as follows: shortening the GOP length and inserting I-frames more frequently in the video; changing the reference frames of P-frames at fixed intervals (referred to as enhanced frames), and these reference P-frames are directly encoded in an inter-frame coding manner with reference to the I-frame or the previous enhanced P-frame. The above two methods can indeed solve the problem of error propagation within the GOP to a certain extent.

[0004] However, the above two methods do not consider that the data volume after encoding the I-frame and the enhanced frame is much larger than that of ordinary P-frames, cannot effectively control the video bit rate, and cannot solve the problem of wasted spectrum resources. Summary of the Invention

[0005] In order to solve the problem that the existing methods for containing error propagation cannot control the video bit rate, the present invention proposes a method and apparatus for suppressing video error propagation for high-definition video, which takes advantage of the characteristic that the inter-frame prediction residual data volume of frames with high mutual correlation in various aspects is small, so as to have a smaller bit rate in the coding system while fully suppressing the error propagation phenomenon to solve the above problems.

[0006] The present application discloses a method for suppressing video error propagation for high-definition video, including the following steps: S1. Obtain the original video information from the shooting end and send it to the encoder; S2. Adjust the GOP structure of the video data in the encoder; S3. Encode the adjusted video data to generate a video bitstream; S4. Send the encoded video bitstream.

[0007] Preferably, the S1 includes the following steps: S11. Identify the original video data format; S12, if the recognized format is RGB or YUV, it is directly sent to the encoder; S13. If the recognized format is RAWData, the data is first ISP processed, converted into RGB data and then sent to the encoder.

[0008] Preferably, S2 comprises the following steps: S21, taking the image groups divided by the encoder as units, starting from the video frame that is m frames apart from the I frame, calculating the cross-correlation between each frame and the I frame; S22, selecting a frame where the local maximum value of the cross-correlation is located, and recording it as an enhanced P frame, that is, an EP frame; S23, starting from the video frame that is m frames away from the EP frame, calculate the cross-correlation between each frame and the EP frame and the cross-correlation between each frame and the I frame; S24, selecting a frame where the local maximum value of the cross-correlation with the previous EP frame is located, and recording it as an enhanced P frame, that is, an EP frame; S25, repeat S23 and S24 until the interval between the last EP frame and the last frame of the picture group is less than or equal to m; S26, adjusting the reference relationship of the EP frames so that the first EP frame directly references the I frame. If the cross-correlation between the remaining EP frames and the previous EP frame is greater than the cross-correlation between the EP frame and the I frame, the previous EP frame is directly referenced; if the cross-correlation between the remaining EP frames and the previous EP frame is less than the cross-correlation between the EP frame and the I frame, the I frame is directly referenced. S27, the frames other than the I frame and the EP frame in the image group keep the original reference relationship unchanged; S28. Repeat S21-S27 for the next image group until the video ends.

[0009] Preferably, the cross-correlation metrics include cosine similarity, Euclidean distance, Mahalanobis distance, SSIM, PSNR, MSE, structural element matching, texture feature matching and shape feature matching.

[0010] Preferably, S3 comprises the following steps: The I frame directly performs inter-frame prediction on the macroblocks in the EP frame. The motion vector and motion compensation are for the macroblocks in the I frame to directly point to the corresponding macroblocks in the EP frame. The residual information is the residual of the EP frame and the reconstructed EP frame with the prediction information.

[0011] The present application also discloses a video error propagation suppression device for high-definition video, which is used to implement the video error propagation suppression method for high-definition video, including: Data receiving module: used to obtain the original video data from the shooting end and identify the original video data format; Image group adjustment module: used to adjust the image transmitted by the data receiving module; Video encoding module: used to encode the data transmitted from the image group adjustment module; Data transmission module: used to transmit the data encoded by the video encoding module.

[0012] Beneficial effects of the present invention: (1) The present invention utilizes the fact that the amount of inter-frame prediction residual data of frames with high correlation is small, and selects coding enhancement reference frames, so that the coding system has a lower bit rate while fully suppressing the error propagation phenomenon. Specifically, the present invention compares the correlation between two frames with a certain distance in the video image group before coding, changes the original reference structure, interrupts the error propagation within the image group, and minimizes the bit rate cost.

[0013] (2) The present invention can control the error propagation within the image group within an acceptable range while the bit rate fluctuates only within a small range, thereby improving the quality of high-definition video during transmission while minimizing the bit rate cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flow chart of a method for suppressing video error propagation for high-definition video according to an embodiment of the present invention.

[0015] Figure 2 The present invention is a flowchart of adjusting the picture group structure of video data in an encoder according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and examples.

[0017] The embodiment of the present invention discloses a method for suppressing the propagation of video errors for high-definition video. The process is as follows: Figure 1 As shown, the following steps are included: S1. Obtain the original video information from the shooting end and send it to the encoder.

[0018] S11, identifying the format of the original video data; S12, if the recognized format is RGB or YUV, it is directly sent to the encoder; S13. If the recognized format is RAWData, the data is first ISP processed, converted into RGB data and then sent to the encoder.

[0019] In a specific embodiment, the specific method for obtaining the original video data is: a video shooting terminal sends the original video data, and the original video data can be sent and received through a wired serial port module or through a wireless communication method such as a WiFi-based local area network, ZigBee node, Bluetooth module, LORA terminal or NB-IoT terminal.

[0020] Specifically, a piece of data obtained from the OV2640 camera module is received through the WiFi transceiver module using the TCP / IP protocol Socket programming. The OV2640 camera module is connected to the WiFi module. The WiFi module is set up to establish a TCP / IP client, and the IP address is set to 192.168.5.3; at the same time, a TCP / IP server is established on the PC, and the local address is set to 192.168.5.3, the subnet mask is 255.255.255.0, and the default gateway is 192.168.5.1, and the WiFi signal is connected.

[0021] After the connection is completed, once the OV2640 camera module and WiFi module are powered on, they will continuously send out signals, and the PC can then establish a connection with them, receive the data and unpack it. The unpacked data is in RGB format.

[0022] S2. Adjust the image group structure of the video data in the encoder. Divide the video into several image groups, and pre-set the video reference relationship within the image group according to the image inter-correlation. The set reference relationship will be substituted into the video encoding module to encode the video. Figure 2 As shown, the following steps are included: S21. Taking the image groups divided by the encoder as units, starting from the first image group divided by the encoder and starting from the video frame that is m frames apart from the I frame, calculate the correlation between each frame and the I frame.

[0023] The cross-correlation metrics include cosine similarity, Euclidean distance, Mahalanobis distance, SSIM, PSNR, MSE, structural element matching, texture feature matching, and shape feature matching.

[0024] S22. Select a frame where the local maximum value of the cross-correlation exists, and record it as an enhanced P frame, that is, an EP frame.

[0025] S23. Starting from the video frames that are m frames apart from the EP frame, calculate the cross-correlation between each frame and the EP frame and the cross-correlation between each frame and the I frame.

[0026] S24, selecting a frame where the local maximum value of the cross-correlation with the previous EP frame is located, and recording it as an enhanced P frame, that is, an EP frame.

[0027] S25, repeat S23 and S24 until the interval between the last EP frame and the last frame of the GOP is less than or equal to m.

[0028] S26. Adjust the reference relationship of the EP frames so that the first EP frame directly references the I frame. If the cross-correlation between the remaining EP frames and the previous EP frame is greater than the cross-correlation between the EP frame and the I frame, then the previous EP frame is directly referenced; if the cross-correlation between the remaining EP frames and the previous EP frame is less than the cross-correlation between the EP frame and the I frame, then the I frame is directly referenced.

[0029] S27. The frames other than the I frame and the EP frame in the image group keep the original reference relationship unchanged.

[0030] S28. Repeat S21-S27 for the next image group until the video ends.

[0031] In a specific embodiment, a segment of video data sent from an OV2640 camera module and received through the method in S1 is adjusted.

[0032] The data obtained is based on the requirements, with the cross-correlation between frames as the standard, starting from the I frame, so that the frame with the highest cross-correlation with the frame in the subsequent frames refers to the frame, and so on until the end of the image group. The cross-correlation measure here is MSE. MSE refers to the mean square error. The smaller the MSE, the higher the cross-correlation between the two frames. The MSE between two frames can be calculated according to the following formula:

[0033] in, and Respectively represent the two images in The pixel value at , the video resolution is .

[0034] S3, encoding the adjusted video data to generate a video code stream. Methods for encoding video data include layered encoding and scalable video encoding.

[0035] In a specific embodiment, video data is sent to the encoder as required, and the reference structure is set to the reference structure adjusted by S2 during inter-frame prediction, so that the I frame directly performs inter-frame prediction on the macroblocks in the EP frame, the motion vector and motion compensation are the macroblocks in the I frame directly pointing to the corresponding macroblocks in the EP frame, the residual information is the residual between the EP frame and the reconstructed EP frame whose prediction information is obtained by the above method, and the encoding output format is a video bitstream.

[0036] S4. Send the encoded video code stream.

[0037] Another embodiment of the present application discloses a video error propagation suppression device for high-definition video, which is used to implement the above-mentioned video error propagation suppression method for high-definition video, including: Data receiving module: used to select the corresponding communication method according to different usage scenarios, obtain the original video data from the shooting end, and identify the original video data format.

[0038] Image group adjustment module: used to adjust the images transmitted by the data receiving module according to the cross-correlation between video frames. The image groups may not be of equal length.

[0039] Video encoding module: used to encode the data transmitted by the image group adjustment module using the corresponding encoding method according to the specific needs of different projects.

[0040] Data transmission module: used to transmit the data encoded by the video encoding module according to the specific requirements of different projects.

[0041] In this implementation, the data receiving module uses the WiFi transmission module to establish a TCP server on the PC side, and uses the WiFi transmission module on the PC to receive the original video data sent by the OV2640 camera module and unpack it. Then the image group adjustment module uses MSE to find the P frame with the highest correlation with the I frame and an interval of not less than 8 frames, and makes it the EP frame. Use MSE to continue to find the P frame with the highest correlation with the previous EP frame and an interval of not less than m frames in the subsequent P frames, and repeat the above steps until the number of remaining P frames is less than m+1 frames. Next, the video encoding module encodes the adjusted data. During encoding, the inter-frame prediction structure is executed according to the result adjusted by the image group adjustment module, so that the I frame or the previous EP frame directly performs inter-frame prediction on the macroblock in the EP frame, the motion vector and motion compensation are the macroblocks in the I frame or the previous EP frame directly pointing to the corresponding macroblocks in the EP frame, and the residual information is the residual of the EP frame and the reconstructed EP frame obtained by the above method. Finally, the data transmission module transmits the encoded video code stream to the receiving end for decoding.

[0042] The above modules can be distributed in one device or in multiple devices. The above modules can be combined into one module or further divided into multiple submodules. The video error propagation suppression device for high-definition video proposed in the embodiment of the present application can be independently set as a separate device or integrated in a base station and a mobile terminal device.

[0043] In summary, to achieve the purpose of fully suppressing the in-group error propagation of video images and solve the problem that the existing methods for containing error propagation cannot control the video bit rate, so as to enable the high-definition video coding and transmission system to work more effectively, the present application utilizes the characteristic that the inter-frame prediction residual data volume of frames with relatively high mutual correlation in all aspects is smaller, so that while the coding system has a smaller bit rate, the error propagation phenomenon is fully suppressed.

[0044] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for suppressing video error propagation for high-definition video, characterized in that: The following steps are involved: S1, obtain the original video information from the shooting end and send it to the encoder; S2, adjusting the picture group structure of the video data in the encoder; S3, encoding the adjusted video data to generate a video code stream; S4. Send the encoded video code stream.

2. The method for suppressing video error propagation for high-definition video according to claim 1, characterized in that: The S1 comprises the following steps: S11, identifying the format of the original video data; S12, if the recognized format is RGB or YUV, it is directly sent to the encoder; S13. If the recognized format is RAWData, the data is first ISP processed, converted into RGB data and then sent to the encoder.

3. The method for suppressing video error propagation for high-definition video according to claim 2, characterized in that: The S2 comprises the following steps: S21, taking the image groups divided by the encoder as units, starting from the video frame that is m frames apart from the I frame, calculating the cross-correlation between each frame and the I frame; S22, selecting a frame where the local maximum value of the cross-correlation is located, and recording it as an enhanced P frame, that is, an EP frame; S23, starting from the video frame that is m frames away from the EP frame, calculate the cross-correlation between each frame and the EP frame and the cross-correlation between each frame and the I frame; S24, selecting a frame where the local maximum value of the cross-correlation with the previous EP frame is located, and recording it as an enhanced P frame, that is, an EP frame; S25, repeat S23 and S24 until the interval between the last EP frame and the last frame of the picture group is less than or equal to m; S26, adjusting the reference relationship of the EP frames so that the first EP frame directly references the I frame. If the cross-correlation between the remaining EP frames and the previous EP frame is greater than the cross-correlation between the EP frame and the I frame, the previous EP frame is directly referenced; if the cross-correlation between the remaining EP frames and the previous EP frame is less than the cross-correlation between the EP frame and the I frame, the I frame is directly referenced. S27, the frames other than the I frame and the EP frame in the image group keep the original reference relationship unchanged; S28. Repeat S21-S27 for the next image group until the video ends.

4. The method for suppressing video error propagation for high-definition video according to claim 3, characterized in that: The cross-correlation metrics include cosine similarity, Euclidean distance, Mahalanobis distance, SSIM, PSNR, MSE, structural element matching, texture feature matching, and shape feature matching.

5. The method for suppressing video error propagation for high-definition video according to claim 4, characterized in that: The S3 comprises the following steps: The I frame directly performs inter-frame prediction on the macroblocks in the EP frame. The motion vector and motion compensation are for the macroblocks in the I frame to directly point to the corresponding macroblocks in the EP frame. The residual information is the residual of the EP frame and the reconstructed EP frame with the prediction information.

6. A video error propagation suppression device for high-definition video, characterized in that: The method for suppressing the propagation of video errors for high-definition video according to any one of claims 1 to 5 comprises: Data receiving module: used to obtain the original video data from the shooting end and identify the original video data format; Image group adjustment module: used to adjust the image transmitted by the data receiving module; Video encoding module: used to encode the data transmitted from the image group adjustment module; Data transmission module: used to transmit the data encoded by the video encoding module.