A low-latency ultra-high-definition video encoding method, decoding method and their devices
By establishing a prediction template by pixel-by-pixel processing and combining dictionary compression coding, the problems of high delay and large hardware overhead in the prior art are solved, and low-latency and lossless video compression effects are achieved.
Patent Information
- Application Number
- CN202211479889.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-11-23
AI Technical Summary
The existing video compression technology has high delay and high hardware overhead in ultra-high-definition video streaming data processing, and the data output after decompression is lossy and affects the user's viewing experience.
The low-latency ultra-high-definition video encoding method is adopted to establish a prediction template through pixel-by-pixel processing, calculate the prediction value and obtain the residual coefficient, perform data compression processing, and combine dictionary compression coding to reduce data redundancy.
It improves the encoding and decoding efficiency of ultra-high-definition video stream data, reduces delay, and realizes lossless compression processing, which is suitable for real-time display devices.
Smart Images

Figure CN115834907B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a low-latency ultra-high definition video encoding method, a decoding method and an apparatus thereof. Background Art
[0002] With the development of video technology, the amount of data carried by current ultra-high definition video streams is extremely large, which has a huge impact on video transmission and storage. Moreover, the video stream data contains redundant data. If these data can be compressed, the transmission volume of the video stream data can be reduced. Therefore, video compression has great practical significance.
[0003] In the existing video compression technology, compression processing is mainly performed through a relatively complex compression coding framework, and the video stream can be compressed to a relatively small level, but its latency is high and it will also bring additional hardware overhead costs in actual applications. The data output after decompression is lossy and affects the user viewing experience, and it is not suitable for application in ultra-high definition real-time display devices using wired transmission. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a low-latency ultra-high definition video encoding method, a decoding method and an apparatus thereof, so as to improve the encoding and decoding efficiency and low latency of ultra-high definition video stream data.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A low-latency ultra-high definition video encoding method includes the following steps:
[0006] Step S10: Obtain the video stream data to be compressed;
[0007] Step S20: Establish a prediction template according to the current pixel to be processed, and calculate the predicted value of the pixel to be predicted;
[0008] Step S30: Obtain a residual coefficient according to the predicted value and the original value of the pixel to be predicted;
[0009] Step S40: Perform data compression processing according to the residual coefficient and the original value of the pixel to be predicted, and according to the above situation;
[0010] Step S50: Remap the data address of the first compressed encoding data stream to generate a second compressed encoding data stream;
[0011] Step S60: Perform dictionary compression encoding on the second compressed encoding data stream to generate a third compressed encoding data stream.
[0012] In a preferred embodiment, the step S20 includes:
[0013] Step S21: Calculate the image feature values of the prediction template, including horizontal feature coefficients, vertical feature coefficients, right diagonal feature coefficients, and left diagonal feature coefficients;
[0014] Step S22: Determine the texture structure of the prediction template according to the image feature values of the prediction template, divide and cut the prediction template to obtain a division and cutting result;
[0015] Step S23: Calculate the predicted value of the pixel to be predicted according to the division and cutting result;
[0016] The algorithm formula adopted in Step S21 is:
[0017]
[0018] where x is the horizontal feature coefficient, y is the vertical feature coefficient, z1 is the right diagonal feature coefficient, z2 is the left diagonal feature coefficient, and A1 to A7, B1 to B6, C1, and C2 are the pixel values of the indicated reference pixels.
[0019] In a preferred embodiment, Step S22 includes the following situations:
[0020] In the first situation, when the horizontal feature coefficient is simultaneously less than a certain value of the vertical feature coefficient, the right diagonal feature coefficient, and the left diagonal feature coefficient, that is, when x < (y - 4), x < (z1 - 4), x < (z2 - 4), it is determined that the texture structure of the prediction template is a horizontal texture structure, and it is determined that the prediction template is divided in a horizontal cutting manner;
[0021] In the second situation, when the vertical feature coefficient is simultaneously less than a certain value of the horizontal feature coefficient, the right diagonal feature coefficient, and the left diagonal feature coefficient, that is, when y < (x - 4), y < (z1 - 4), y < (z2 - 4), it is determined that the texture structure of the prediction template is a vertical texture structure, and it is determined that the prediction template is divided in a vertical cutting manner;
[0022] In the third situation, when the right diagonal feature coefficient is simultaneously less than a certain value of the vertical feature coefficient, the horizontal feature coefficient, and the left diagonal feature coefficient, that is, when z1 < (y - 4), z1 < (x - 4), z1 < (z2 - 4), it is determined that the texture structure of the prediction template is a right diagonal texture structure, and it is determined that the prediction template is divided in a right diagonal cutting manner;
[0023] In the fourth case, when the left oblique feature coefficient is simultaneously less than the vertical feature coefficient, the right oblique feature coefficient, and the horizontal feature coefficient by a certain value, that is, when z2 < (y - 4), z2 < (z1 - 4), z2 < (x - 4), it is determined that the texture structure of the prediction template is a left oblique texture structure, and it is determined that the prediction template is divided in a left oblique cutting manner;
[0024] In the fifth case, when the horizontal feature coefficient and the vertical feature coefficient are simultaneously less than the right oblique feature coefficient and less than the left oblique feature coefficient by a certain value and the difference between the horizontal feature coefficient and the vertical feature coefficient is within a certain range, that is, when |x - y| < 4, x < (z1 - 4), x < (z2 - 4), y < (z1 - 4), y < (z2 - 4), it is determined that the texture structure of the prediction template is a cross texture structure, and it is determined that the prediction template is divided in a cross cutting manner;
[0025] In the sixth case, when the right oblique feature coefficient and the left oblique feature coefficient are simultaneously less than the horizontal feature coefficient and less than the vertical feature coefficient by a certain value and the difference between the left oblique feature coefficient and the right oblique feature coefficient is within a certain range, that is, when |z1 - z2| < 4, z1 < (x - 4), z1 < (y - 4), z2 < (x - 4), z2 < (y - 4), it is determined that the texture structure of the prediction template is a cross texture structure, and it is determined that the prediction template is divided in a cross cutting manner;
[0026] In the seventh case, when the differences between the horizontal feature coefficient, the vertical feature coefficient, the right oblique feature coefficient, and the left oblique feature coefficient are within a certain range, that is, when |x - y| < 4, |x - z1| < 4, |x - z2| < 4, |y - z1| < 4, |y - z2| < 4, |z1 - z2| < 4, it is determined that the texture structure of the prediction template is a flat texture structure, and it is determined that the prediction template does not need to be divided and cut;
[0027] The step S23 includes:
[0028] In the first case, when it is horizontal cutting, the prediction value calculation formula of the pixel to be predicted is:
[0029]
[0030] In the second case, when it is vertical cutting, the prediction value calculation formula of the pixel to be predicted is:
[0031]
[0032] In the third case, when it is right oblique cutting, the prediction value calculation formula of the pixel to be predicted is:
[0033]
[0034] In the fourth case, when it is a left diagonal cut, the calculation formula for the predicted value of the pixel to be predicted is as follows:
[0035]
[0036] In the fifth case, when it is a cross cut, the calculation formula for the predicted value of the pixel to be predicted is as follows:
[0037]
[0038] In the sixth case, when it is a crosswise cut, the calculation formula for the predicted value of the pixel to be predicted is as follows:
[0039]
[0040] In a preferred embodiment, the step S40 includes:
[0041] Step S41: When receiving the data stream of the residual coefficient and the original value of the pixel to be predicted, judge the residual coefficient to determine whether the residual coefficient is less than the first threshold and greater than the second threshold. If it is established, execute step S42; otherwise, execute step S43, where the first threshold includes, but is not limited to, 14, and the second threshold includes, but is not limited to, -14;
[0042] Step S42: Truncate and encode the residual coefficient to generate an output first compression code, and then jump to execute step S44;
[0043] Step S43: Encode the original value of the pixel to be predicted to generate an output first compression code;
[0044] Step S44: Repeat steps S41 to S43 until the data stream of the residual coefficient and the original value of the pixel to be predicted stops being input;
[0045] The step S60 includes:
[0046] Step S61: Initialize, restore the coding dictionary to its initial state. At this time, the dictionary only includes all default items, and clear the first character, the first string, and the second string;
[0047] Step S62: Receive the second compression code data stream, and read the data from the second compression code data stream as the first character;
[0048] Step S63: Combine the first character and the first string to form the second string;
[0049] Step S64: Check whether the second string exists in the coding dictionary. If it exists, execute step S65; otherwise, execute step S66;
[0050] Step S65: Determine that the storage address of the second string corresponding to the encoding dictionary matches successfully, output the encoding dictionary address mark corresponding to the second string, generate the third compressed encoding data, and update the first string, then execute step S67;
[0051] Step S66: Determine that the second string fails to match in the encoding dictionary storage. First, check whether the encoding dictionary has reached the storage limit. If so, output the mark of the first string to generate the third compressed encoding data and then update the first string; if not, first store the second character in the encoding dictionary and mark it, then output the mark of the first string to generate the third compressed encoding data and update the first string;
[0052] Step S67: Check whether the first string has reached the splicing limit. If it has reached the limit, clear the first string; if not, there is no need to update the first string;
[0053] Step S68: Repeat steps S61 to S67 until the output of the second compressed encoding data is completed.
[0054] The present invention also provides a low-latency ultra-high-definition video decoding method, which is characterized by including the following steps:
[0055] Step S70: Obtain the video stream data to be decoded;
[0056] Step S80: Perform dictionary decoding on the video data bitstream to be decoded to generate the first decoded data bitstream;
[0057] Step S90: Perform data address inverse mapping on the first decoded data bitstream to generate the second decoded data bitstream;
[0058] Step S100: Perform structural decoding operations on the second decoded data bitstream to obtain the original pixel values or residual coefficients;
[0059] Step S110: According to the original pixel values or residual coefficients and in combination with the template feature coefficients, perform restoration to finally obtain the original data bitstream.
[0060] In a preferred embodiment, step S80 includes:
[0061] Step S81: Initialize, restore the decoding dictionary to its initial state. At this time, the decoding dictionary only includes all default items, and clear the third character, third string, fourth string, and fifth string;
[0062] Step S82: Receive the video data bitstream to be decoded, and read data from the video data bitstream to be decoded as the first token;
[0063] Step S83: Check whether there is a corresponding string or character in the decoding dictionary for the first token. If so, execute Step S84; otherwise, execute Step S85;
[0064] Step S84: Determine that there is a corresponding character or string in the decoding dictionary for the first token, read the character or string stored in the decoding dictionary, update it as the third string and output it, generate the first decoded data bitstream, and then determine whether the fourth string has reached the splicing limit. If it has reached the limit, only update the fourth string without updating the decoding dictionary. If it has not reached the limit, splice the first character of the third string and the fourth string to form the fifth string, store the fifth string in the decoding dictionary, and add a new token mapping; then jump to execute Step S86;
[0065] Step S85: Determine that there is no corresponding character or string in the decoding dictionary for the first token, and determine whether the fourth string has reached the splicing limit. If it has reached the limit, only update the fourth string without updating the decoding dictionary. If it has not reached the limit, splice the fourth string and its first character to form the fifth string, store the fifth string in the decoding dictionary, and add a new token mapping, where the new token is the first token; update the third string, with the updated value being the fifth string, and then output the fifth string to generate the first decoded data bitstream;
[0066] Step S86: Update the fourth string to the value of the third string;
[0067] Step S87: Repeat Steps S81 to S86 until the input of the video data bitstream to be decoded stops.
[0068] The present invention also provides a video encoding device, which includes:
[0069] An encoding control signal synchronization module, configured to obtain the control signal of the video of the data to be encoded and synchronously output the control signal when outputting the compressed encoded data;
[0070] A video data encoding processing module, configured to perform encoding processing operations on video data;
[0071] Among them, the video data encoding processing module includes:
[0072] A data acquisition and caching sub-module, configured to receive the video data stream to be compressed and perform data caching operations;
[0073] A predictor sub-module for calculating the predicted value of the pixel to be compressed and processed;
[0074] A data compression and reconstruction sub-module for generating first compressed encoded data based on the predicted value of the pixel to be compressed and processed and the original value of the compressed processed pixel, and then generating second compressed encoded data according to the first compressed encoded data;
[0075] A data cache sub-module for receiving the second compressed encoded data output by the data compression and reconstruction sub-module, performing cache control and outputting the second compressed encoded data;
[0076] A dictionary compression sub-module for processing the second compressed encoded data, dynamically establishing a content search value memory for implementing string search and matching functions, and finally generating third compressed encoded data;
[0077] An output control sub-module for receiving the third compressed encoded data and performing cache processing, and logically controlling the output of the complete third compressed encoded data.
[0078] The present invention also provides a video decoding device, which includes:
[0079] A decoding control signal synchronization module for obtaining the control signal of the video of the data to be decoded and synchronizing the decoded output control signal when outputting the decoded data stream;
[0080] A video data decoding and processing module for performing decoding processing operations on the compressed encoded data;
[0081] Wherein, the video data decoding and processing module includes:
[0082] A decoded data cache control sub-module for receiving the video data stream to be decoded and performing cache operations;
[0083] A data dictionary decoding sub-module for performing dictionary decoding operations on the video data stream to be decoded to obtain a first decoded data stream;
[0084] A data decoding and restoring sub-module for processing the first decoded data stream, first performing data inverse mapping to obtain a second decoded data stream, and then performing structure decoding operations on it to finally obtain the decoded data stream;
[0085] A decoded data output control sub-module for receiving the decoded data stream and performing cache processing, and logically controlling the output of the complete decoded data stream.
[0086] The present invention also provides a video encoding and transmitting device, which includes: a processor, a memory, a receiver and a transmitter. Program codes are stored in the memory, and the processor executes the program codes to perform the low-latency ultra-high-definition video encoding method described above.
[0087] The present invention also provides a video decoding and receiving device, which includes: a processor, a memory, a receiver and a transmitter. Program codes are stored in the memory, and the processor executes the program codes to perform the low-latency ultra-high-definition video decoding method described above.
[0088] Compared with the prior art, the present invention has the following beneficial effects: The present invention performs pixel-by-pixel processing on video stream data and establishes a prediction template for prediction based on each pixel, improving the accuracy of the prediction model. At the same time, through a lightweight data reconstruction coding model, lossless compression processing is performed on the residuals and the original pixels. By taking advantage of the characteristics of the data reconstruction coding model, it is combined with dictionary compression coding. After data reconstruction, dictionary compression coding is further performed to further compress the code stream and eliminate redundant information between data. And through multi-module parallel processing, the latency of the overall compression coding algorithm is reduced. Therefore, the present invention can perform low-latency lossless video compression processing on ultra-high-definition video data code streams. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 is a flowchart of the low-latency ultra-high-definition video encoding method provided by an embodiment of the present invention;
[0090] Figure 2 is a flowchart of the low-latency ultra-high-definition video decoding method provided by an embodiment of the present invention;
[0091] Figure 3 is a block diagram of the video encoding device provided by an embodiment of the present invention;
[0092] Figure 4 is a block diagram of the video decoding device provided by an embodiment of the present invention;
[0093] Figure 5 is a schematic structural diagram of the video encoding and transmitting device provided by an embodiment of the present invention;
[0094] Figure 6 is a schematic structural diagram of the video decoding and receiving device provided by an embodiment of the present invention;
[0095] Figure 7 is a schematic diagram of the prediction template structure provided by an embodiment of the present invention;
[0096] Figure 8 is a schematic diagram of the prediction template division and cutting provided by an embodiment of the present invention;
[0097] Figure 9 It is a flowchart of a data compression and encoding processing method provided by an embodiment of the present invention based on the above situation;
[0098] Figure 10 It is a flowchart of data address remapping for compressed and encoded data provided by an embodiment of the present invention;
[0099] Figure 11 It is a flowchart of data address inverse mapping for decoded data provided by an embodiment of the present invention. Detailed implementation manners
[0100] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0101] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0102] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0103] To solve the above technical problems, according to a first aspect, an embodiment of the present invention provides a low-latency ultra-high video encoding method, which will be described in detail below with reference to the accompanying drawings.
[0104] S10. Obtain the video stream data to be compressed and process it based on a per-pixel manner, that is, each pixel is processed separately;
[0105] S20. Establish a prediction template based on the currently processed pixel and calculate the predicted value of the pixel to be predicted;
[0106] Further, the step S20 specifically includes:
[0107] S21. Calculate the predicted template image feature values of the pixel to be predicted, including horizontal feature coefficients, vertical feature coefficients, right oblique feature coefficients, and left oblique feature coefficients;
[0108] As Figure 7 shown is a schematic diagram of the prediction template of the pixel to be predicted. The prediction template is a trapezoidal structure and contains a total of 15 pixels. x MED For the pixel to be predicted (A1 to A7, B1 to B6, C1, C2) as the reference pixels.
[0109] Further, the algorithm formula adopted in step S21 is as follows:
[0110]
[0111] where x is the horizontal feature coefficient, y is the vertical feature coefficient, z1 is the right oblique feature coefficient, z2 is the left oblique feature coefficient, and A1 to A7, B1 to B6, C1, and C2 are the pixel values of the indicated reference pixels;
[0112] S22. Determine the texture structure of the prediction template according to the image feature value of the prediction template, divide and cut the prediction template to obtain a division and cutting result;
[0113] In a possible case, when the horizontal feature coefficient is simultaneously less than a certain value of the vertical feature coefficient, the right oblique feature coefficient, and the left oblique feature coefficient, that is, when x < (y - 4), x < (z1 - 4), x < (z2 - 4), it is determined that the texture structure of the prediction template is a horizontal texture structure, and it is determined that the prediction template is divided in a horizontal cutting manner;
[0114] In a possible case, when the vertical feature coefficient is simultaneously less than a certain value of the horizontal feature coefficient, the right oblique feature coefficient, and the left oblique feature coefficient, that is, when y < (x - 4), y < (z1 - 4), y < (z2 - 4), it is determined that the texture structure of the prediction template is a vertical texture structure, and it is determined that the prediction template is divided in a vertical cutting manner;
[0115] In a possible case, when the right oblique feature coefficient is simultaneously less than a certain value of the vertical feature coefficient, the horizontal feature coefficient, and the left oblique feature coefficient, that is, when z1 < (y - 4), z1 < (x - 4), z1 < (z2 - 4), it is determined that the texture structure of the prediction template is a right oblique texture structure, and it is determined that the prediction template is divided in a right oblique cutting manner;
[0116] In a possible case, when the left oblique feature coefficient is simultaneously less than a certain value of the vertical feature coefficient, the right oblique feature coefficient, and the horizontal feature coefficient, that is, when z2 < (y - 4), z2 < (z1 - 4), z2 < (x - 4), it is determined that the texture structure of the prediction template is a left oblique texture structure, and it is determined that the prediction template is divided in a left oblique cutting manner;
[0117] A possible situation is that when the horizontal feature coefficient and the vertical feature coefficient are both less than the right - oblique feature coefficient and less than the left - oblique feature coefficient by a certain value, and the difference between the horizontal feature coefficient and the vertical feature coefficient is within a certain range, that is, when |x - y| < 4, x < (z1 - 4), x < (z2 - 4), y < (z1 - 4), y < (z2 - 4), then it is determined that the texture structure of the prediction template is a cross texture structure, and it is determined that the prediction template is divided in a cross - cutting manner;
[0118] A possible situation is that when the right - oblique feature coefficient and the left - oblique feature coefficient are both less than the horizontal feature coefficient and less than the vertical feature coefficient by a certain value, and the difference between the left - oblique feature coefficient and the right - oblique feature coefficient is within a certain range, that is, when |z1 - z2| < 4, z1 < (x - 4), z1 < (y - 4), z2 < (x - 4), z2 < (y - 4), then it is determined that the texture structure of the prediction template is a cross - intersection texture structure, and it is determined that the prediction template is divided in a cross - intersection cutting manner;
[0119] A possible situation is that when the difference between the horizontal feature coefficient, the vertical feature coefficient, the right - oblique feature coefficient and the left - oblique feature coefficient is within a certain range, that is, when |x - y| < 4, |x - z1| < 4, |x - z2| < 4, |y - z1| < 4, |y - z2| < 4, |z1 - z2| < 4, then it is determined that the texture structure of the prediction template is a flat texture structure, and it is determined that the prediction template does not need to be divided and cut.
[0120] As Figure 8 shown are the schematic diagrams of the horizontal cutting, vertical cutting, right - oblique cutting, left - oblique cutting, cross - cutting, and cross - intersection cutting;
[0121] S23. According to the division and cutting result and the horizontal feature coefficient, vertical feature coefficient, right - oblique feature coefficient and left - oblique feature coefficient, calculate the predicted value of the pixel to be predicted;
[0122] A possible situation is that when it is horizontal cutting, the calculation formula for the predicted value of the pixel to be predicted is:
[0123]
[0124] A possible situation is that when it is vertical cutting, the calculation formula for the predicted value of the pixel to be predicted is:
[0125]
[0126] A possible situation is that when it is right - oblique cutting, the calculation formula for the predicted value of the pixel to be predicted is:
[0127]
[0128] A possible situation is that when it is a horizontal cut, the calculation formula for the predicted value of the pixel to be predicted is as follows:
[0129]
[0130] A possible situation is that when it is a cross cut, the calculation formula for the predicted value of the pixel to be predicted is as follows:
[0131]
[0132] A possible situation is that when it is a crosswise cut, the calculation formula for the predicted value of the pixel to be predicted is as follows:
[0133]
[0134] S30. According to the predicted value and the original value of the pixel to be predicted, subtract them to obtain the residual coefficient. The formula is
[0135] resid = x og - x MED ;
[0136] where resid is the residual coefficient, x og is the original value of the pixel to be predicted and x MED is the predicted value of the pixel to be predicted;
[0137] S40. According to the residual coefficient and the original pixel value, and perform the first-step data compression encoding process according to the above situation, as Figure 9 shown;
[0138] Furthermore, the specific steps of step S40 include:
[0139] S41. When receiving the data stream of the residual coefficient and the original value of the pixel to be predicted, judge the residual coefficient to determine whether the residual coefficient is less than the first threshold and greater than the second threshold. If it holds, execute step S42; otherwise, execute step S43. The first threshold includes, but is not limited to, 14, and the second threshold includes, but is not limited to, -14;
[0140] S42. Truncate and encode the residual coefficient to generate the first compressed encoding output, and then jump to execute step S44;
[0141] S43. Encode the original value of the pixel to be predicted to generate the first compressed encoding output;
[0142] S44. Repeat steps S41 to S43 until the data stream of the residual coefficient and the original value of the pixel to be predicted stops inputting;
[0143] S50. Remap the data addresses of the first compressed encoded data stream to generate a second compressed encoded data stream. The specific process is as follows Figure 10 shown, where data_press_1 is the first compressed encoded data, data_press_2 is the second compressed encoded data, and i is a variable required to be calculated in the process;
[0144] S60. Perform dictionary compression encoding on the second compressed encoded data stream to generate a third compressed encoded data stream;
[0145] Further, step S60 includes:
[0146] S61. Initialize, restore the encoding dictionary to its initial state. At this time, the dictionary only includes all default items, and clear the first character, the first string, and the second string;
[0147] S62. Receive the second compressed encoded data stream, and read data from the second compressed encoded data stream as the first character;
[0148] S63. Combine the first character and the first string to form a second string;
[0149] S64. Check whether the second string exists in the encoding dictionary. If it exists, execute step S65; otherwise, execute step S66;
[0150] S65. Determine that the storage address of the second string in the encoding dictionary matches successfully, output the encoding dictionary address mark corresponding to the second string, generate the third compressed encoded data, and update the first string, then execute step S67;
[0151] S66. Determine that the second string does not match in the encoding dictionary storage. First, check whether the encoding dictionary has reached the storage limit. If so, output the mark of the first string to generate the third compressed encoded data and then update the first string; if not, first store the second character in the encoding dictionary and mark it, then output the mark of the first string to generate the third compressed encoded data and then update the first string;
[0152] S67. Check whether the first string has reached the splicing limit. If it has reached the limit, clear the first string; if not, there is no need to update the first string;
[0153] S68. Repeat steps S61 to S67 until the output of the second compressed encoded data is completed.
[0154] By performing the above steps, the low-latency ultra-high video encoding method provided by the embodiments of the present invention obtains the video stream data to be compressed, calculates the pixel prediction value based on the prediction template established for each pixel, obtains the residual coefficient using the pixel prediction value and the original value, and performs data compression processing according to the context. Then, the obtained compressed data is notified to the dictionary compression for further compression processing to reduce the redundancy between data. In the whole process, not only the image features around each pixel in the video stream data are fully considered to establish a prediction template and the prediction value is obtained in an efficient and accurate calculation manner, but also lossless compression processing is performed using the context relationship and the relationship of the data structure when compressing the data for the first time. Finally, in the dictionary compression processing, by dynamically establishing a dictionary, the correlation between data is better utilized for compression processing. Therefore, the present invention can efficiently perform lossless compression encoding processing on ultra-high video under low-latency requirements.
[0155] According to a second aspect, an embodiment of the present invention provides a video decoding method, the method comprising:
[0156] S70. Obtain the video data code stream to be decoded, where the video data code stream to be decoded is a data code stream that has been encoded multiple times, so multiple decoding operations are required;
[0157] S80. Perform dictionary decoding on the video data code stream to be decoded to generate a first decoded data code stream;
[0158] Further, the step S80 includes:
[0159] S81. Initialize, restore the decoding dictionary to its initial state. At this time, the decoding dictionary only includes all default items, and clear the third character, the third string, the fourth string, and the fifth string;
[0160] S82. Receive the video data code stream to be decoded, and read the data from the video data code stream to be decoded as the first token;
[0161] S83. Check whether there is a corresponding string or character in the decoding dictionary for the first token. If so, execute step S84; otherwise, execute S85;
[0162] S84. Determine that there is a corresponding character or string in the decoding dictionary for the first token, read the character or string stored in the decoding dictionary, update it as the third string and output it to generate a first decoded data code stream. Then, determine whether the fourth string has reached the splicing limit. If it has reached the limit, only update the fourth string without updating the decoding dictionary. If it has not reached the limit, splice the first character of the third string and the fourth string to form the fifth string, and store the fifth string in the decoding dictionary to add a new token mapping. Jump to execute step S106;
[0163] S85. Determine that there is no corresponding stored character or string for the first mark in the decoding dictionary, and judge whether the fourth string has reached the splicing limit. If it has reached the limit, do not update the decoding dictionary and only update the fourth string. If it has not reached the limit, splice the fourth string and its first character to form a fifth string, store the fifth string in the decoding dictionary, add a new mark mapping, and the new mark is the first mark. Then update the third string, with the updated value being the fifth string, and then output the fifth string to generate the first decoded data bitstream;
[0164] S86. Update the fourth string to the value of the third string;
[0165] S87. Repeat steps S81 to S86 until the input of the video data bitstream to be decoded stops;
[0166] S90. Perform data address inverse mapping on the first decoded data bitstream to generate the second decoded data bitstream. The specific process is as Figure 11 shown, where data_re_1 is the first decoded data, data_re_2 is the second decoded data, and j is the variable to be calculated in the process;
[0167] S100. Perform a decoding operation on the second decoded data bitstream to obtain the original pixel value or decoded residual coefficient, and form the third decoded data stream;
[0168] In a possible case, first judge the highest bit of the second decoded data. If the highest bit is 1, the pixel original value is the lower data bits. If the highest bit is not 1, split the lower data into two halves. The first half is the first decoded residual coefficient, and for the second half, if all bits are 1, discard this part. If not all bits are 1, this part is the second decoded residual coefficient. Finally, package according to the possible obtained pixel original value, first decoded residual coefficient, and second decoded residual coefficient to form the third decoded data stream;
[0169] S110. According to the obtained third decoded data stream composed of the original pixel value or decoded residual coefficient and in combination with the template feature coefficient, perform restoration to finally obtain the original data bitstream;
[0170] Further, the step S110 includes:
[0171] S111. First judge whether the third decoded data is received. If it is the original pixel value, directly output it. If it is the decoded residual coefficient, execute step S142;
[0172] S112. Establish a template based on the current original pixel value to be decoded. The template is the same as that in step S20 Figure 7Consistently convert the to-be-predicted pixels therein into to-be-decoded pixels, and similarly calculate the template image feature values, determine the division cutting result, and obtain the predicted value of the to-be-decoded pixels. The specific implementation is the same as that of steps S21, S22, and step S23, and will not be repeated here;
[0173] S113. Calculate the original value of the to-be-decoded pixel according to the obtained decoding residual coefficient and the predicted value of the to-be-decoded pixel;
[0174] An embodiment of the present invention provides a video encoding device, such as Figure 3 shown. This device can be used to implement the functional modules of the first aspect or any of the methods in the first aspect. The device includes:
[0175] An encoding control signal synchronization module 100, configured to obtain the control signals of the video of the to-be-encoded data, specifically including but not limited to field synchronization signals, line synchronization signals, data valid signals, and reset signals, and synchronously output the control signals when outputting compressed encoded data;
[0176] A video data encoding and processing module 200, configured to perform encoding processing operations on video data;
[0177] Among them, the video data encoding and processing module includes:
[0178] A data acquisition and caching sub-module 201, configured to receive the video data stream to be compressed and perform data caching operations;
[0179] A predictor sub-module 202, configured to calculate the predicted value of the pixel to be compressed, specifically to implement the related descriptions of steps S21, S22, and step S23, and will not be repeated here;
[0180] A data compression and reconstruction sub-module 203, configured to generate first compressed encoded data based on the predicted value of the pixel to be compressed and the original value of the pixel to be compressed, and then regenerate second compressed encoded data according to the first compressed encoded data. Specifically, it is to implement the related descriptions of steps S30, S40, and step S50, and will not be repeated here;
[0181] A data caching sub-module 204, configured to receive the second compressed encoded data output by the data compression and reconstruction sub-module, perform caching control, and output the second compressed encoded data;
[0182] A dictionary compression sub-module 205, configured to process the second compressed encoded data, dynamically establish a content search value memory for implementing string search and matching functions, and finally generate third compressed encoded data. Specifically, it is to implement the related description of step S60, and will not be repeated here;
[0183] The output control sub-module 206 is configured to receive the third compressed and encoded data, perform caching processing, and output the complete third compressed and encoded data under the control of logic.
[0184] According to a fourth aspect, an embodiment of the present invention provides a video decoding device, which can be used to implement the functional modules of the second aspect or any of the methods in the second aspect. The device includes:
[0185] An S300 decoding control signal synchronization module, configured to obtain the control signal of the video of the data to be decoded and synchronize the decoded output control signal when outputting the decoded data stream;
[0186] An S400 video data decoding and processing module, configured to perform decoding processing operations on the compressed and encoded data;
[0187] Wherein, the video data decoding and processing module includes:
[0188] An S401 decoded data caching control sub-module, configured to receive the video data stream to be decoded and perform caching operations;
[0189] An S402 data dictionary decoding sub-module, configured to perform dictionary decoding operations on the video data stream to be decoded to obtain a first decoded data stream, specifically implementing the relevant descriptions of step S80. For the convenience and conciseness of description, it will not be repeated here;
[0190] An S403 data decoding and restoring sub-module, configured to first perform data address inverse mapping on the first decoded data stream to obtain a second decoded data stream, and then perform decoding operations on it to finally obtain the decoded data stream, specifically implementing the relevant descriptions of steps S90, S100, and S110. For the convenience and conciseness of description, it will not be repeated here;
[0191] An S404 decoded data output control sub-module, configured to receive the decoded data stream and perform caching processing, and output the complete decoded data stream under the control of logic;
[0192] An embodiment of the present invention provides a video encoding and transmitting device 500, as Figure 5 shown. The device may include, but is not limited to, a first processor 501, a first memory 502, a first data receiver 503, and a first data transmitter 504, wherein the first processor 501, the first memory 502, the first data receiver 503, and the first data transmitter 504 are connected through a bus 505.
[0193] The processor 501 is used to provide control and computing processing to ensure the operation of the video encoding and transmitting device 500. The processor 501 can be, but is not limited to, a central processing unit (CPU), other general-purpose processors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., chips, or a combination of the above types of chips.
[0194] The memory 502, as a non-transitory readable storage medium, can be used to store non-transitory software programs, non-transitory executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. The memory 502 can be, but is not limited to, a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The processor 501 executes various functional applications and data processing of the processor 501 by running the non-transitory software programs, instructions, and modules stored in the memory 502, that is, implements the low-latency ultra-high-definition video encoding method in the above embodiments.
[0195] According to the video encoding device, one or more modules are stored in the memory 502 and, when executed by the processor 501, execute the low-latency ultra-high-definition video encoding method in the above embodiments.
[0196] The first data receiver 503 is used to receive and obtain the video stream data to be compressed, and it can be, but is not limited to, an HDMI interface, a UART interface, an IIC interface, an SPI interface, a VBONE interface, etc.;
[0197] The first data transmitter 502 is used to transmit the video compression encoding calculated by the first processor 501, and it can be, but is not limited to, an HDMI interface, a UART interface, an IIC interface, an SPI interface, a VBONE interface, etc.;
[0198] The bus 505 is used to build an information transmission path among the first processor 501, the first memory 502, the first data receiver 503, and the first data transmitter 504 in the low-latency ultra-high-definition video decoding and transmitting device 500 and perform information transmission.
[0199] An embodiment of the present invention provides a video decoding and receiving device 600, such as Figure 6As shown, the device may include, but is not limited to, a second processor 601, a second memory 602, a second data receiver 603, and a second data transmitter 604, wherein the second processor 601, the second memory 602, the second data receiver 603, and the second data transmitter 604 are connected by a bus 605.
[0200] The definitions and explanations of the various modules of the above video encoding and transmitting device 500 also apply to the video decoding and receiving device 600, and will not be described in detail here.
[0201] The processor 601 is used to provide control and computing processing to ensure the operation of the video codec receiving device 600. The processor 601 executes various functional applications and data processing of the processor 601 by running non-transitory software programs, instructions, and modules stored in the memory 602, that is, implements the low-latency ultra-high-definition video decoding method in the above embodiments.
[0202] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0203] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A low-latency ultra-high-definition video encoding method, characterized in that, It includes the following steps: Step S10: Obtain the video stream data to be compressed; Step S20: Establish a prediction template based on the current pixel to be processed, and calculate the predicted value of the pixel to be predicted; Step S30: Obtain the residual coefficient according to the predicted value and the original value of the pixel to be predicted; Step S40: Perform data compression processing according to the residual coefficient and the original value of the pixel to be predicted, and according to the above situation; Step S50: Remap the data address of the first compressed coding data stream to generate a second compressed coding data stream; Step S60: Perform dictionary compression coding on the second compressed coding data stream to generate a third compressed coding data stream; Step S20 includes: Step S21: Calculate the image feature values of the prediction template, including horizontal feature coefficient, vertical feature coefficient, right oblique feature coefficient, and left oblique feature coefficient; Step S22: Determine the texture structure of the prediction template according to the image feature values of the prediction template, divide and cut the prediction template to obtain the division and cutting result; Step S23: Calculate the predicted value of the pixel to be predicted according to the division and cutting result; The algorithm formula adopted in Step S21 is: where x is the horizontal feature coefficient, y is the vertical feature coefficient, z1 is the right oblique feature coefficient, z2 is the left oblique feature coefficient, A1 to A7, B1 to B6, C1, and C2 are the pixel values of the indicated reference pixels; Steps S22 and S23 include the following situations: When x < (y - 4), x < (z1 - 4), x < (z2 - 4), it is determined that the texture structure of the prediction template is a horizontal texture structure, and it is determined that the prediction template is divided in a horizontal cutting manner; the calculation formula for the predicted value of the pixel to be predicted is: When y < (x - 4), y < (z1 - 4), y < (z2 - 4), it is determined that the texture structure of the prediction template is a vertical texture structure, and it is determined that the prediction template is divided in a vertical cutting manner; the calculation formula for the predicted value of the pixel to be predicted is: When z1 < (y - 4), z1 < (x - 4), z1 < (z2 - 4), it is determined that the texture structure of the prediction template is a right oblique texture structure, and it is determined that the prediction template is divided in a right oblique cutting manner; the calculation formula for the predicted value of the pixel to be predicted is: When z2 < (y - 4), z2 < (z1 - 4), z2 < (x - 4), it is determined that the texture structure of the prediction template is a left oblique texture structure, and it is determined that the prediction template is divided in a left oblique cutting manner; the calculation formula for the predicted value of the pixel to be predicted is: When |x - y| < 4, x < (z1 - 4), x < (z2 - 4), y < (z1 - 4), y < (z2 - 4), it is determined that the texture structure of the prediction template is a cross texture structure, and it is determined that the prediction template is divided in a cross cutting manner; the calculation formula for the predicted value of the pixel to be predicted is: When |z1 - z2| < 4, z1 < (x - 4), z1 < (y - 4), z2 < (x - 4), z2 < (y - 4), it is determined that the texture structure of the prediction template is a cross texture structure, and it is determined that the prediction template is divided in a cross-cutting manner; the calculation formula for the predicted value of the pixel to be predicted is: When |x - y| < 4, |x - z1| < 4, |x - z2| < 4, |y - z1| < 4, |y - z2| < 4, |z1 - z2| < 4, it is determined that the texture structure of the prediction template is a flat texture structure, and it is determined that the prediction template does not need to be divided and cut.
2. The low-latency ultra-high-definition video encoding method according to claim 1, characterized in that, The step S40 includes: Step S41: When receiving the data stream of the residual coefficient and the original value of the pixel to be predicted, judge the residual coefficient to determine whether the residual coefficient is less than the first threshold and greater than the second threshold. If it holds, execute step S42; otherwise, execute step S43; Step S42: Truncate and encode the residual coefficient to generate an output of the first compressed code, and then jump to execute step S44; Step S43: Encode the original value of the pixel to be predicted to generate an output of the first compressed code; Step S44: Repeat steps S41 to S43 until the data stream of the residual coefficient and the original value of the pixel to be predicted stops being input; The step S60 includes: Step S61: Initialize, restore the coding dictionary to its initial state. At this time, the dictionary only includes all default items, and clear the first character, the first string, and the second string; Step S62: Receive the second compressed code data stream, and read the data from the second compressed code data stream as the first character; Step S63: Combine the first character and the first string to form the second string; Step S64: Check whether the second string exists in the coding dictionary. If it exists, execute step S65; otherwise, execute step S66; Step S65: Determine that the storage address of the second string corresponding to the coding dictionary matches successfully, output the coding dictionary address mark corresponding to the second string, generate the third compressed code data, and update the first string, and then execute step S67; Step S66: Determine that the second string does not match in the coding dictionary storage. First, check whether the coding dictionary has reached the storage limit. If so, output the mark of the first string to generate the third compressed code data and then update the first string; if the storage limit has not been reached, first store the second character in the coding dictionary and mark it, then output the mark of the first string to generate the third compressed code data and update the first string; Step S67: Check whether the first string has reached the splicing limit. If it has reached the limit, clear the first string; if not, there is no need to update the first string; Step S68: Repeat steps S61 to S67 until the output of the second compressed code data is completed.
3. A low-latency ultra-high-definition video decoding method, characterized in that, It is characterized by including the following steps: The video stream data to be decoded is the video stream data obtained by the encoding method as described in claim 1. Step S70: Obtain the video stream data to be decoded; Step S80: Perform dictionary decoding on the video data bitstream to be decoded to generate a first decoded data bitstream; Step S90: Perform data address inverse mapping on the first decoded data bitstream to generate a second decoded data bitstream; Step S100: Perform structure decoding operation on the second decoded data bitstream to obtain the original pixel values or residual coefficients; Step S110: Perform restoration based on the original pixel values or residual coefficients and in combination with the template feature coefficients to finally obtain the original data bitstream.
4. The low-latency ultra-high-definition video decoding method according to claim 3, characterized in that, The said Step S80 includes: Step S81: Initialization, restore the decoding dictionary to its initial state. At this time, the decoding dictionary only includes all default items, and clear the third character, the third string, the fourth string, and the fifth string; Step S82: Receive the video data bitstream to be decoded, and read data from the video data bitstream to be decoded as the first token; Step S83: Check whether there is a corresponding string or character in the decoding dictionary for the first token. If so, execute Step S84; otherwise, execute Step S85; Step S84: Determine that there is a corresponding character or string in the decoding dictionary for the first token, read out the character or string stored in the decoding dictionary, update it to the third string and output it to generate a first decoded data bitstream. Then, check whether the fourth string has reached the splicing upper limit. If it has reached the upper limit, only update the fourth string without updating the decoding dictionary. If it has not reached the upper limit, splice the first character of the third string and the fourth string to form the fifth string, store the fifth string in the decoding dictionary to add a new token mapping; jump to execute Step S86; Step S85: Determine that there is no corresponding character or string in the decoding dictionary for the first token, check whether the fourth string has reached the splicing upper limit. If it has reached the upper limit, only update the fourth string without updating the decoding dictionary. If it has not reached the upper limit, splice the fourth string and its first character to form the fifth string, store the fifth string in the decoding dictionary to add a new token mapping, and its new token is the first token; update the third string, and the updated value is the fifth string, then output the fifth string to generate a first decoded data bitstream; Step S86: Update the fourth string, and update it to the value of the third string; Step S87: Repeat Steps S81 to S86 until the input of the video data bitstream to be decoded stops.
5. A video encoding device, characterized in that, The said video encoding device is applied to the encoding method described in Claim 1. The said video encoding device includes: An encoding control signal synchronization module, which is used to obtain the control signal of the video of the data to be encoded and synchronously output the control signal when outputting the compressed encoded data; A video data encoding and processing module, which is used to perform encoding processing operations on the video data; Among them, the said video data encoding and processing module includes: A data acquisition and caching sub-module, which is used to receive the video data stream to be compressed and perform data caching operations; A predictor sub-module, which is used to calculate the predicted value of the pixel to be compressed; A data compression and reconstruction sub-module, which is used to generate first compressed coding data based on the predicted value of the pixel to be compressed and the original value of the compressed pixel, and then generate second compressed coding data according to the first compressed coding data; A data caching sub-module, which is used to receive the second compressed coding data output by the data compression and reconstruction sub-module, perform caching control and output the second compressed coding data; A dictionary compression sub-module, which is used to process the second compressed coding data, dynamically establish a content search value memory for implementing string search and matching functions, and finally generate third compressed coding data; An output control sub-module, which is used to receive the third compressed coding data and perform caching processing, and logically control the output of the complete third compressed coding data.
6. A video decoding device, characterized in that, The video decoding device is applied to the decoding method described in claim 3, and the video decoding device includes: A decoding control signal synchronization module, which is used to obtain the control signal of the video of the data to be decoded and synchronize the decoded output control signal when outputting the decoded data stream; A video data decoding and processing module, which is used to perform decoding processing operations on the compressed coding data; Among them, the video data decoding and processing module includes: A decoded data caching control sub-module, which is used to receive the data stream of the video data to be decoded and perform caching operations; A data dictionary decoding sub-module, which is used to perform dictionary decoding operations on the data stream of the video data to be decoded to obtain a first decoded data stream; A data decoding and restoration sub-module, which is used to process the first decoded data stream, first perform data inverse mapping to obtain a second decoded data stream, and then perform structural decoding operations on it to finally obtain the decoded data stream; A decoded data output control sub-module, which is used to receive the decoded data stream for caching processing, and logically control the output of the complete decoded data stream.
7. A video encoding and transmitting device, characterized in that, The video encoding and sending device includes: a processor, a memory, a receiver, and a transmitter. The memory stores program codes, and the processor executes the program codes to execute a low-latency ultra-high-definition video encoding method according to any one of claims 1-2.
8. A video decoding and receiving device, characterized in that, The video decoding and receiving device includes: a processor, a memory, a receiver, and a transmitter. The memory stores program codes, and the processor executes the program codes to execute a low-latency ultra-high-definition video decoding method according to any one of claims 3-4.
Citation Information
Patent Citations
Hardware realization system and method of video display stream compression coding
CN109618157A
Video encoding method and device
CN111107345A