Image decoding and encoding method, apparatus, device, and storage medium
By employing a flexible polygon scanning region division method in image encoding, the problem of redundant coefficient encoding caused by insufficient flexibility in scanning region division is solved, thereby improving image compression performance and encoding efficiency.
Patent Information
- Application Number
- CN202311262593.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-09-27
AI Technical Summary
In existing technologies, the division of scanning regions during image encoding is not flexible enough, resulting in redundant coefficient encoding and unsatisfactory compression performance.
A flexible scanning region segmentation method is adopted. By extracting the region syntax parameters of the transform block to be decoded from the image bitstream, the scanning region is determined to be a polygonal region that corresponds to the non-zero data in the quantization residual data, including points and/or polygons with a side length greater than or equal to three, and decoding is performed according to the region scanning order.
It improves image compression performance, reduces encoding bit cost, and achieves a more compact representation of the scan area.
Smart Images

Figure CN119728974B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to an image decoding and encoding method, device, equipment and storage medium. BACKGROUND
[0002] In the process of encoding an image (such as a video), a prediction process is needed to remove spatial and temporal redundancies. The encoder obtains a prediction value through prediction, and the original pixel is subtracted from the prediction value to obtain a residual error. The residual error is processed by transformation and quantization to obtain quantized residual data (also referred to as a coefficient). The quantized residual data and other information (such as block division information, mode information, transformation parameters, quantization parameters, etc.) are then entropy encoded to obtain a code stream.
[0003] In the process of encoding the coefficient, the scanning area is generally used to identify the range of the coefficient to be encoded. The current scanning area is a rectangular area, and this method is not optimal in expressing the scanning area. For example, when the lower right corner of the rectangle is all zeros, the encoding of the redundant zero coefficients is increased. SUMMARY
[0004] The main purpose of the present application is to provide an image decoding and encoding method, device, equipment and storage medium, which aims to solve the technical problems of the prior art that the division of the scanning area is not flexible enough, the encoding of the redundant coefficients is included, and the compression performance is not ideal.
[0005] To achieve the above-mentioned purpose, the present application provides an image decoding method, which comprises the following steps:
[0006] extracting a to-be-decoded transform block from an image code stream, and obtaining region syntax parameters corresponding to the to-be-decoded transform block, the region syntax parameters comprising a horizontal parameter, a vertical parameter and a diagonal parameter;
[0007] determining a scanning area in the to-be-decoded transform block according to the region syntax parameters, the scanning area being an area adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block, and the scanning area comprising a point and / or a polygon with a side length greater than or equal to three;
[0008] sequentially decoding the scanning area according to a region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data;
[0009] constructing a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data.
[0010] In a possible implementation of the present application, the horizontal parameter is the horizontal coordinate of the rightmost non-zero quantized residual data in the scanning area, or the length of the horizontal straight edge.
[0011] The vertical parameter is a longitudinal coordinate of the lowest non-zero quantized residual data in the scanning region, or a length of a straight side in the vertical direction;
[0012] The hypotenuse parameter is a sum of a horizontal coordinate and a longitudinal coordinate of the lowest right non-zero quantized residual data in the scanning region except the point determined according to the horizontal parameter and the vertical parameter, or a horizontal coordinate of an intersection point of an extension line of a hypotenuse of the region shape and an upper boundary of the to-be-decoded transform block.
[0013] In a possible implementation of the present application, the polygon includes at least one hypotenuse.
[0014] In a possible implementation of the present application, the scanning region includes a pentagon with a 45-degree hypotenuse and a right lower corner point of a circumscribed rectangle of the pentagon.
[0015] In a possible implementation of the present application, the hypotenuse parameter is in a range of [0, k], and k is a sum of the horizontal parameter and the vertical parameter minus one.
[0016] In a possible implementation of the present application, the obtaining of the region syntax parameter corresponding to the to-be-decoded transform block includes:
[0017] extracting the horizontal parameter and the vertical parameter from an image code stream;
[0018] determining diagonal information according to the image code stream, the diagonal information being an original value of a hypotenuse parameter, or a difference between the hypotenuse parameter and a preset function value, the preset function value being set based on the horizontal parameter and / or the vertical parameter;
[0019] determining the hypotenuse parameter according to the diagonal information.
[0020] In a possible implementation of the present application, the determining of the diagonal information according to the image code stream includes:
[0021] extracting the diagonal information from the image code stream according to a diagonal coding mode through a diagonal context model;
[0022] or,
[0023] obtaining a target condition in a parsing condition set, the parsing condition set including at least one first condition, the first condition being set according to at least one of the horizontal parameter, the vertical parameter, a width of the to-be-decoded transform block, and a height of the to-be-decoded transform block;
[0024] obtaining the diagonal information according to an information extraction mode corresponding to the target condition.
[0025] In a possible implementation of the present application, the obtaining of the diagonal information according to the information extraction mode corresponding to the target condition includes:
[0026] if the information extraction mode corresponding to the target condition is a direct extraction type, extracting diagonal information from the image code stream according to a diagonal context model and a diagonal coding mode;
[0027] or,
[0028] if the information extraction mode corresponding to the target condition is a conversion acquisition type, searching for a diagonal conversion expression corresponding to the information extraction mode;
[0029] determining diagonal information based on the diagonal conversion expression and the horizontal parameter and / or the vertical parameter.
[0030] In a possible implementation of the present application, the diagonal context model is a context model related to diagonal information.
[0031] The diagonal context model is set according to indication information and a preset threshold.
[0032] The indication information includes at least one of channel information, transform block information, a quantization parameter, a prediction mode, a horizontal parameter and a vertical parameter of the to-be-decoded transform block, and the transform block information includes at least one of a transform type, a transform block width, a transform block height and an area of the transform block.
[0033] In a possible implementation of the present application, the diagonal coding mode is a coding mode when diagonal information is coded.
[0034] The diagonal coding mode is a truncated unary code.
[0035] In a possible implementation of the present application, the diagonal information is used as a hypotenuse parameter.
[0036] The diagonal information is used as a hypotenuse parameter.
[0037] or,
[0038] The diagonal information is parsed to extract a diagonal difference value.
[0039] A hypotenuse parameter is calculated according to the diagonal difference value, the horizontal parameter and / or the vertical parameter according to a difference expression.
[0040] In a possible implementation of the present application, the region scanning order corresponding to the to-be-decoded transform block includes at least one of horizontal scanning, vertical scanning, zigzag scanning and reverse zigzag scanning.
[0041] The scanning order is a preset order or is acquired from an image code stream.
[0042] In a possible implementation of the present application, the sequentially decoding the scan region according to the region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data comprises:
[0043] obtaining a context model corresponding to the scan region, wherein the context model corresponding to the scan region comprises a context model used when extracting a binarization component of each quantized residual data in the scan region;
[0044] sequentially decoding the scan region according to the region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data based on the context model.
[0045] In a possible implementation of the present application, the context model corresponding to the scan region is related to at least one of the horizontal parameter, the vertical parameter, the diagonal parameter, the region shape of the scan region, the region area of the scan region, and the position of the quantized residual data in the scan region.
[0046] In a possible implementation of the present application, the obtaining the context model corresponding to the scan region comprises:
[0047] obtaining an area of a polygon in the scan region;
[0048] determining the context model corresponding to the scan region according to the area of the polygon, the horizontal parameter, and the vertical parameter.
[0049] In a possible implementation of the present application, the sequentially decoding the scan region according to the region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data comprises:
[0050] when decoding a zero-value flag component, selecting to-be-decoded residual data from the scan region according to the region scanning order corresponding to the to-be-decoded transform block;
[0051] if the to-be-decoded residual data is on a target edge of the scan region, detecting whether the to-be-decoded residual data is the last undecoded residual data on the target edge;
[0052] if yes, detecting whether all other zero-value flag components on the target edge except the to-be-decoded residual data are 0;
[0053] if all are 0, setting the zero-value flag component of the to-be-decoded residual data to 1;
[0054] if not all are 0, extracting the zero-value flag component of the to-be-decoded residual data from an image code stream according to a context model of the to-be-decoded residual data.
[0055] Furthermore, to achieve the above object, the present application also provides an image coding method, which comprises:
[0056] transforming a to-be-coded transform block corresponding to a target image to obtain quantized residual data corresponding to the to-be-coded transform block;
[0057] marking a scanning region in the to-be-coded transform block according to non-zero data in the quantized residual data, wherein the scanning region comprises a point and / or a polygon with a side length greater than or equal to three;
[0058] encoding quantized residual data corresponding to the scanning region in sequence according to a region scanning order corresponding to the to-be-coded transform block to generate an image code stream corresponding to the target image;
[0059] constructing a region syntax parameter corresponding to the to-be-coded transform block according to the scanning region and writing the region syntax parameter into the image code stream.
[0060] Furthermore, to achieve the above object, the present application also provides an image decoding device, which comprises:
[0061] a decoding module, configured to extract a to-be-decoded transform block from an image code stream and acquire a region syntax parameter corresponding to the to-be-decoded transform block, wherein the region syntax parameter comprises a horizontal parameter, a vertical parameter and a diagonal parameter;
[0062] a constructing module, configured to determine a scanning region in the to-be-decoded transform block according to the region syntax parameter, wherein the scanning region is a region adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block, and the scanning region comprises a point and / or a polygon with a side length greater than or equal to three;
[0063] a scanning module, configured to decode the scanning region in sequence according to a region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data;
[0064] a reconstructing module, configured to construct a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data.
[0065] Furthermore, to achieve the above object, the present application also provides an image coding device, which comprises:
[0066] a transforming module, configured to transform a to-be-coded transform block corresponding to a target image to obtain quantized residual data corresponding to the to-be-coded transform block;
[0067] a marking module, configured to mark a scanning region in the to-be-coded transform block according to non-zero data in the quantized residual data, wherein the scanning region comprises a point and / or a polygon with a side length greater than or equal to three;
[0068] generating module, configured to sequentially encode quantized residual data corresponding to the scanning areas according to the region scanning order corresponding to the to-be-encoded transform block, and generate an image code stream corresponding to the target image;
[0069] The encoding module is configured to construct region syntax parameters corresponding to the to-be-encoded transform block according to the scanning areas, and write the region syntax parameters into the image code stream.
[0070] In addition, to achieve the above object, the present application further provides a decoding device, which comprises a processor, a memory, and an image decoding program stored in the memory and executable on the processor, and the image decoding program is executed by the processor to implement the image decoding method as described above.
[0071] In addition, to achieve the above object, the present application further provides an encoding device, which comprises a processor, a memory, and an image encoding program stored in the memory and executable on the processor, and the image encoding program is executed by the processor to implement the image encoding method as described above.
[0072] In addition, to achieve the above object, the present application further provides a storage medium, which stores an image decoding program and / or an image encoding program, the image decoding program is executed to implement the image decoding method as described above, and the image encoding program is executed to implement the image encoding method as described above.
[0073] The present application extracts a to-be-decoded transform block from an image code stream, and obtains region syntax parameters corresponding to the to-be-decoded transform block; determines scanning areas in the to-be-decoded transform block according to the region syntax parameters, wherein the scanning areas are regions adapted to non-zero quantized residual data in the to-be-decoded transform block; sequentially decodes the scanning areas according to a region scanning order corresponding to the to-be-decoded transform block, and obtains quantized residual data; and constructs a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data. Since the scanning areas can be regions including points and / or polygons with a side length greater than or equal to three, a more flexible and compact scanning area expression mode can be selected according to the distribution of non-zero quantized residual data when compressing an image, thereby reducing the bit cost of image encoding and improving the compression performance. BRIEF DESCRIPTION OF DRAWINGS
[0074] Figure 1 is a structural schematic diagram of an electronic device of a hardware running environment related to an embodiment scheme of the present application;
[0075] Figure 2 is a flowchart of a first embodiment of the image decoding method of the present application;
[0076] Figure 3A video encoding framework structure diagram of an embodiment of the present application;
[0077] Figure 4 A scanning region division diagram of an embodiment of the present application;
[0078] Figure 5 An encoding skip diagram of an embodiment of the present application;
[0079] Figure 6 A flow diagram of a second embodiment of the image decoding method of the present application;
[0080] Figure 7 A standard code text example diagram of an embodiment of the present application;
[0081] Figure 8 A flow diagram of a first embodiment of the image encoding method of the present application;
[0082] Figure 9 A structure block diagram of a first embodiment of the image decoding device of the present application;
[0083] Figure 10 A structure block diagram of a first embodiment of the image encoding device of the present application.
[0084] The implementation, functional features and advantages of the present application will be further explained with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0085] It should be understood that the specific embodiments described herein merely exemplify the present application and are not intended to limit the present application.
[0086] Reference Figure 1 , Figure 1 A decoding device or encoding device structure diagram of a hardware running environment involved in the embodiment scheme of the present application.
[0087] As Figure 1As shown, the electronic device can include a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to realize the connection and communication between the components. The user interface 1003 can include a display, an input unit such as a keyboard, and can also include a standard wired interface, a wireless interface. The network interface 1004 can optionally include a standard wired interface, a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 can be a high-speed random access memory (RAM), and can also be a stable non-volatile memory (NVM), such as a disk memory. The memory 1005 can also be a storage device independent of the aforementioned processor 1001.
[0088] Those skilled in the art can understand that Figure 1 The structure shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than the figure, or combine certain components, or different component arrangements.
[0089] As Figure 1 As shown, the memory 1005 as a storage medium can include an operating system, a network communication module, a user interface module, and an image decoding program and / or an image encoding program.
[0090] In Figure 1 As shown in the electronic device, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present application can be arranged in a decoding device or an encoding device, the electronic device calls the image decoding program stored in the memory 1005 through the processor 1001, and executes the image decoding method provided by the embodiment of the present application; the electronic device calls the image encoding program stored in the memory 1005 through the processor 1001, and executes the image encoding method provided by the embodiment of the present application.
[0091] The embodiment of the present application provides an image decoding method, which refers to Figure 2 , Figure 2 The flowchart of the first embodiment of the image decoding method of the present application.
[0092] In this embodiment, the image decoding method includes the following steps:
[0093] Step S10: Extract the transform block to be decoded from the image bitstream and obtain the region syntax parameters corresponding to the transform block to be decoded.
[0094] It should be noted that the execution subject of this embodiment can be a decoding device when encoding image data. The decoding device can be a personal computer, server or other electronic device. Of course, it can also be other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the image decoding method of the present invention is described using a decoding device as an example.
[0095] Since the encoding device typically decodes the encoded image stream after encoding is completed during the image encoding process, and determines whether the parameters used in the encoding need to be adjusted based on the image quality of the decoded image, the execution subject in this embodiment can also be the encoding device.
[0096] It should be noted that an image bitstream can be a bitstream generated by an encoding device after encoding image data that needs to be compressed. When compressing image data, the image data is generally divided into at least one image transform block (TB), and then the divided image transform blocks are encoded / decoded sequentially. The transform block to be decoded can be the image transform block currently being decoded extracted from the image bitstream.
[0097] Among them, the region syntax parameters can be parameters used to indicate the region extent of the scanned region, including horizontal parameters, vertical parameters, and diagonal parameters;
[0098] The horizontal parameter is the x-coordinate of the rightmost non-zero quantized residual data in the scan area, or the length of the right-angled side in the horizontal direction; the vertical parameter is the y-coordinate of the bottommost non-zero quantized residual data in the scan area, or the length of the right-angled side in the vertical direction; the hypotenuse parameter includes the sum of the x-coordinate and y-coordinate of the bottom rightmost non-zero quantized residual data in the scan area (excluding the points determined by the horizontal and vertical parameters, or the x-coordinate of the intersection point of the extension of the hypotenuse of the polygon in the scan area and the upper boundary of the transform block to be decoded), and the hypotenuse angle value, which is the angle between the hypotenuse of the polygon area and the upper boundary of the transform block to be decoded.
[0099] Figure 3 This is a schematic diagram of the video coding framework structure in this embodiment. The operational flow of a commonly used video coding framework during image encoding or decoding is as follows: Figure 3 As shown, it mainly involves modules such as prediction, transformation, quantization, entropy coding, and filtering, among which:
[0100] The prediction module includes intra-frame prediction (e.g.) Figure 3 Intra-prediction units (IMUs) and inter-prediction units (IPUs)Figure 3 Intra prediction is to predict using the reconstructed pixels around the current block to remove spatial redundancy; Inter prediction is to predict using the reconstructed pixels on the temporal reference frame to remove temporal redundancy.
[0101] The transform module is to linearly map the spatial residual information to the transform domain (e.g. frequency domain), aiming to concentrate the energy and remove the frequency domain correlation of the signal. The theoretical transform matrix is reversible and does not cause signal loss.
[0102] The quantization module is a "many-to-one" mapping process, which is irreversible and causes signal loss; the advantage is that it can greatly reduce the value range of the signal, so that the encoder can give a good approximation of the original signal with a small number of symbols, thereby improving the compression rate. The inverse quantization module and the inverse transform module perform the inverse processes of the quantization module and the transform module, respectively.
[0103] The entropy coding module (e.g. the entropy encoder shown in Figure 3 ) is a lossless coding method based on the information entropy principle, which converts a series of element symbols (e.g. transform coefficients and mode information, etc.) used to represent the video sequence into a binary code stream, removing the statistical redundancy of these video element symbols.
[0104] The filter module (e.g. the in-loop filter module shown in Figure 3 ) will enhance the reconstructed image, aiming to make the reconstructed image closer to the original image, while reducing the impact of blocking effect and ringing effect, and improving the quality of the reconstructed image.
[0105] During the encoding process, the encoding device will try to perform image reconstruction, and after the reconstructed image is processed by the in-loop filter, it will pass through the reference image buffer, perform motion estimation / motion compensation, and determine the inter mode information according to the motion estimation / motion compensation. Finally, after determining that the performance of the encoding meets the standard, the entropy encoder is used for entropy coding to generate an image code stream (i.e. bit stream).
[0106] In this process, some technical terms may be involved, including residual, coefficient coding, scan region-based coefficient coding (SRCC), entropy coding, etc., which will be described as follows:
[0107] Residual: In the process of image (e.g. video) coding, the spatial and temporal redundancy is removed by prediction. The encoder gets the predicted value by prediction, and the residual is obtained by subtracting the predicted value from the original pixel. The residual is transformed and quantized to get the coefficients, and the coefficients and other information (including block partition information, mode information, transform parameters, quantization parameters, etc.) are entropy coded to get the bitstream.
[0108] Coefficient Coding: The quantized coefficients are coded, including determining the range, order and method of coding the coefficients, where the coding method includes binarization method, context model and entropy coding. The range of the coded coefficients refers to only the coefficients in this range are coded, and the coefficients outside the range are not coded (default is zero). The order of coding refers to all the coefficients to be coded are arranged into a one-dimensional array in this order, and then coded one by one. The binarization method refers to each coefficient is represented as a binary symbol string, and then each bit (0 / 1 binary number, called bin) of the binary symbol string is written into the bitstream by the entropy encoder. The context model estimates the probability of each bin taking value 0 or 1 to assist the entropy encoder to better compress the binary symbol string.
[0109] Scan Region-based Coefficient Coding (SRCC): A scan region is used to represent the range of coefficients to be coded. The scan region is a rectangular region, which is expressed by coordinates (sr_x, sr_y), where sr_x is the horizontal coordinate of the rightmost non-zero coefficient in the transform block (TB), and sr_y is the vertical coordinate of the lowermost non-zero coefficient in the transform block. In the coefficient coding process, only the coefficients inside and on the boundary of the scan region are coded, and the coefficients outside the scan region are not coded (default is zero). The coding order (also called scan order) uses reverse zigzag scanning.
[0110] Entropy Coding: The commonly used entropy coding method is Context-based Adaptive Binary Arithmetic Coding (CABAC), which is to arithmetic code each bin of the binary symbol after binarization according to its context model to get the final output bitstream. In coding, the coding of each symbol is related to the results of previous coding, and the code word is assigned to each symbol according to the statistical characteristics of the bitstream. CABAC is especially suitable for symbols with non-equal probability, which can remove the correlation between symbols and further compress the code rate.
[0111] Step S20: determining a scanning region in the to-be-decoded transform block according to the region syntax parameter.
[0112] It should be noted that the scanning region is a region adapted to the non-zero data in the quantized residual data corresponding to the to-be-decoded transform block, and the scanning region includes a point and / or a polygon with a side length greater than or equal to three, and the polygon includes at least one oblique side, wherein the polygon used can be a pentagon.
[0113] It should be noted that the scanning region can be composed of a point and / or a polygon region combined according to the region syntax parameter.
[0114] In a possible implementation of the embodiment, in order to ensure that the region can be accurately divided and avoid the phenomenon that the points on the oblique side of the region need to be rounded, the oblique side angle value can be set to 45°, and at this time, the scanning region can be a region including a pentagon with a 45° oblique side and a point at the right lower corner of the circumscribed rectangle of the pentagon.
[0115] In actual use, in order to facilitate the division of the scanning region, a coordinate system can be constructed in the to-be-decoded transform block in advance, for example, taking the upper left corner of the to-be-decoded transform block as the coordinate system origin, taking the upper boundary of the to-be-decoded transform block to the right as the positive direction of the x-axis of the coordinate system, and taking the left boundary of the to-be-decoded transform block to the down as the positive direction of the y-axis of the coordinate system, and the unit length of the coordinate system can be the side length of a 1x1 sub-block.
[0116] Then, the intersection of the polygon and the x-axis of the coordinate system is determined according to the horizontal parameter, the intersection of the polygon and the y-axis of the coordinate system is determined according to the vertical parameter, and the oblique side of the polygon is marked in the coordinate system according to the oblique side parameter, then the intersection and the oblique side are connected to obtain the polygon region, and then the polygon region and the point constructed according to the horizontal parameter and the vertical parameter are combined to divide the scanning region.
[0117] In order to facilitate understanding, the following will be described in combination with Figure 4 , but the scheme is not limited thereto. Figure 4 A scanning region division schematic diagram of the embodiment is shown in Figure 4 , the oblique side angle value is 45°, the obtained horizontal parameter sr_x=4, the vertical parameter sr_y=6, and the diagonal sr_d=7, and at this time, the intersection of the pentagon and the x-axis is (4, 0) according to the horizontal parameter, the intersection of the pentagon and the y-axis is (0, 6) according to the vertical parameter, and the end points of the oblique side of the pentagon are (4, 3) and (1, 6) according to sr_x, sr_y, sr_d and the oblique side angle value, and then the coordinate origin (0, 0) of the upper left corner is added, and then the five points are connected to obtain the pentagon region, and then the complete scanning region can be obtained by combining the pentagon region and the point Y(4, 6) determined according to the horizontal parameter sr_x and the vertical parameter sr_y.
[0118] In a specific implementation, the value range of the oblique side parameter is [0, k], k is the sum of the horizontal parameter and the vertical parameter minus one, in the case of using a pentagon as the polygon, the value range of the horizontal parameter sr x is [0, min(32, tb w-1)], the value range of the vertical parameter sr y is [0, min(32, tb h-1)], and the value range of the oblique side parameter sr d is [0, sr x+sr y-1], where tb w and tb h are respectively the width and the height of the to-be-decoded transform block.
[0119] For a given sr x and sr y, the scanning region changes when sr d takes different values.
[0120] For example, if sr d==0, the pentagon degenerates into a point, and the scanning region is two points (one point degenerated from the pentagon, and one point determined according to sr x and sr y);
[0121] If 0<sr d<=min(sr x, sr y), the pentagon degenerates into an isosceles right triangle, and the scanning region is “isosceles right triangle+point determined according to sr x and sr y”;
[0122] If min(sr x, sr y)<sr d<=max(sr x, sr y), the pentagon degenerates into a right trapezoid, and the scanning region is “right trapezoid+point determined according to sr x and sr y”;
[0123] If max(sr x, sr y)<sr d<sr x+sr y-1, the scanning region is “pentagon+point determined according to sr x and sr y”;
[0124] If sr d==sr x+sr y-1, the scanning region is “pentagon+point determined according to sr x and sr y”, and “pentagon+point” at this time will be just spliced into a rectangle.
[0125] Step S30: sequentially decoding the scanning region according to the region scanning order corresponding to the to-be-decoded transform block, to obtain quantized residual data.
[0126] It should be noted that the region scanning order can include at least one of horizontal scanning, vertical scanning, zigzag scanning and reverse zigzag scanning.
[0127] In a specific implementation, the management personnel of the encoding device and / or the decoding device can set the plurality of to-be-decoded transform blocks in the image code stream to the same region scanning order, and at this time, the preset order set by the management personnel is taken as the region scanning order corresponding to the to-be-decoded transform block.
[0128] In actual use, when encoding image data, the corresponding region scanning order can be set for each transform block into which the image data is split, and the set region scanning order is written into the image code stream, and at this time, the region scanning order corresponding to the to-be-decoded transform block can be extracted from the image code stream through syntax. The syntax used here can be syntax at the sequence level, frame level, patch level, slice level, tile level, coding tree unit (CTU) level, luma coding tree block (CTB) level, coding unit (CU) level, coding block (CB) level, or transform block (TB) level.
[0129] In actual use, when the quantized residual data is written into the image code stream, the quantized residual data is generally binarized, and then the binarized components are written into the image code stream. The binarization formula for each quantized residual data (i.e., the coefficient obtained after transform and quantization) is:
[0130] coef = sign * (sig + gt1 + gt2 + remain)
[0131] In the formula, coef is the quantized residual data, sig indicates whether the quantized residual data is non-zero, gt1 indicates whether the absolute value of the quantized residual data is greater than 1, gt2 indicates whether the absolute value of the quantized residual data is greater than 2, remain indicates the absolute value of the quantized residual data minus 3 (usually using exponential Golomb for binarization), and sign indicates the sign bit of the quantized residual data. Among them, sig, gt1, and gt2 usually use context-based encoding, and remain and sign usually use bypass encoding.
[0132] At this time, the scanning regions are sequentially decoded according to the region scanning order corresponding to the to-be-decoded transform block to obtain the quantized residual data, which can be to obtain the context model corresponding to the scanning region, to decode the image code stream according to the context model, to sequentially obtain the binarized components of each quantized residual data in the scanning region according to the region scanning order corresponding to the to-be-decoded transform block, and then to calculate the corresponding quantized residual data according to the above binarization formula.
[0133] The context model corresponding to the scanning region can include all context models required for extracting the binarization components of each quantized residual data in the scanning region from the image code stream. For example, assuming that the scanning region contains four quantized residual data, A, B, C, and D, and that the sig, gt1, and gt2 are usually encoded based on the context, the context model is required for parsing from the image code stream, and three context models are required for each quantized residual data. Therefore, the context model corresponding to the scanning region contains a total of 12 context models.
[0134] In a specific implementation, the context model corresponding to the scanning region can be related to at least one of the horizontal parameter, the vertical parameter, the diagonal parameter, the shape of the scanning region, the area of the scanning region, and the position of the quantized residual data in the scanning region.
[0135] For example, the context model of sig, gt1, and gt2 can be designed based on at least one of the following:
[0136] (1) sr_x, sr_y, sr_d, etc. For example, the relationship between the coefficient position (x, y) and sr_d indicates that the coefficient falls on the diagonal of the pentagon, or the upper left, or the lower right.
[0137] (2) The area of the pentagonal region, or the area of the scanning region, or the area of the circumscribed rectangle of the pentagon.
[0138] (3) The processed “area”. The area can be any of the area of the pentagonal region, the area of the scanning region, or the area of the circumscribed rectangle of the pentagon. The processing method can be to find the maximum value max, the minimum value min, or clipping clip. For example, if the pentagon is not degenerated, the area of the pentagonal region is used; otherwise, the area of the largest trapezoid of the degenerated pentagon is used.
[0139] Therefore, the standard code text of the context setting can be as follows:
[0140] scanArea=(ScanRegionX+1)*(ScanRegionY+1)
[0141] deltaSrd=ScanRegionX+ScanRegionY–Max(ScanRegionD,Max(ScanRegionX,ScanRegionY))
[0142] scanArea=scanArea–((deltaSrd*(deltaSrd+1))>>1)
[0143] ctxIndexInc = ctxIndexInc + scanArea <= 4? 0 : 17 « (scanArea <= 16? 0 : 1)
[0144] In the standard code text, the scanArea is the area of the processed pentagon region, and the processing method is: if the pentagon is not degenerated, the area of the pentagon region is used; otherwise, the area of the largest trapezoid degenerated from the pentagon is used, the deltaSrd is the diagonal information, the ScanRegionX is the horizontal parameter, the ScanRegionY is the vertical parameter, and the ctxIndexInc is the context.
[0145] In actual use, since only the scanning region in the to-be-decoded transform block is scanned, the obtained quantized residual data at this time only contains the quantized residual data corresponding to the scanning region in the to-be-decoded transform block, and the quantized residual data outside the scanning region of the to-be-decoded transform block is zero by default (of course, it can also be other similar values, which is not limited here), therefore, after obtaining the quantized residual data corresponding to the scanning region in the to-be-decoded transform block, the zero is used to complete, that is, the quantized residual data corresponding to the to-be-decoded transform block can be obtained.
[0146] Step S40: constructing a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data.
[0147] It should be noted that constructing a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data can be dequantizing the quantized residual data to obtain a reconstructed image feature, de-transforming the reconstructed image feature to obtain a reconstructed residual value, adding the reconstructed residual value to a predicted value obtained according to a prediction mode to obtain a reconstructed pixel value, and constructing a reconstructed image block corresponding to the to-be-decoded transform block.
[0148] In actual use, if the to-be-decoded transform block is multiple, the reconstructed image blocks corresponding to each to-be-decoded transform block in the image code stream can be obtained, and the obtained reconstructed image blocks can be aggregated to obtain a complete reconstructed image.
[0149] It should be noted that in order to reduce the parsing and / or decoding steps of the quantized residual data, the management personnel of the encoding device or the decoding device can pre-set the encoding skip rule according to actual needs.
[0150] For example, the coding skip rule is set as: in a given scan order, when quantized residual data on a target edge (a right straight edge, a lower straight edge, a right lower diagonal edge) of a coding scan region is coded, if all the quantized residual data on the target edge except the last quantized residual data are 0, then the last quantized residual data on the target edge must be non-zero, or the sig (a flag information indicating whether the current quantized residual data is non-zero) of the quantized residual data must be 1, and the coding of the sig of the quantized residual data can be skipped at this time;
[0151] Similarly, when decoding, if it is detected that all the quantized residual data on the target edge of the scan region are 0, then when the last quantized residual data on the target edge is decoded, the sig of the quantized residual data can be directly set to 1 without parsing the sig from the image code stream through the context model, that is, the parsing of the sig of the quantized residual data is skipped.
[0152] In a specific application, if such a coding skip rule is set, then when the coding skip rule is met, the parsing and / or decoding steps of the quantized residual data can be reduced, and at this time, the step S30 described in the embodiment can include:
[0153] When decoding the zero value flag component, the to-be-decoded residual data is selected from the scan region according to the region scan order corresponding to the to-be-decoded transform block;
[0154] If the to-be-decoded residual data is on a target edge of the scan region, it is detected whether the to-be-decoded residual data is the last undecoded residual data on the target edge;
[0155] If yes, it is detected whether all the zero value flag components on the target edge except the to-be-decoded residual data are 0;
[0156] If all are 0, the zero value flag component of the to-be-decoded residual data is set to 1;
[0157] If not all are 0, the zero value flag component of the to-be-decoded residual data is extracted from the image code stream according to the context model of the to-be-decoded residual data.
[0158] It should be noted that the zero value flag component can be the sig in the above-mentioned binary component, which is used to indicate whether the quantized residual data is non-zero.
[0159] For the convenience of understanding, the coding skip rule will be described below in combination with Figure 5 , but the scheme is not limited thereto. Figure 5 The coding skip rule is shown in the following figure. As shown in the figure, the coding scan region is divided into a plurality of scan regions, and the scan regions are scanned in a scan order. Figure 5As shown, when sig coding is performed, the feature of the scanning region is that the coefficients on the target side (the right straight side, the lower straight side, and the oblique side) are not all 0, and if the reverse zigzag scanning order is used:
[0160] For the coefficients on the lower straight side, the A-point coefficient is scanned last, so if the coefficients other than the A-point on the lower straight side (the scanning order of these coefficients is all before the A-point) are all 0, then the sig of the A-point coefficient must be 1, and the coding and decoding of the sig of the A-point coefficient can be skipped;
[0161] For the coefficients on the right straight side, the B-point coefficient is scanned last, so if the coefficients other than the B-point on the right straight side (the scanning order of these coefficients is all before the B-point) are all 0, then the sig of the B-point coefficient must be 1, and the coding and decoding of the sig of the B-point coefficient can be skipped;
[0162] For the coefficients on the oblique side, according to the feature of the reverse zigzag scanning order and the parity of sr_d, it can be determined whether the last scanned coefficient position on the oblique side is the C-point or the D-point. Assuming that the C-point coefficient is scanned last, if the coefficients other than the C-point on the oblique side (the scanning order of these coefficients is all before the C-point) are all 0, then the sig of the C-point coefficient must be 1, and the coding and decoding of the sig of the C-point coefficient can be skipped. Similarly, the sig skipping mode of the D-point coefficient can be obtained.
[0163] The embodiment extracts a to-be-decoded transform block from an image code stream, and obtains region syntax parameters corresponding to the to-be-decoded transform block; determines a scanning region in the to-be-decoded transform block according to the region syntax parameters, the scanning region being a region adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block; decodes the scanning region in turn according to a region scanning order corresponding to the to-be-decoded transform block, to obtain the quantized residual data; and constructs a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data. Since the scanning region can be a region including a polygon having a point and / or a side length greater than or equal to three, when an image is compressed, a more flexible and compact scanning region expression mode can be selected according to the distribution of non-zero quantized residual data, thereby reducing the bit cost of image coding and improving the compression performance.
[0164] Reference Figure 6 , Figure 6 The flowchart of a second embodiment of the image decoding method of the present application is shown in FIG. 2.
[0165] Based on the first embodiment, the step S10 of the image decoding method of the present embodiment comprises:
[0166] Step S101: Extracting a to-be-decoded transform block from an image code stream, and extracting a horizontal parameter and a vertical parameter from the image code stream.
[0167] It should be noted that when the horizontal parameter and the vertical parameter are encoded into the code stream, the encoding manner adopted can be group encoding, and the implementation manner of the group encoding can be that a code word includes two parts of a prefix and a suffix, the prefix represents a group number, and the suffix represents an offset in the group. Of course, the group encoding can also be implemented in other manners, and the present embodiment does not limit this.
[0168] The horizontal parameter and the vertical parameter extracted from the image code stream can be the horizontal parameter and the vertical parameter corresponding to the to-be-decoded transform block extracted from the image code stream based on a context model through syntax. The syntax used here can be sequence-level, frame-level, patch-level, slice-level, tile-level, CTU-level, CTB-level, CU-level, CB-level or TB-level syntax. The context model can be related to at least one of the width, height, area, color channel, luminance component channel, chrominance component channel, block transform type, sub-block transform type of the to-be-decoded transform block.
[0169] Step S102: determining diagonal information according to the image code stream.
[0170] It should be noted that the diagonal information can be the original value of the diagonal parameter, or the difference between the diagonal parameter and a preset function value (which can be calculated according to sr_x and / or sr_y). The preset function value can be obtained based on the horizontal parameter and / or the vertical parameter.
[0171] For example, the diagonal information can be the original value of the diagonal parameter, or the diagonal information can be sr_d-f(sr_x, sr_y) or f(sr_x, sr_y)-sr_d, where f(sr_x, sr_y) is a function related to sr_x and / or sr_y. The expression of the function can be any of the following expressions:
[0172] f(sr_x, sr_y) = sr_x + sr_y + a1;
[0173] f(sr_x, sr_y) = max(sr_x, sr_y) + a2;
[0174] f(sr_x, sr_y) = min(sr_x, sr_y) + a3;
[0175] f(sr_x, sr_y) = max(a4, sr_x + sr_y + a5);
[0176] f(sr_x, sr_y) = min(a6, sr_x + sr_y + a7);
[0177] f(sr_x, sr_y) = clip(a8, a9, sr_x + sr_y + a10);
[0178] wherein a1-a10 are all preset constants.
[0179] Of course, the expression of the function can also be other types of expressions, for example: other binary linear functions or binary multiple functions constructed according to the horizontal parameter and / or the vertical parameter, which are not limited in the embodiment.
[0180] In a possible implementation of the embodiment, the diagonal information can be directly coded into the image code stream, and therefore the step S102 can include the following steps:
[0181] extracting the diagonal information from the image code stream according to the diagonal coding mode through the diagonal context model.
[0182] It should be noted that the diagonal context model can be a context model related to the diagonal information, which is used to assist in extracting the diagonal information from the image code stream. The diagonal coding mode can be any one of the fixed-length code, the unary code, the truncated unary code, the grouped code, the Rice code, the truncated Rice code, the Golomb code, the exponential Golomb code, and the Huffman code.
[0183] The implementation of the grouped code can be that the code word includes two parts, a prefix and a suffix, the prefix represents a group number, and the suffix represents an offset value in the group.
[0184] The implementation of the fixed-length code can be that the code length of the fixed-length code is calculated according to the value range of the diagonal information, and then the diagonal information is encoded by using the fixed-length code.
[0185] The implementation of the truncated unary code can be that the maximum value of the diagonal information is calculated according to the value range of the diagonal information, and then the diagonal information is encoded by using the truncated unary code.
[0186] Of course, the above examples of the coding mode are only illustrative, and in actual application, similar coding modes can also be implemented in other ways, which are not limited in the embodiment.
[0187] In actual use, the diagonal context model can be set according to the indication information and in combination with a preset threshold, for example: for discrete indication information, according to the value, the indication information can be divided into several categories, and different categories correspond to different context models; for continuous indication information, according to the value and the preset threshold, the indication information can be divided into several intervals, and different intervals correspond to different context models.
[0188] The indication information can include at least one of channel information of the to-be-decoded transform block, transform block information, a quantization parameter, a prediction mode, a horizontal parameter, and a vertical parameter.
[0189] For example, the channel index, whether it is a luminance channel, or whether it is a chrominance channel, is used as the indication information.
[0190] The transform type, the transform block width, the transform block height, and the area of the transform block of the to-be-decoded transform block are used as the indication information.
[0191] The quantization mode or the prediction mode (such as inter-frame prediction or intra-frame prediction) is used as the indication information.
[0192] The combination information of the horizontal parameter sr_x and / or the vertical parameter sr_y is used as the indication information, and the combination information can be the following combination information:
[0193] sr_x+c1;
[0194] sr_y+c2;
[0195] sr_x+sr_y+c3;
[0196] min(sr_x,sr_y)+c4;
[0197] max(sr_x,sr_y)+c5;
[0198] (sr_x+c6)*(sr_y+c7)+c8.
[0199] wherein c1-c8 are all preset constants.
[0200] In order to facilitate understanding, the acquisition of the context is described here, but the present scheme is not limited thereto, for example:
[0201] The context of the diagonal information can be acquired by the following method:
[0202] If the horizontal parameter (ScanRegionX) is less than 4 or the vertical parameter (ScanRegionY) is less than 4, the context model ctxIndexInc is equal to 0;
[0203] Otherwise, the area of the outer rectangle of the scanning region is calculated as follows:
[0204] rectArea=(ScanRegionX+1)*(ScanRegionY+1)
[0205] If the current block is a luminance coding block, then:
[0206] ctxIndexInc=rectArea<64?1:(rectArea<256?2:3)
[0207] If the current block is a chroma-coded block, then:
[0208] ctxIndexInc=rectArea<64?4:5.
[0209] In one possible implementation of this embodiment, the method for obtaining diagonal information can be determined by judging whether the set conditions are met. In this case, step S102 of this embodiment may include:
[0210] Obtain the target conditions that are satisfied in the set of parsing conditions;
[0211] Obtain diagonal information according to the information extraction method corresponding to the target conditions.
[0212] It should be noted that the parsing condition set includes at least one first condition, which is set based on at least one of the following: a horizontal parameter, a vertical parameter, the width of the transform block to be decoded, and the height of the transform block to be decoded. Information extraction can be performed by obtaining diagonal information.
[0213] For example, the first condition can be set to any of the following forms:
[0214] The horizontal parameter sr_x >= b1;
[0215] Vertical parameter sr_y >= b2;
[0216] The width of the block to be decoded, tb_w, is greater than or equal to b3;
[0217] The height of the block to be decoded, tb_h, is greater than or equal to b4;
[0218] Wherein, b1 to b4 are all preset constants. The expression >= can be modified to <=, ==, >, <, != according to actual needs. The first condition can also be a combination of at least two of the above (the combination can be "or" or "and"). This embodiment does not limit this.
[0219] The standard code example for obtaining the hypotenuse parameter at this time can be as follows: Figure 7 As shown, Figure 7 In the diagram, blockWidth and blockHeight are the width and height of the transform block to be decoded, respectively; ScanRegionX is the horizontal parameter; ScanRegionY is the vertical parameter; ScanRegionD is the diagonal parameter; and DeltaScanRegionD is the diagonal information.
[0220] In a specific implementation, the information extraction manner can include two categories of direct extraction and conversion acquisition. In this case, the step of acquiring the diagonal information according to the information extraction manner corresponding to the target condition can include:
[0221] If the information extraction manner corresponding to the target condition is of the direct extraction type, the diagonal information is extracted from the image code stream according to the diagonal coding manner through the diagonal context model.
[0222] Or,
[0223] If the information extraction manner corresponding to the target condition is of the conversion acquisition type, the diagonal conversion expression corresponding to the information extraction manner is searched.
[0224] The diagonal information is determined based on the diagonal conversion expression and the horizontal parameter and / or the vertical parameter.
[0225] It should be noted that if the information extraction manner corresponding to the target condition is of the direct extraction type, it means that the manager of the decoding device or the encoding device directly encodes the diagonal information into the image code stream. Therefore, the diagonal information can be extracted from the image code stream according to the diagonal coding manner through the diagonal context model.
[0226] If the information extraction manner corresponding to the target condition is of the conversion acquisition type, it means that the manager of the decoding device or the encoding device does not encode the diagonal information into the image code stream, but sets to calculate the diagonal information through the horizontal parameter and / or the vertical parameter. The conversion acquisition type can further include multiple different information extraction manners, and each different information extraction manner can correspond to a different diagonal conversion expression. In this case, in order to correctly determine the diagonal information, the diagonal conversion expression corresponding to the information extraction manner is searched, and then the horizontal parameter and / or the vertical parameter is substituted into the diagonal conversion expression, so as to calculate the diagonal information.
[0227] For example, the conversion acquisition type can include four information extraction manners, namely A, B, C, and D, and each information extraction manner can correspond to a different diagonal conversion expression.
[0228] In a specific implementation, the correspondence between the information extraction manner and the diagonal conversion expression can be stored by pre-setting the manner expression information table in the encoding device or the decoding device.
[0229] Step S103: determining the hypotenuse parameter according to the diagonal information.
[0230] In a specific implementation, the diagonal information can be the original value of the hypotenuse parameter. In this case, the hypotenuse parameter can be directly determined according to the diagonal information.
[0231] And if the diagonal information is the difference between the hypotenuse parameter and the preset function value, then according to the diagonal information to determine the hypotenuse parameter can be to analyze the diagonal information, extract the diagonal difference value, and then calculate the hypotenuse parameter according to the diagonal difference value, the horizontal parameter and / or the vertical parameter according to the difference expression;
[0232] It should be noted that the difference expression can be an expression used when calculating the diagonal information according to the hypotenuse parameter, and according to the difference expression, the hypotenuse parameter can be calculated according to the diagonal difference value, the horizontal parameter and / or the vertical parameter by substituting the diagonal difference value, the horizontal parameter and / or the vertical parameter into the difference expression.
[0233] For example: Assuming that the difference expression used when calculating the diagonal information is s=sr_d-(max(sr_x,sr_y)+a2), where s is the diagonal difference value, sr_d is the hypotenuse parameter, sr_x is the horizontal parameter, and sr_y is the vertical parameter, then the hypotenuse parameter sr_d can be calculated by substituting the diagonal difference value, the horizontal parameter, and the vertical parameter into it.
[0234] The embodiment extracts the horizontal parameter and the vertical parameter from the image code stream, determines the diagonal information according to the image code stream, and determines the hypotenuse parameter according to the diagonal information. Since the diagonal information is determined according to the image code stream when the hypotenuse parameter is obtained, and the diagonal information can be the original value of the hypotenuse parameter or the difference value of other values, it is ensured that when the hypotenuse parameter is complex, it can be converted into other values that are easier to encode by calculating the difference value, and the diagonal information can be set to be calculated by combining the horizontal parameter and the vertical parameter, which can ensure that when necessary, the action of analyzing the image code stream can be reduced, thereby improving the speed of image decoding.
[0235] The embodiment of the present application provides an image encoding method, referring to Figure 8 , Figure 8 The flowchart of the first embodiment of the image encoding method of the present application is shown.
[0236] In the embodiment, the image encoding method comprises the following steps:
[0237] Step S100: converting the target image corresponding to the to-be-encoded transform block to obtain the quantized residual data corresponding to the to-be-encoded transform block.
[0238] It should be noted that the target image can be image data to be encoded, and can be set by a manager of the encoding device in advance. The target image is converted into a to-be-encoded transform block, and the quantized residual data corresponding to the to-be-encoded transform block is obtained by converting the to-be-encoded transform block, dividing the target image into at least one to-be-encoded transform block, extracting features of the to-be-encoded transform block, subtracting a predicted mean value, obtaining a residual, and performing transform and quantization on the residual.
[0239] Step S200: Marking a scanning region in the to-be-encoded transform block according to non-zero data in the quantized residual data.
[0240] It should be noted that the scanning region can be composed of points and / or polygonal regions divided according to region syntax parameters.
[0241] Step S300: Encoding the quantized residual data corresponding to the scanning region in sequence according to the region scanning order corresponding to the to-be-encoded transform block, to generate an image code stream corresponding to the target image.
[0242] It should be noted that the region scanning order can include at least one of horizontal scanning, vertical scanning, Z-shaped scanning, and reverse Z-shaped scanning. The quantized residual data corresponding to the scanning region is encoded in sequence according to the region scanning order corresponding to the to-be-encoded transform block, to generate an image code stream corresponding to the target image, which can be a binaryization process of the quantized residual data contained in the scanning region according to the region scanning order corresponding to the to-be-encoded transform block, and the binaryized components are written into the image code stream, thereby generating the image code stream corresponding to the target image.
[0243] Step S400: Constructing region syntax parameters corresponding to the to-be-encoded transform block according to the scanning region, and writing the region syntax parameters into the image code stream.
[0244] It should be noted that the region syntax parameters corresponding to the to-be-encoded transform block are generated according to the horizontal parameter (sr_x), the vertical parameter (sr_y), and the diagonal parameter (sr_d) generated according to the scanning region. In order to facilitate decoding, the horizontal parameter and the vertical parameter can be written into the image code stream, the diagonal information can be generated according to the diagonal parameter, and the diagonal information can also be written into the image code stream or the diagonal information.
[0245] The encoding method used to write the horizontal parameter and the vertical parameter into the image code stream can be group encoding, specifically, the group number can be encoded by using a truncated unary code, and the group internal offset can be encoded by using a fixed-length code. The truncated unary code uses context-based encoding, and the fixed-length code uses bypass encoding.
[0246] In actual use, if the region scanning order is not set in a preset manner, in order to ensure consistency of the region scanning order used in encoding and decoding, the region scanning order can be written into the image code stream through syntax, and the syntax used herein can be sequence level, frame level, patch level, slice level, tile level, CTU level, CTB level, CU level, CB level or TB level syntax.
[0247] In a specific implementation, the image encoding method can be an inverse operation corresponding to the image decoding method, and specific implementation can be referred to the specific description in any of the above image decoding method embodiments, which will not be described here.
[0248] The embodiment converts a to-be-encoded transform block corresponding to a target image to obtain quantized residual data corresponding to the to-be-encoded transform block; marks a scanning region in the to-be-encoded transform block according to non-zero data in the quantized residual data, the scanning region including a non-rectangular region; encodes quantized residual data corresponding to the scanning region in turn according to a region scanning order corresponding to the to-be-encoded transform block to generate an image code stream corresponding to the target image; constructs a region syntax parameter corresponding to the to-be-encoded transform block according to the scanning region, and writes the region syntax parameter into the image code stream. Since the scanning region in the to-be-encoded transform block is marked according to the non-zero data in the quantized residual data after the quantized residual data is obtained, the scanning region can cover the non-zero data in the quantized residual data as much as possible, and since the scanning region can include a non-rectangular region, a more flexible and compact scanning region expression mode can be selected according to actual needs during encoding, thereby reducing the bit cost of image encoding and improving compression performance.
[0249] In addition, the embodiment of the present application also provides a storage medium, the storage medium stores an image decoding program and / or an image encoding program, the image decoding program is executed by a processor to implement the steps of the image decoding method as described above, and the image encoding program is executed by the processor to implement the steps of the image encoding method as described above.
[0250] Reference Figure 9 , Figure 9 is a structural block diagram of a first embodiment of an image decoding device of the present application.
[0251] As Figure 9 shown, the image decoding device provided by the embodiment of the present application comprises:
[0252] The decoding module 10 is configured to extract a to-be-decoded transform block from the image code stream and obtain a region syntax parameter corresponding to the to-be-decoded transform block, the region syntax parameter comprising a horizontal parameter, a vertical parameter and a diagonal parameter.
[0253] The constructing module 20 is configured to determine a scanning region in the to-be-decoded transform block according to the region syntax parameter, the scanning region being a region adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block, and the scanning region including a point and / or a polygon with a side length greater than or equal to three.
[0254] The scanning module 30 is configured to sequentially decode the scanning region according to a region scanning order corresponding to the to-be-decoded transform block, to obtain quantized residual data.
[0255] The reconstructing module 40 is configured to construct a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data.
[0256] The embodiment extracts a to-be-decoded transform block from an image code stream and acquires region syntax parameters corresponding to the to-be-decoded transform block, determines a scanning region in the to-be-decoded transform block according to the region syntax parameters, the scanning region being a region adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block, sequentially decodes the scanning region according to a region scanning order corresponding to the to-be-decoded transform block, to obtain quantized residual data, and constructs a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data. Since the scanning region can be a region including a point and / or a polygon with a side length greater than or equal to three, a more flexible and compact scanning region expression manner can be selected according to the distribution of non-zero quantized residual data when an image is compressed, thereby reducing the bit cost of image coding and improving compression performance.
[0257] In a possible implementation manner of the embodiment, the horizontal parameter is a horizontal coordinate of rightmost non-zero quantized residual data of the scanning region, or a length of a horizontal straight side.
[0258] The vertical parameter is a vertical coordinate of lowermost non-zero quantized residual data of the scanning region, or a length of a vertical straight side.
[0259] The diagonal parameter is a sum of a horizontal coordinate and a vertical coordinate of lowermost non-zero quantized residual data of the scanning region other than the point determined according to the horizontal parameter and the vertical parameter, or a horizontal coordinate of an intersection point of an extension line of a diagonal side of the region shape and an upper boundary of the to-be-decoded transform block.
[0260] In a possible implementation manner of the embodiment, the polygon includes at least one diagonal side.
[0261] In a possible implementation manner of the embodiment, the scanning region includes a pentagon with a 45° diagonal side and a right lower corner point of a circumscribed rectangle of the pentagon.
[0262] In a possible implementation manner of the embodiment, the diagonal parameter is in a range of [0, k], and k is a sum of the horizontal parameter and the vertical parameter minus one.
[0263] In a possible implementation of the present embodiment, the decoding module 10 is further configured to extract the horizontal parameter and the vertical parameter from the image code stream; determine diagonal information according to the image code stream, the diagonal information being an original value of a diagonal parameter or a difference value between the diagonal parameter and a preset function value, the preset function value being set based on the horizontal parameter and / or the vertical parameter; and determine the diagonal parameter according to the diagonal information.
[0264] In a possible implementation of the present embodiment, the decoding module 10 is further configured to extract the diagonal information from the image code stream according to a diagonal coding mode through a diagonal context model; or, obtain a target condition in a parsing condition set, the parsing condition set including at least one first condition, the first condition being set according to at least one of the horizontal parameter, the vertical parameter, a width of the to-be-decoded transform block and a height of the to-be-decoded transform block; and extract the diagonal information according to an information extraction mode corresponding to the target condition.
[0265] In a possible implementation of the present embodiment, the decoding module 10 is further configured to, if the information extraction mode corresponding to the target condition is a direct extraction type, extract the diagonal information from the image code stream according to the diagonal coding mode through the diagonal context model; or, if the information extraction mode corresponding to the target condition is a conversion extraction type, search for a diagonal conversion expression corresponding to the information extraction mode; and determine the diagonal information based on the diagonal conversion expression and the horizontal parameter and / or the vertical parameter.
[0266] In a possible implementation of the present embodiment, the diagonal context model is a context model related to the diagonal information.
[0267] The diagonal context model is set according to indication information and a preset threshold.
[0268] The indication information includes at least one of channel information of the to-be-decoded transform block, transform block information, a quantization parameter, a prediction mode, the horizontal parameter and the vertical parameter, and the transform block information includes at least one of a transform type, a transform block width, a transform block height and an area of the transform block.
[0269] In a possible implementation of the present embodiment, the diagonal coding mode is a coding mode when the diagonal information is coded.
[0270] The diagonal coding mode is a truncated unary code.
[0271] In a possible implementation of the present embodiment, the decoding module 10 is further configured to: take the diagonal information as a hypotenuse parameter; or parse the diagonal information to extract a diagonal difference value; and calculate the hypotenuse parameter according to the diagonal difference value, the horizontal parameter and / or the vertical parameter according to a difference value expression.
[0272] In a possible implementation of the present embodiment, the region scanning order corresponding to the to-be-decoded transform block comprises at least one of horizontal scanning, vertical scanning, zigzag scanning, and reverse zigzag scanning.
[0273] The scanning order is a preset order or is obtained from an image code stream.
[0274] In a possible implementation of the present embodiment, the scanning module 30 is further configured to: obtain a context model corresponding to the scanning region, the context model corresponding to the scanning region comprising a context model used when extracting a binarization component of each quantized residual data in the scanning region; and sequentially decode the scanning region according to the region scanning order corresponding to the to-be-decoded transform block based on the context model, to obtain quantized residual data.
[0275] In a possible implementation of the present embodiment, the context model corresponding to the scanning region is related to at least one of the horizontal parameter, the vertical parameter, the hypotenuse parameter, a region shape of the scanning region, a region area of the scanning region, and a position of quantized residual data in the scanning region.
[0276] In a possible implementation of the present embodiment, the scanning module 30 is further configured to: obtain an area of a polygon in the scanning region; and determine the context model corresponding to the scanning region according to the area of the polygon, the horizontal parameter and the vertical parameter.
[0277] In a possible implementation of the present embodiment, the scanning module 30 is further configured to: when decoding a zero-value flag component, select to-be-decoded residual data from the scanning region according to the region scanning order corresponding to the to-be-decoded transform block; if the to-be-decoded residual data is on a target edge of the scanning region, detect whether the to-be-decoded residual data is the last undecoded residual data on the target edge; if yes, detect whether other zero-value flag components on the target edge except the to-be-decoded residual data are all 0; if yes, set the zero-value flag component of the to-be-decoded residual data to 1; and if not, extract the zero-value flag component of the to-be-decoded residual data from an image code stream according to a context model of the to-be-decoded residual data.
[0278] Refer to Figure 10 , Figure 10 is a structure block diagram of a first embodiment of an image coding device of the present application.
[0279] As shown in Figure 10 the image coding device provided by the embodiment of the present application comprises:
[0280] a conversion module 100, configured to convert a to-be-coded transform block corresponding to a target image, to obtain quantized residual data corresponding to the to-be-coded transform block;
[0281] a marking module 200, configured to mark a scanning region in the to-be-coded transform block according to non-zero data in the quantized residual data, the scanning region comprising a point and / or a polygon with a side length greater than or equal to three;
[0282] a generation module 300, configured to sequentially code quantized residual data corresponding to the scanning region according to a region scanning order corresponding to the to-be-coded transform block, to generate an image code stream corresponding to the target image;
[0283] an encoding module 400, configured to construct a region syntax parameter corresponding to the to-be-coded transform block according to the scanning region, and write the region syntax parameter into the image code stream.
[0284] The embodiment converts a to-be-coded transform block corresponding to a target image, to obtain quantized residual data corresponding to the to-be-coded transform block; marks a scanning region in the to-be-coded transform block according to non-zero data in the quantized residual data, the scanning region comprising a point and / or a polygon with a side length greater than or equal to three; sequentially codes quantized residual data corresponding to the scanning region according to a region scanning order corresponding to the to-be-coded transform block, to generate an image code stream corresponding to the target image; and constructs a region syntax parameter corresponding to the to-be-coded transform block according to the scanning region, and writes the region syntax parameter into the image code stream. Since the scanning region is marked according to non-zero data in the quantized residual data after the quantized residual data is obtained, the scanning region can cover as much non-zero data in the quantized residual data as possible, and since the scanning region can comprise a point and / or a polygon with a side length greater than or equal to three, a more flexible and compact scanning region expression mode can be selected according to actual needs during coding, thereby reducing the bit cost of image coding and improving compression performance.
[0285] It should be understood that the above is only for illustration, and does not constitute any limitation on the technical solutions of the present application. In specific applications, those skilled in the art can set them up according to needs, and the present application does not limit this.
[0286] It should be noted that the above-described workflow is merely illustrative and does not limit the scope of protection of the present application. In actual applications, a person skilled in the art can select part or all of the above-described workflow to achieve the purpose of the embodiment according to actual needs, which is not limited herein.
[0287] In addition, technical details not described in detail in the embodiment can refer to the image decoding method or image encoding method provided by any embodiment of the present application, which will not be described here.
[0288] In addition, it should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or system. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or system including the element.
[0289] The above-mentioned embodiment numbers of the present application are only for description, not representing the advantages and disadvantages of the embodiments.
[0290] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk) and includes a number of instructions to make a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0291] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. An image decoding method characterized by, The image decoding method comprises: extracting a to-be-decoded transform block from an image code stream, and obtaining region syntax parameters corresponding to the to-be-decoded transform block, the region syntax parameters comprising a horizontal parameter, a vertical parameter and a diagonal parameter; determining a scanning region in the to-be-decoded transform block according to the region syntax parameters, the scanning region being a region adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block, the scanning region comprising a point and / or a polygon with a side length greater than or equal to three; decoding the scanning region in sequence according to a region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data; constructing a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data. The horizontal parameter is the horizontal coordinate of the rightmost non-zero quantized residual data in the scanning region, or the length of the horizontal straight edge. The vertical parameter is the vertical coordinate of the lowermost non-zero quantized residual data in the scanning region, or the length of the vertical straight edge. The diagonal parameter is the sum of the horizontal coordinate and the vertical coordinate of the lowermost non-zero quantized residual data in the scanning region other than the point determined according to the horizontal parameter and the vertical parameter, or the horizontal coordinate of the intersection point of the extension line of the diagonal of the region shape and the upper boundary of the to-be-decoded transform block.
2. The image decoding method of claim 1, wherein, The polygon comprises at least one diagonal.
3. The image decoding method of claim 1, wherein, The scanning region comprises a pentagon with a 45° diagonal and a lower right corner point of a circumscribed rectangle of the pentagon.
4. The image decoding method of claim 1, wherein, The value range of the diagonal parameter is [0, k], and k is the sum of the horizontal parameter and the vertical parameter minus one.
5. The image decoding method of claim 1, wherein, The method for obtaining the region syntax parameters corresponding to the to-be-decoded transform block comprises: extracting the horizontal parameter and the vertical parameter from the image code stream; determining diagonal information according to the image code stream, the diagonal information being a diagonal parameter original value, or a difference value between the diagonal parameter and a preset function value, the preset function value being set based on the horizontal parameter and / or the vertical parameter; determining the diagonal parameter according to the diagonal information.
6. The image decoding method of claim 5, wherein, The method for determining the diagonal information according to the image code stream comprises: extracting the diagonal information from the image code stream according to a diagonal coding mode through a diagonal context model; or, obtaining a target condition satisfied in a parsing condition set, the parsing condition set comprising at least one first condition, the first condition being set according to at least one of the horizontal parameter, the vertical parameter, the width of the to-be-decoded transform block and the height of the to-be-decoded transform block; obtaining the diagonal information according to an information extraction mode corresponding to the target condition.
7. The image decoding method of claim 6, wherein, The method for obtaining the diagonal information according to the information extraction mode corresponding to the target condition comprises: if the information extraction mode corresponding to the target condition is a direct extraction type, extracting the diagonal information from the image code stream according to a diagonal coding mode through a diagonal context model; or, if the information extraction mode corresponding to the target condition is a conversion acquisition type, searching for a diagonal conversion expression corresponding to the information extraction mode; determining the diagonal information based on the diagonal conversion expression and the horizontal parameter and / or the vertical parameter.
8. The image decoding method of claim 6, wherein, The diagonal context model is a context model related to the diagonal information. The diagonal context model is set according to the indication information and a preset threshold. The indication information includes at least one of channel information, transform block information, a quantization parameter, a prediction mode, a horizontal parameter and a vertical parameter of the to-be-decoded transform block, and the transform block information includes at least one of a transform type, a transform block width, a transform block height and an area of the transform block.
9. The image decoding method of claim 6, wherein, The diagonal coding mode is a coding mode for coding the diagonal information. The diagonal coding mode is a truncated unary code.
10. The image decoding method of claim 5, wherein, The diagonal information is used as the hypotenuse parameter. Or, The diagonal information is parsed to extract a diagonal difference value. The hypotenuse parameter is calculated according to the diagonal difference value, the horizontal parameter and / or the vertical parameter according to a difference value expression. The region scanning order corresponding to the to-be-decoded transform block includes at least one of horizontal scanning, vertical scanning, zigzag scanning and reverse zigzag scanning.
11. The image decoding method of claim 1, wherein, The scanning order is a preset order or is obtained from an image code stream. The scanning regions are sequentially decoded according to the region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data, including:
12. The image decoding method of claim 1, wherein, A context model corresponding to the scanning region is obtained, and the context model corresponding to the scanning region includes a context model used when extracting a binary component of each quantized residual data in the scanning region. The scanning regions are sequentially decoded according to the region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data based on the context model. The context model corresponding to the scanning region is related to at least one of the horizontal parameter, the vertical parameter, the hypotenuse parameter, a region shape of the scanning region, a region area of the scanning region and a position of the quantized residual data in the scanning region.
13. The image decoding method of claim 12, wherein, The context model corresponding to the scanning region is obtained, including:
14. The image decoding method of claim 13, wherein, An area of a polygon in the scanning region is obtained. The context model corresponding to the scanning region is determined according to the area of the polygon, the horizontal parameter and the vertical parameter. The scanning regions are sequentially decoded according to the region scanning order corresponding to the to-be-decoded transform block to obtain quantized residual data, including:
15. The image decoding method according to any one of claims 1 to 14, wherein When decoding a zero-value flag component, to-be-decoded residual data is selected from the scanning region according to the region scanning order corresponding to the to-be-decoded transform block; If the to-be-decoded residual data is on a target edge of the scanning region, it is detected whether the to-be-decoded residual data is the last undecoded residual data on the target edge; If yes, it is detected whether other zero-value flag components except the to-be-decoded residual data on the target edge are all 0; If all are 0, the zero-value flag component of the to-be-decoded residual data is set to 1; If not all are 0, the zero-value flag component of the to-be-decoded residual data is extracted from an image code stream according to a context model of the to-be-decoded residual data. The image coding method includes:
16. An image coding method characterized by, A to-be-encoded transform block corresponding to a target image is transformed to obtain quantized residual data corresponding to the to-be-encoded transform block. According to the non-zero data in the quantized residual data, a scanning region in the to-be-encoded transform block is marked, the scanning region including a point and / or a polygon with a side length greater than or equal to three; According to a region scanning order corresponding to the to-be-encoded transform block, the quantized residual data corresponding to the scanning region is sequentially encoded, to generate an image code stream corresponding to the target image; According to the scanning region, a region syntax parameter corresponding to the to-be-encoded transform block is constructed, and the region syntax parameter is written into the image code stream, the region syntax parameter including a horizontal parameter, a vertical parameter and a diagonal parameter; The horizontal parameter is a horizontal coordinate of the rightmost non-zero quantized residual data in the scanning region, or a length of a horizontal straight side; The vertical parameter is a vertical coordinate of the lowermost non-zero quantized residual data in the scanning region, or a length of a vertical straight side; The diagonal parameter is a sum of a horizontal coordinate and a vertical coordinate of the lowermost non-zero quantized residual data in the scanning region except for a point determined according to the horizontal parameter and the vertical parameter, or a horizontal coordinate of an intersection point of an extension line of a diagonal side of a region shape and an upper boundary of the to-be-decoded transform block.
17. An image decoding apparatus characterized by comprising: The image decoding apparatus includes: A decoding module is configured to extract a to-be-decoded transform block from an image code stream, and obtain a region syntax parameter corresponding to the to-be-decoded transform block, the region syntax parameter including a horizontal parameter, a vertical parameter and a diagonal parameter; A constructing module is configured to determine a scanning region in the to-be-decoded transform block according to the region syntax parameter, the scanning region being a region adapted to non-zero data in quantized residual data corresponding to the to-be-decoded transform block, the scanning region including a point and / or a polygon with a side length greater than or equal to three; A scanning module is configured to sequentially decode the scanning region according to a region scanning order corresponding to the to-be-decoded transform block, to obtain quantized residual data; A reconstructing module is configured to construct a reconstructed image block corresponding to the to-be-decoded transform block according to the quantized residual data. The horizontal parameter is a horizontal coordinate of the rightmost non-zero quantized residual data in the scanning region, or a length of a horizontal straight side; The vertical parameter is a vertical coordinate of the lowermost non-zero quantized residual data in the scanning region, or a length of a vertical straight side; The diagonal parameter is a sum of a horizontal coordinate and a vertical coordinate of the lowermost non-zero quantized residual data in the scanning region except for a point determined according to the horizontal parameter and the vertical parameter, or a horizontal coordinate of an intersection point of an extension line of a diagonal side of a region shape and an upper boundary of the to-be-decoded transform block.
18. An image coding apparatus characterized by comprising: The image encoding apparatus includes: A converting module is configured to convert a to-be-encoded transform block corresponding to a target image, to obtain quantized residual data corresponding to the to-be-encoded transform block; A marking module is configured to mark a scanning region in the to-be-encoded transform block according to non-zero data in the quantized residual data, the scanning region including a point and / or a polygon with a side length greater than or equal to three; A generating module is configured to sequentially encode quantized residual data corresponding to the scanning region according to a region scanning order corresponding to the to-be-encoded transform block, to generate an image code stream corresponding to the target image. The encoding module is configured to construct a region syntax parameter corresponding to the to-be-encoded transform block according to the scanning region, and write the region syntax parameter into the image code stream, wherein the region syntax parameter comprises a horizontal parameter, a vertical parameter and a diagonal parameter. The horizontal parameter is the horizontal coordinate of the rightmost non-zero quantized residual data in the scanning region, or the length of the horizontal straight side. The vertical parameter is the vertical coordinate of the lowermost non-zero quantized residual data in the scanning region, or the length of the vertical straight side. The diagonal parameter is the sum of the horizontal coordinate and the vertical coordinate of the lowermost right non-zero quantized residual data in the scanning region except the point determined according to the horizontal parameter and the vertical parameter, or the horizontal coordinate of the intersection point of the extension line of the diagonal side of the region shape and the upper boundary of the to-be-decoded transform block.
19. A decoding device, comprising: The decoding device comprises a processor, a memory and an image decoding program stored on the memory and executable on the processor, and the image decoding program is executed by the processor to implement the image decoding method according to any one of claims 1-15.
20. An encoding device, comprising: The encoding device comprises a processor, a memory and an image encoding program stored on the memory and executable on the processor, and the image encoding program is executed by the processor to implement the image encoding method according to claim 16.
21. A storage medium, characterized by The storage medium stores an image decoding program and / or an image encoding program, the image decoding program is executed to implement the image decoding method according to any one of claims 1-15, and the image encoding program is executed to implement the image encoding method according to claim 16.
Citation Information
Patent Citations
Residual coefficient coding and decoding
CN114365492A
Video coding and decoding method and device, computer readable medium and electronic equipment
CN114979642A