Word-containing Image Encoding Method and Apparatus, and Word-containing Image Decoding Method and Apparatus
By recognizing and aligning word regions to match CU sizes, the method improves the accuracy and efficiency of word content compression in video encoding, addressing the issue of pixels being split across multiple CUs.
Patent Information
- Application Number
- JP2025501681
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-15
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Existing video encoding methods for images containing words, such as those with text content, suffer from low accuracy and efficiency due to pixels of the same word being classified into different Coding Units (CUs), leading to suboptimal compression results.
A method and apparatus for encoding and decoding word-containing images that involve recognizing word regions, filling them to match the size of CUs, and aligning with CU boundaries to avoid classification into multiple CUs, thereby improving compression accuracy and efficiency.
This approach enhances the accuracy and efficiency of word content compression by ensuring pixels of the same character are not split across different CUs, resulting in improved encoding and decoding performance.
Smart Images

Figure 2025522097000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to Chinese Patent Application No. 202210829379.1, titled "Word-Containing Image Encoding Method and Apparatus, and Word-Containing Image Decoding Method and Apparatus", filed with the China National Intellectual Property Administration on July 15, 2022, which is incorporated herein by reference in its entirety. This application relates to video encoding and decoding technologies, and particularly to a word-containing image encoding method and apparatus, and a word-containing image decoding method and apparatus.
Background Art
[0002] In a High Efficiency Video Coding (HEVC) encoder, an input image is first divided into a plurality of Coding Tree Units (CTUs) that have the same size and do not overlap with each other. The size of the CTU may be determined by the encoder, and the maximum size of the CTU may be 64×64. One CTU may be directly used as one Coding Unit (CU), or may be further divided into a plurality of CUs of smaller sizes in a quadtree recursive partitioning manner. The CU partitioning depth is determined based on the rate-distortion cost obtained by calculation. The encoder compares different partitioning modes to balance and selects the partitioning mode with the lowest rate-distortion cost for encoding. Finally, the CU located at the leaf node of the quadtree is the basic unit for the encoder to perform subsequent prediction, transformation, and encoding.
[0003] However, when the above encoding method is applied to an image containing word content, pixels of the same word may be classified into different CTUs or CUs, resulting in low accuracy and efficiency of word content compression during prediction, transformation, and quantization.
Summary of the Invention
Means for Solving the Problems
[0004] This application provides a word-containing image encoding method, an apparatus for encoding a word-containing image, a word-containing image decoding method, and an apparatus for decoding a word-containing image in order to improve the accuracy and efficiency of word content compression and to avoid classifying pixels of the same character into different CUs as much as possible.
[0005] According to a first aspect, this application provides a method for encoding a word-containing image, the method including: obtaining a word region in a first image, the word region including at least one character; filling the word region to obtain a word filling region, the height or width of the word filling region being n times a preset size, where n≥1; obtaining a second image based on the word filling region; and encoding the second image and word alignment side information to obtain a first bitstream, the word alignment side information including the height or width of the word filling region.
[0006] In this embodiment of this application, the word region obtained by recognition in the original image is filled so that the size of the obtained word filling region matches the size of the CU for encoding, implementing alignment between the word filling region and the CU, avoiding classifying pixels of the same character into different CUs as much as possible, and improving the accuracy and efficiency of word content compression.
[0007] The first image may be an independent image or any video frame within a video. This is not specifically limited. In the present embodiment of the present application, the first image may include a word region. There may be only one or more word regions. The number of word regions may depend on the number of characters included in the first image, the character distribution rule, the character size, etc. This is not specifically limited. In addition, the word region can include one character, that is, the word region can include only one character, or the word region can include one line of characters, the line of characters includes a plurality of characters, or the word region can include one string of characters, and the string of characters includes a plurality of characters. The aforementioned characters may be Chinese characters, Chinese phonetic characters, English characters, numbers, etc. This is not specifically limited.
[0008] In this embodiment of the present application, the word region in the first image can be obtained using the following four methods.
[0009] In the first method, character recognition is performed on the first image to obtain the word region.
[0010] Binarization and dilation can be first performed on the first image. Usually, in the binarized image, the samples corresponding to the word region are white, and the samples corresponding to the non-word region are black. Dilation can be used to handle the pixel adhesion between characters. Then, horizontal projection is performed on the image obtained through binarization and dilation, and the number of white samples in the horizontal direction (corresponding to the word region) is calculated. The amount of white samples on the horizontal line having characters is not zero. Therefore, all lines of characters in the image can be obtained based on this.
[0011] For each line of text, the following operations are performed: extracting connected components (a connected component refers to a continuous white sample area, and one character is a connected component, and correspondingly, one line of text can have multiple connected components), and selecting a character box of an appropriate size. If the character box contains only one character, the size of the character box is greater than or equal to the size of the largest connected component within the line of text, and all characters within the same line of text use character boxes of the same size. If the character box contains one line of text or a string of characters, the size of the character box is greater than or equal to the sum of the sizes of all connected components within the line of text or the string of characters. For example, the connected component of character 1 within a line of text occupies a pixel space of 6 (width) × 10 (height), and the connected component of character 2 occupies a pixel space of 11 × 8. In this case, at least a character box of 11 × 10 should be selected to cover the two characters, or at least a character box of 17 × 10 should be selected to cover the two characters of the line of text. In the present embodiment of the present application, the area covered by the character box in the first image can be referred to as a word area, that is, the word area and the character box refer to the same area range in the first image.
[0012] Therefore, in order to obtain the word area, the first image can be recognized using the method described above.
[0013] In the second method, character recognition is performed on the first image to obtain an intermediate word area, the intermediate word area contains multiple lines of text, and then line splitting is performed for each line of text in the intermediate word area to obtain the word area.
[0014] The difference from the first method is that in the second method, the intermediate word area can be obtained by character recognition, the intermediate word area can contain multiple lines of text, and then line splitting is performed on the intermediate word area to obtain multiple word areas, each of which contains only one line of text. This method may be used when the characters are arranged horizontally. For example, multiple characters are arranged from left to right and from top to bottom.
[0015] In the third method, character recognition is performed on the first image to obtain an intermediate word area, the intermediate word area includes a plurality of character strings, and column splitting is performed for each character string in the intermediate word area to obtain word areas.
[0016] The difference from the first method is that in the third method, the intermediate word area can be obtained by character recognition, the intermediate word area can include a plurality of character lines, and then column splitting is performed on the intermediate word area to obtain a plurality of word areas each containing only one character string. This method may be used when characters are arranged vertically. For example, a plurality of characters are arranged from top to bottom and from left to right.
[0017] In the fourth method, character recognition is performed on the first image to obtain an intermediate word area, the intermediate word area includes a plurality of character lines or a plurality of character strings, and line splitting and column splitting are performed for each character in the intermediate word area to obtain word areas.
[0018] The difference from the first method is that in the fourth method, the intermediate word area can be obtained by character recognition, the intermediate word area can include a plurality of character lines or a plurality of character strings, and then line splitting and column splitting are performed on the intermediate word area to obtain a plurality of word areas each containing only one character. This method may be used when characters are arranged horizontally or vertically. For example, a plurality of characters are arranged from top to bottom and from left to right. In another example, a plurality of characters are arranged from top to bottom and from left to right.
[0019] It should be noted that in addition to the four methods described above, in this embodiment of the present application, another method may be used to obtain the word area in the first image. This is not specifically limited.
[0020] The preset size may be any one of 8, 16, 32, and 64. In this embodiment of the present application, related encoding and decoding methods can be used to encode an image. In the related encoding and decoding methods, an image usually needs to be divided to obtain coding units (CUs), and subsequent processing is performed based on the CUs. The size of a CU may be one of 8, 16, 32, and 64. In this embodiment of the present application, the preset size may be determined based on the size of the CU. For example, when the size of the CU is 8, the preset size may be 8 or 16, or when the size of the CU is 32, the preset size may be 32.
[0021] In this embodiment of the present application, the word filling area can be obtained by filling the word area. Therefore, the height of the word filling area is greater than or equal to the height of the word area, and / or the width of the word filling area is greater than or equal to the width of the word area. In the aforementioned case where the height of the word filling area is equal to the height of the word area, and / or the width of the word filling area is equal to the width of the word area, it may mean that the height and / or width of the word area is exactly n times the preset size. Therefore, the word area does not need to be filled, and the subsequent steps are performed by directly using the word area as the word filling area. In addition, as described in step 801, when the word area is obtained, the height and / or width of the word areas of the characters in the same row or the same column are kept constant. Correspondingly, when the word filling area is obtained, the height and / or width of the word areas in the same row or the same column are also kept constant. In this way, it can be guaranteed that the height of the word filling area coincides with the height of the CU obtained by the division in the related encoding and decoding methods, and / or the width of the word filling area coincides with the width of the CU obtained by the division in the related encoding and decoding methods. In this embodiment of the present application, the aforementioned effect may also be referred to as the word filling area being aligned with the CU obtained by the division in the related encoding and decoding methods.
[0022] In a possible embodiment, the height of the word filling area is n times the preset size. For example, when the height of the word area is 10 and the preset size is 8, the height of the word filling area may be 16 (2 times 8). As another example, when the height of the word area is 10 and the preset size is 16, the height of the word filling area may be 16 (1 times 16).
[0023] In a possible embodiment, the width of the word filling area is n times the preset size. For example, when the width of the word area is 17 and the preset size is 8, the width of the word filling area may be 24 (3 times 8). As another example, when the width of the word area is 17 and the preset size is 16, the width of the word filling area may be 32 (2 times 16).
[0024] In a possible embodiment, the height of the word filling area is n1 times the preset size, the width of the word filling area is n2 times the preset size, and n1 and n2 may or may not be equal. For example, when the size of the word area is 6 (width) × 10 (height) and the preset size is 8, the size of the word filling area may be 8 (1 times 8) × 16 (2 times 8). As another example, when the size of the word area is 17 × 10 and the preset size is 16, the size of the word filling area may be 32 (2 times 16) × 16 (1 times 16).
[0025] In this embodiment of the present application, in order to obtain the word filling area, the following three methods can be used to fill the word area.
[0026] In the first method, the height of the word filling area is obtained based on the height of the word area. The height of the word filling area is n times the preset size. The word area is filled in the vertical direction of the word area to obtain the word filling area, and the filling height is the difference between the height of the word filling area and the height of the word area.
[0027] As described above, the height of the word filling area is greater than the height of the word area, and the height of the word filling area is n times the preset size. Based on this, the height of the word filling area can be obtained based on the height of the word area. Refer to the above example. Optionally, when the height of the word area is exactly n times the preset size, the word area does not need to be filled. In this case, the height of the word filling area is equal to the height of the word area. After the height of the word filling area is obtained, the word area may be filled in the vertical direction of the word area (for example, the upper side, the middle, or the lower side of the word area, which is not specifically limited), and the filling height is the difference between the height of the word filling area and the height of the word area. For example, when the height of the word filling area is 16 and the height of the word area is 10, the filling height is 6. The pixel value of the filling may be the pixel value of the background pixel other than the character pixel in the word area, or alternatively, the preset pixel value, such as 0 or 255. This is not specifically limited.
[0028] In the second method, the width of the word filling area is obtained based on the width of the word area, the width of the word filling area is n times the preset size, the word area is filled in the horizontal direction of the word area to obtain the word filling area, and the filling width is the difference between the width of the word filling area and the width of the word area.
[0029] As described above, the width of the word filling area is greater than the width of the word area, and the width of the word filling area is n times the preset size. Based on this, the width of the word filling area can be obtained based on the width of the word area. Refer to the above example. Optionally, when the width of the word area is exactly n times the preset size, the word area does not need to be filled. In this case, the width of the word filling area is equal to the width of the word area. After the width of the word filling area is obtained, the word area may be filled in the horizontal direction of the word area (for example, on the left side, in the center, or on the right side of the word area, which is not specifically limited), and the width of the filling is the difference between the width of the word filling area and the width of the word area. For example, when the width of the word filling area is 24 and the width of the word area is 17, the width of the filling is 7. The pixel value of the filling may be the pixel value of the background pixel other than the character pixel in the word area, or alternatively, the preset pixel value, such as 0 or 255. This is not specifically limited.
[0030] The third method may be a combination of the first method and the second method, that is, after the height and width of the word filling area are obtained, in order to obtain the word filling area, the first two methods are respectively used to fill the word area in the vertical and horizontal directions of the word area.
[0031] It can be found that the size of the word filling area obtained by filling has a plurality of relationships with the size of the CU. Therefore, after such a word filling area is divided, the word filling area can cover the complete CU, that is, the case where the word filling area covers a part of the CU does not occur, and the same CU does not cover two word filling areas, thereby implementing the alignment between the word area and the CU during division.
[0032] In this embodiment of the present application, a new image, i.e., a second image, can be obtained based on the word filling area. The second image contains the characters in the original first image. Obtaining the second image includes the following four cases.
[0033] 1. When there is only one word filling area, the word filling area is used as the second image.
[0034] When there is only one word filling area, the word filling area may be directly used as the second image.
[0035] 2. When there are multiple word filling areas and each of the word filling areas contains one line of characters, the multiple word filling areas are joined from top to bottom to obtain the second image.
[0036] One word filling area contains one line of characters. In this case, the word filling areas to which the multiple lines of characters belong may be joined from top to bottom based on the distribution of the lines of characters in the first image. Specifically, the word filling area to which the line of characters arranged at the top of the first image belongs is also joined at the top, the word filling area to which the line of characters arranged in the second line of the first image belongs is also joined in the second line, and the rest may be estimated by analogy. It should be noted that in this embodiment of the present application, the joining may alternatively be performed from bottom to top, and the order of joining is not specifically limited.
[0037] 3. When there are multiple word filling areas and each of the word filling areas contains one string of characters, the multiple word filling areas are joined from left to right to obtain the second image.
[0038] One word filling area contains one string. In this case, the word filling areas to which a plurality of strings respectively belong may be joined from left to right based on the distribution of the strings in the first image. Specifically, the word filling area to which the string arranged on the leftmost side in the first image belongs is also joined on the leftmost side, the word filling area to which the string arranged in the second column in the first image belongs is also joined in the second column, and the rest may be estimated by analogy. It should be noted that in this embodiment of the present application, the joining may alternatively be performed from right to left, and the joining order is not specifically limited.
[0039] 4. When there are multiple word filling areas and each of the word filling areas contains one character, the multiple word filling areas are joined from left to right and from top to bottom to obtain the second image, and the multiple word filling areas in the same row have the same height or the multiple word filling areas in the same column have the same width.
[0040] One word filling area contains one character. In this case, based on the distribution of the characters in the first image, the word filling areas to which a plurality of characters respectively belong may be joined from left to right and from top to bottom. The arrangement order of the word filling areas is consistent with the arrangement order of the characters included in the word filling areas in the first image, and the multiple word filling areas in the same row have the same height or the multiple word filling areas in the same column have the same width.
[0041] It should be noted that in this embodiment of the present application, the joining may alternatively be performed in another order, and the joining order is not specifically limited.
[0042] In this embodiment of the present application, standard video encoding or IBC encoding can be performed on the second image. This is not specifically limited. When IBC encoding is used, if the characters within the currently encoded CU already appear in the encoded CU, the encoded CU can be determined as the predicted block of the currently encoded CU. In this way, the probability that the obtained residual block is 0 becomes very high, thereby improving the encoding efficiency of the current CU. The word alignment side information may be encoded in an exponential Golomb coding scheme or in another coding scheme. This is also not specifically limited. In this embodiment of the present application, the height or width of the word filling area may be directly encoded, or n times the preset size of the height or width of the word filling area may be encoded. This is not specifically limited.
[0043] Optionally, the word alignment side information may further include the height and width of the word area, and the horizontal and vertical coordinates of the pixel at the upper left corner of the word area in the first image. This information can assist the decoder side in restoring the word area in the reconstructed image, thereby improving the decoding efficiency.
[0044] In a possible implementation, to obtain the third image, the pixel values within the word area in the first image may be filled with preset pixel values, and the third image is encoded to obtain the second bitstream.
[0045] In the above steps, the word area is recognized from the first image, and then the word area is filled to obtain the word filling area. The second image is obtained by joining the word filling area. It can be found that the second image only includes the character content in the first image, and the second image is encoded separately. However, the first image includes other content in addition to the characters. Therefore, in order to maintain the integrity of the image, other image content other than the characters needs to be further processed and encoded.
[0046] In this embodiment of the present application, the pixel values of the pixels in the recognized word region in the original first image can be filled with a preset pixel value (for example, 0, 1, or 255, but not limited thereto). This corresponds to removing the word region from the first image and replacing the pixel values with the pixel values of non-character pixel values. The difference between the third image obtained in this way and the first image lies in the different pixel values of the word region. Standard video encoding or IBC encoding can also be performed on the third image. This is not specifically limited.
[0047] According to a second aspect, the present application provides a method for decoding a word-containing image. The method includes the steps of obtaining a bitstream, decoding the bitstream to obtain a first image and word alignment side information, where the word alignment side information includes the height or width of a word filling region included in the first image, and the height or width of the word filling region is n times a preset size, where n≥1, obtaining a word filling region based on the first image and the height or width of the word filling region, and obtaining a word region in the image to be reconstructed based on the word filling region.
[0048] In this embodiment of the present application, the alignment information of the word region can be obtained by analyzing the bitstream to obtain a character filling region containing characters. The size of the character filling region matches the size of the CU for decoding in order to improve the accuracy and efficiency of word content compression and to implement the alignment between the word filling region and the CU.
[0049] The decoder side can obtain the first image and word alignment side information through decoding. The decoding method can correspond to the encoding method used on the encoder side. The encoder side can encode an image using standard video encoding or IBC encoding, and the decoder side can decode the bitstream using standard video decoding or IBC decoding to obtain the reconstructed image. The encoder side encodes the word alignment side information using Golomb coding, and the decoder side can decode the bitstream using Golomb decoding to obtain the word alignment side information.
[0050] The height or width of the word filling area is n times the preset size, where n ≥ 1, and the preset size is any one of 8, 16, 32, and 64. For details, please refer to the following description of step 802. Details will not be described again here.
[0051] After determining the height or width of the word filling area, the decoder side can extract the corresponding pixel values from the first image based on the height or width to obtain the word filling area.
[0052] In this embodiment of the present application, the word alignment side information further includes the height and width of the word area. Therefore, based on the height and width of the word area, the corresponding pixel values can be extracted from the word filling area to obtain the word area in the image to be reconstructed. The word area corresponds to the word area in step 801 and can contain one character, or can contain one line of characters, where the line of characters contains multiple characters, or can contain one string, and the string contains multiple characters.
[0053] In a possible implementation, the decoder side can decode the bitstream to obtain the second image and obtain the image to be reconstructed based on the word area and the second image.
[0054] The word alignment side information further includes the height and width of the word area, as well as the horizontal and vertical coordinates of the pixel at the upper left corner of the word area. Based on this, the replacement area in the second image can be determined based on the height and width of the word area, as well as the horizontal and vertical coordinates of the pixel at the upper left corner of the word area, and the pixel values in the replacement area are filled with the pixel values in the word area to obtain the image to be reconstructed. This process may be the reverse of the process of obtaining the third image in the embodiment of the first aspect, that is, the process of filling the word area in the second image (corresponding to the first image in the embodiment shown in FIG. 8) to obtain the reconstructed image (corresponding to the third image in the embodiment of the first aspect). After the horizontal and vertical coordinates of the pixel at the upper left corner of the word area in the second image, as well as the height and width of the word area are obtained, the replacement area in the second image may be determined, and then the pixel values in the filling area are filled with the pixel values of the word area. This corresponds to removing the replacement area from the second image and replacing the pixel values in the filling area with the pixel values of the word area.
[0055] According to a third aspect, the present application provides an encoding device, which includes an acquisition module configured to acquire a word area in a first image, where the word area includes at least one character; a filling module configured to fill the word area to acquire a word filling area, where the height or width of the word filling area is n times the preset size, and n≥1; and an encoding module configured to encode the second image and the word alignment side information to acquire a first bit stream, where the word alignment side information includes the height or width of the word filling area.
[0056] In a possible implementation, the filling module obtains the height of the word filling area based on the height of the word area. The height of the word filling area is n times the preset size. To obtain the word filling area, the word area is filled in the vertical direction of the word area, and the filling height is the difference between the height of the word filling area and the height of the word area. Specifically, it is configured in this way.
[0057] In a possible implementation, the filling module obtains the width of the word filling area based on the width of the word area. The width of the word filling area is n times the preset size. To obtain the word filling area, the word area is filled in the horizontal direction of the word area, and the filling width is the difference between the width of the word filling area and the width of the word area. Specifically, it is configured in this way.
[0058] In a possible implementation, the word area can contain one character, or the word area can contain one line of characters, the line of characters contains multiple characters, or the word area can contain one string, and the string contains multiple characters.
[0059] In a possible implementation, the acquisition module is specifically configured to perform character recognition on the first image to obtain the word area.
[0060] In a possible implementation, the acquisition module performs character recognition on the first image to obtain the intermediate word area. The intermediate word area contains multiple lines of characters. To obtain the word area, line splitting is performed on each line of the intermediate word area. Specifically, it is configured in this way.
[0061] In a possible implementation, the acquisition module performs character recognition on the first image to obtain the intermediate word area. The intermediate word area contains multiple strings. To obtain the word area, column splitting is performed on each string of the intermediate word area. Specifically, it is configured in this way.
[0062] In a possible implementation, the acquisition module is specifically configured to perform character recognition on the first image to acquire an intermediate word area, where the intermediate word area includes a plurality of character lines or a plurality of character strings, and perform line division and column division for each character in the intermediate word area to acquire a word area.
[0063] In a possible implementation, when there is only one word filling area, the filling module is specifically configured to use the word filling area as the second image.
[0064] In a possible implementation, when there are a plurality of word filling areas and each of the word filling areas includes one character line, the filling module is specifically configured to join the plurality of word filling areas from top to bottom to acquire the second image.
[0065] In a possible implementation, when there are a plurality of word filling areas and each of the word filling areas includes one character string, the filling module is specifically configured to join the plurality of word filling areas from left to right to acquire the second image.
[0066] In a possible implementation, when there are a plurality of word filling areas and each of the word filling areas includes one character, the filling module is specifically configured to join the plurality of word filling areas from left to right and from top to bottom, and the plurality of word filling areas in the same row have the same height or the plurality of word filling areas in the same column have the same width.
[0067] In a possible implementation, the word alignment side information further includes the height and width of the word area, and the horizontal coordinate and vertical coordinate of the pixel at the upper left corner of the word area in the first image.
[0068] In a possible implementation, the preset size is one of 8, 16, 32, and 64.
[0069] In a possible implementation, the filling module is further configured to fill pixel values within a word region in the first image with preset pixel values in order to obtain a third image, and the encoding module is further configured to encode the third image in order to obtain a second bitstream.
[0070] According to a fourth aspect, the present application provides a decoding device, which includes an acquisition module configured to acquire a bitstream, and a decoding module configured to decode the bitstream in order to acquire a first image and word alignment side information, where the word alignment side information includes the height or width of a word filling region included in the first image, the height or width of the word filling region is n times a preset size, and n≥1, and a reconstruction module configured to acquire the word filling region based on the first image and the height or width of the word filling region, and acquire a word region in a target image to be reconstructed based on the word filling region.
[0071] In a possible implementation, the word alignment side information further includes the height and width of the word region, and the reconstruction module is specifically configured to extract corresponding pixel values from the word filling region based on the height and width of the word region in order to acquire the word region.
[0072] In a possible implementation, the decoding module is further configured to decode the bitstream in order to obtain a second image, and the reconstruction module is further configured to obtain a target image to be reconstructed based on the word region and the second image.
[0073] In a possible implementation, the word alignment side information further includes the horizontal and vertical coordinates of the pixel at the upper left corner of the word region, and the reconstruction module is specifically configured to determine a replacement region in the second image based on the height and width of the word region and the horizontal and vertical coordinates of the pixel at the upper left corner of the word region, and fill the pixel values in the replacement region with the pixel values in the word region to obtain the image to be reconstructed.
[0074] In a possible implementation, the word region includes one character, or the word region includes one line of characters, the line of characters includes a plurality of characters, or the word region includes one string of characters, and the string of characters includes a plurality of characters.
[0075] In a possible implementation, the preset size is one of 8, 16, 32, and 64.
[0076] According to a fifth aspect, this application provides an encoder including one or more processors and a memory configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are capable of implementing the method according to any one of the possible implementations of the first aspect.
[0077] According to a sixth aspect, this application provides a decoder including one or more processors and a memory configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are capable of implementing the method according to any one of the possible implementations of the second aspect.
[0078] According to a seventh aspect, this application provides a computer-readable storage medium including a computer program. When the computer program is executed on a computer, the computer is capable of implementing the method according to any one of the possible implementations of the first aspect and the second aspect.
[0079] According to an eighth aspect, the present application provides a computer program product. The computer program product includes instructions. When the instructions are executed on a computer or a processor, the computer or the processor can implement a method according to any one of the possible embodiments of the first aspect and the second aspect.
Brief Description of the Drawings
[0080]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11a
Figure 11b
Figure 11c
Figure 12a
Figure 12b
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0081] To make the object, technical solution, and advantages of the present application clearer, the following will clearly and completely describe the technical solution of the present application with reference to the accompanying drawings of the present application. It is obvious that the described embodiments are only a part, not all, of the embodiments of the present application. Any other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0082] In the specification, embodiments, claims, and accompanying drawings of the present application, terms such as "first" and "second" are merely intended for distinction and description, and should not be understood as indicating or implying relative importance, or indicating or implying order. In addition, terms such as "including", "having", and any variations thereof are intended to cover non-exclusive inclusion, for example, including a series of steps or units. A method, system, product, or device is not necessarily limited to the explicitly listed steps or units, and may include other steps or units not explicitly listed or inherent to such a process, method, product, or device.
[0083] In this application, it should be understood that "at least one" means one or more, and "a plurality" means two or more. The term "and / or" is used to describe the relationship between related objects and represents that three relationships can exist. For example, "A and / or B" can represent the following three cases, namely, only A exists, only B exists, and both A and B exist, and A and B can be singular or plural. The symbol " / " generally indicates an "or" relationship between associated objects. "At least one of the following items" or a similar expression indicates any combination of these items, including any combination of one item or a plurality of items. As an example, at least one of a, b, or c may refer to a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", and a, b, and c can be in singular or plural form.
[0084] Video coding generally refers to the processing of a sequence of images, and the sequence of images forms a video or a video sequence. In the field of video coding, the terms "picture", "frame", and "image" can be used synonymously. Video coding (or generally coding) includes two parts: video coding and video decoding. Video coding is performed on the sender side and usually includes processing the original video image (for more efficient storage and / or transmission) to reduce the amount of data required to represent that video image (e.g., by compression). Video decoding is performed on the receiver side and generally includes performing the reverse process compared to the encoder's processing to reconstruct the video image. The "coding" of a video image (or a general image) in an embodiment should be understood as the "coding" or "decoding" of a video image or a video sequence. The combination of the coding part and the decoding part is also referred to as coding and decoding (CODEC).
[0085] In the case of reversible video coding, the original video image can be reconstructed. In other words, the reconstructed video image has the same quality as the original video image (assuming no transmission loss or other data loss occurs during storage or transmission). In the case of irreversible video coding, in order to reduce the amount of data required to represent the video image, further compression is performed, such as quantization, and the video image cannot be completely reconstructed on the decoder side. In other words, the quality of the reconstructed video image is lower or worse than that of the original video image.
[0086] Some video coding standards are used for "irreversible hybrid video coding and decoding" (i.e., spatial prediction and temporal prediction in the pixel domain are combined with 2D transform coding for applying quantization in the transform domain). Each image of a video sequence is typically divided into a set of non-overlapping blocks, and coding is typically performed at the block level. Specifically, in the encoder, the video is typically processed, i.e., coded, at the block (video block) level. For example, a predicted block is generated by spatial (intra) prediction and temporal (inter) prediction, the predicted block is subtracted from the current block (the block being processed or to be processed) to obtain a residual block, and the residual block is transformed and quantized (compressed) in the transform domain to reduce the amount of data to be transmitted. On the decoder side, the inverse processing part for the encoder is applied to the coded block, or the compressed block, to reconstruct the current block for display. Further, the encoder needs to replicate the processing steps of the decoder so that the encoder and the decoder process, i.e., code, subsequent blocks using the same prediction (e.g., intra prediction or inter prediction) and / or generate the same reconstructed pixels.
[0087] In the following embodiments of the coding system 10, the encoder 20 and the decoder 30 are described with reference to FIGS. 1A to 3.
[0088] FIG. 1A is an exemplary block diagram of an encoding system 10 according to the present application, e.g., a video encoding system 10 (or simply encoding system 10) that can use the techniques in the present application. The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video encoding system 10 represent devices that can be configured to execute techniques according to various examples described in the present application.
[0089] As shown in FIG. 1A, the encoding system 10 includes a source device 12. The source device 12 is configured to provide encoded image data 21, e.g., an encoded image for a destination device 14 for decoding the encoded image data 21.
[0090] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include an image source 16, a preprocessor (or preprocessing unit) 18, e.g., an image preprocessor, and a communication interface (or communication unit) 22.
[0091] The image source 16 can include any type of image capture device, e.g., a camera for capturing real-world images, and / or any type of image generation device, e.g., a computer graphics processing unit for generating computer animation images, or any other device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source can be any type of memory or storage device that stores any of the aforementioned images.
[0092] To distinguish the processing performed by the preprocessor (or preprocessing unit) 18, the image (or image data) 17 can also be referred to as the original image (or original image data) 17.
[0093] The pre-processor 18 is configured to receive the (original) image data 17 and perform pre-processing on the image data 17 to obtain pre-processed images (or pre-processed image data) 19. The pre-processing executed by the pre-processor 18 may include, for example, trimming, color format conversion (e.g., conversion from RGB to YCbCr), color correction, or noise removal. It will be understood that the pre-processing unit 18 may be an optional component.
[0094] The video coder (or coder) 20 is configured to receive the pre-processed image data 19 and provide encoded image data 21 (further explanation will be provided below based on Figure 2 and the like).
[0095] The communication interface 22 of the source device 12 is configured to receive the encoded image data 21 and transmit the encoded image data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0096] The destination device 14 may additionally, i.e., optionally, include a decoder 30, a communication interface (or communication unit) 28, a post-processor (post-processing unit) 32, and a display device 34.
[0097] The communication interface 28 of the destination device 14 is configured to receive the encoded image data 21 (or any further processed version thereof) directly from the source device 12 or from a storage device, such as another source device like an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0098] Communication interfaces 22 and 28 may be configured to transmit and receive encoded image data (or encoded data) 21 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any kind of network, such as a wired or wireless network or any combination thereof, or any kind of private and public network, or any combination of any kind thereof.
[0099] Communication interface 22 may be configured to, for example, package the encoded image data 21 into an appropriate format, such as a packet, and / or process the encoded image data using any kind of transmission encoding or processing for transmission via the communication link or communication network.
[0100] Communication interface 28 forms the counterpart of communication interface 22 and may be configured to, for example, receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpacking to obtain the encoded image data 21.
[0101] Both communication interface 22 and communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow of communication channel 13 in FIG. 1A pointing from the source device 12 to the destination device 14, for example, to send and receive messages, for example, to set up a connection, and to confirm and exchange any other information related to the communication link and / or data transmission, for example, the transmission of the encoded image data.
[0102] The video decoder (or decoder) 30 is configured to receive the encoded image data 21 and provide a decoded image (or decoded image data) 31 (further explanation will be provided below based on, for example, FIG. 3).
[0103] The post-processor 32 is configured to post-process the decoded image data 31 (also referred to as the reconstructed image data) to obtain post-processed image data 33, for example, a post-processed image. The post-processing executed by the post-processing unit 32 includes, for example, color format conversion (e.g., conversion from YCbCr to RGB), color correction, trimming, resampling, or any other processing, for example, processing for preparing the decoded image data 31 for display by the display device 34.
[0104] The display device 34 is configured to receive the post-processed image data 33 for displaying an image to a user or viewer, for example. The display device 34 can be or include any type of display for displaying the reconstructed image, such as an integrated or external display or monitor. For example, the display can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0105] The encoding system 10 further includes a training engine 25. The training engine 25 is configured to train one or more modules of the encoder 20 or the decoder 30 or the encoder 20 or the decoder 30.
[0106] In this embodiment of the present application, the training data can be stored in a database (not shown), and the training engine 25 executes training based on the training data to obtain a target model. It should be noted that the source of the training data is not limited in the embodiments of the present application. For example, the training data may be obtained from the cloud or other locations for performing model training.
[0107] The target model obtained by training with the training engine 25 may be applied to the encoding system 10, for example, to the source device 12 (e.g., the encoder 20) or the destination device 14 (e.g., the decoder 30) shown in FIG. 1A. The training engine 25 may obtain the target model through training on the cloud, and then the encoding system 10 downloads the target model from the cloud and uses the target model. Alternatively, the training engine 25 may obtain the target model through training on the cloud and use the target model, and the encoding system 10 directly obtains the processing result from the cloud. This is not specifically limited.
[0108] FIG. 1A shows that the source device 12 and the destination device 14 are independent devices. However, the device embodiments may include both the source device 12 and the destination device 14, or may include the functions of both the source device 12 and the destination device 14, that is, may include both the source device 12 or the corresponding function and the destination device 14 or the corresponding function. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function may be implemented using the same hardware and / or software, or using separate hardware and / or software, or any combination thereof.
[0109] As will be apparent to those skilled in the art based on the description, the presence of different units and the division of functions into different units in the source device 12 and / or the destination device 14 shown in FIG. 1A may vary depending on the actual device and application.
[0110] The encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30), or both the encoder 20 and decoder 30, may be implemented by a processing circuit such as, for example, one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated, or any combination thereof, as shown in FIG. 1B. The encoder 20 can be implemented via the processing circuit 46 to embody various modules as described for the encoder 20 in FIG. 2 and / or any other encoder system or subsystem described herein. The decoder 30 can be implemented via the processing circuit 46 to embody various modules as described for the decoder 30 in FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit 46 can be configured to perform various operations as described later. As shown in FIG. 5, when the technology is partially implemented in software, the device may store software instructions in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to execute the technology of this application. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of an encoder / decoder (CODEC) in a single device, as shown in FIG. 1B.
[0111] The source device 12 and the destination device 14 may include any of a wide range of devices, such as any type of handheld device or fixed device, for example, a notebook or laptop computer, a mobile phone, a smartphone, a tablet, or a tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video gaming console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiving device, a broadcast transmitting device, etc., and may or may not use an operating system, or may use any type of operating system. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0112] In some cases, the video encoding system 10 shown in FIG. 1A is merely an example. The technology of the present application can be applied to video encoding settings (e.g., video encoding or video decoding) that do not necessarily involve data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory or transmitted over a network. A video encoding device may encode data and store the encoded data in memory, and / or a video decoding device may retrieve data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode data into memory and / or retrieve data from memory and decode the data.
[0113] FIG. 1B is an exemplary block diagram of a video encoding system 40 according to the present application. The video encoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video encoder / decoder implemented by a processing circuit 46), an antenna 42, one or more processors 43, one or more memories 44, and / or a display device 45.
[0114] As shown in FIG. 1B, the imaging device 41, the antenna 42, the processing circuit 46, the video encoder 20, the video decoder 30, the processor 43, the memory 44, and / or the display device 45 can communicate with each other. In different examples, the video encoding system 40 may include only the video encoder 20 or only the video decoder 30.
[0115] In some examples, the antenna 42 may be configured to transmit or receive an encoded bitstream of video data. Further, in some examples, the display device 45 may be configured to present video data. The processing circuit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, and the like. The video encoding system 40 may include an optional processor 43. The optional processor 43 may similarly include application-specific integrated circuit (ASIC) logic, a graphics processing unit, a general-purpose processor, and the like. In addition, the memory 44 may be any type of memory, for example, volatile memory (e.g., static random access memory (SRAM) or dynamic random access memory (DRAM)) or non-volatile memory (e.g., flash memory). In a non-limiting example, the memory 44 may be implemented by cache memory. In another example, the processing circuit 46 may include a memory (e.g., cache) to implement an image buffer.
[0116] In some examples, the video coder 20 implemented using logic circuitry may include an image buffer (implemented, for example, by processing circuitry 46 or memory 44) and a graphics processing unit (implemented, for example, by processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may be included in the video coder 20 implemented by processing circuitry 46 to implement various modules described with reference to FIG. 2 and / or any other coder system or subsystem described herein. The logic circuitry may be configured to perform the various operations described herein.
[0117] In some examples, the video decoder 30 may also be implemented similarly by processing circuitry 46 to implement various modules described with reference to the video decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, the video decoder 30 implemented using logic circuitry may include an image buffer (implemented by processing circuitry 46 or memory 44) and a graphics processing unit (implemented, for example, by processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may be included in the video decoder 30 implemented by processing circuitry 46 to implement various modules described with reference to FIG. 3 and / or any other decoder system or subsystem described herein.
[0118] In some examples, antenna 42 may be configured to receive an encoded bitstream of video data. As described, the encoded bitstream may include data related to video frame encoding described herein, indicators, index values, mode selection data, etc., for example, data related to encoded partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators (as described), and / or data defining the encoded partitions). Video encoding system 40 may further include a video decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. Display device 45 is configured to present video frames.
[0119] In this embodiment of the present application, it should be understood that for the examples described with reference to video encoder 20, video decoder 30 may be configured to perform the reverse process. With respect to the signaling of syntax elements, video decoder 30 may be configured to receive and analyze such syntax elements and, in response, decode the associated video data. In some examples, video encoder 20 may perform entropy encoding on the syntax elements to obtain an encoded video bitstream. In such examples, decoder 30 may analyze the syntax elements and decode the associated video data accordingly.
[0120] For ease of explanation, embodiments of the present invention are described with reference to the Versatile Video Coding (VVC) reference software developed by the Joint Collaboration Team on Video Coding (JCT-VC) constituted by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG), or the reference software of the High-Efficiency Video Coding (HEVC). Those skilled in the art will understand that the embodiments of the present application are not limited to HEVC or VVC.
[0121] Encoder and encoding method FIG. 2 is an exemplary block diagram of a video encoder 20 according to the present application. In the example of FIG. 2, the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output terminal (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a splitting unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.
[0122] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 are said to form the forward signal path of the encoder 20. The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are said to form the reverse signal path of the encoder. The reverse signal path of the encoder 20 corresponds to the signal path of the decoder (see the decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer 230, the inter prediction unit 244, and the intra prediction unit 254 are further said to form the "built-in decoder" of the video encoder 20.
[0123] Image and image segmentation (image and block) The encoder 20 may be configured to receive an image (or image data) 17, for example, an image within a sequence of images forming a video or video sequence, via, for example, the input terminal 201. The received image or image data may also be pre-processed image (or pre-processed image data) 19. For simplicity, image 17 is used in the following description. Image 17 may also be referred to as the current image or the image to be encoded (especially in video encoding to distinguish the current image from other images, for example, previously encoded and / or decoded images of the same video sequence, i.e., the video sequence including the current image).
[0124] (Digital) images may be considered or may be regarded as a two-dimensional array or matrix containing samples with intensity values. Samples within the array may also be referred to as pixels or pels (short forms of picture elements). The number of pixels in the array or image in the horizontal and vertical directions (or axes) determines the size and / or resolution of the image. For color representation, three color components are usually used. Specifically, an image may be represented as three sample arrays or may include an array of three samples. In the RGB format or color space, an image includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space, such as YCbCr, which includes a luminance component indicated by Y (in some cases, indicated by L) and two chrominance components indicated by Cb and Cr. The luminance component Y represents luminance or gray-level intensity (for example, these two are the same in a grayscale image), and both chrominance components Cb and Cr represent chrominance or color information components. Thus, an image in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). An image in the RGB format may be converted or transformed into an image in the YCbCr format, and vice versa. This process is also referred to as color conversion or color transformation. If an image is monochrome, the image may include only a luminance sample array. Thus, an image may be, for example, a monochrome-formatted luminance sample array, or a luminance sample array and two corresponding chrominance sample arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0125] In one embodiment, an embodiment of the video coder 20 can include a splitting unit 262 configured to split the image 17 into a plurality of (usually non-overlapping) image blocks 203 (which may also be abbreviated as block 203 or CTU 203). These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), coding tree blocks (CTB), or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The splitting unit 262 may use the same block size for all images in the video sequence and the corresponding grid that defines the block size, or may be configured to change the block size between images or subsets or groups of images to split each image into corresponding blocks.
[0126] In other embodiments, the video coder may be configured to directly receive the blocks 203 of the image 17, such as one, some, or all of the blocks that form the image 17. The image block 203 may also be referred to as the current image block or the image block to be coded.
[0127] Similar to the image 17, the image block 203 may also be or be regarded as a two-dimensional array or matrix that has a smaller dimension than the image 17 but includes samples with intensity values (sample values). In other words, depending on the color format applied, the block 203 may include one sample array (e.g., a luminance array in the case of a monochrome image 17, a luminance or chrominance array in the case of a color image), or three sample arrays (e.g., in the case of a color image 17, one luminance array and two chrominance arrays), or any other number and / or type of arrays. The number of samples in the block 203 in the horizontal and vertical directions (or axes) defines the size of the block 203. Thus, the block may be an array of M×N (M columns × N rows) samples, an array of M×N transform coefficients, etc.
[0128] In one embodiment, a video coder 20 as shown in FIG. 2 may be configured to encode the image 17 block by block, for example, to encode and predict each block 203.
[0129] In one embodiment, a video coder 20 as shown in FIG. 2 may be further configured to divide and / or encode the image using slices (also referred to as video slices), and the image may be divided or encoded using one or more slices (usually non-overlapping). Each slice may include one or more blocks (e.g., coding tree units CTU) or one or more groups of blocks (e.g., tiles in the H.265 / HEVC and VVC standards, and bricks in the VVC standard).
[0130] In one embodiment, a video coder 20 as shown in FIG. 2 may be further configured to divide and / or encode the image by using slice / tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). The image may be divided or encoded using one or more slice / tile groups (usually non-overlapping), and each slice / tile group may include, for example, one or more blocks (e.g., CTU) or one or more tiles. Each tile may be rectangular or another shape and may include one or more complete blocks or fragment blocks (e.g., CTU).
[0131] Residual calculation The residual calculation unit 204 is configured to calculate the residual block 205 based on the image block 203 and the predicted block 265 (further details regarding the predicted block 265 will be provided later) by subtracting the sample values of the predicted block 265 from the sample values of the image block 203 for each sample (per pixel), for example, to obtain the residual block 205 in the pixel region.
[0132] Transformation The transformation processing unit 206 is configured to apply a transformation such as a discrete cosine transform (DCT) or a discrete sine transform (DST) to the sample values of the residual block 205 in order to obtain the transformation coefficients 207 of the transformation region. The transformation coefficients 207 may also be referred to as transformation residual coefficients and represent the residual block 205 in the transformation region.
[0133] The transformation processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the transformation defined in H.265 / HEVC. Compared with the orthogonal DCT transform, such an integer approximation is usually scaled based on the coefficients. In order to maintain the norm of the residual block processed using the forward transform and the inverse transform, an additional scale factor is applied as part of the transformation process. The scale factor is usually selected based on several constraints. For example, the scale factor is a trade-off between a power of 2 of the shift operation, the bit depth of the transformation coefficients, and the accuracy and implementation cost. For example, a specific scale factor is specified for the inverse transform by the inverse transformation processing unit 212 on the encoder 20 side (and is specified for the corresponding inverse transform by the inverse transformation processing unit 312 etc. on the decoder 30 side). Correspondingly, a scale factor corresponding to the forward transform may be specified by the transformation processing unit 206 on the encoder 20 side.
[0134] In one embodiment, the video encoder 20 (correspondingly, the transformation processing unit 206) is configured to output transformation parameters such as one or more types of transformations, for example, so that the video decoder 30 can receive the transformation parameters and use them for decoding. For example, the transformation parameters are directly output after being encoded or compressed by the entropy encoding unit 270, or are configured to output the transformation parameters.
[0135] Quantization The quantization unit 208 is configured to quantize the transform coefficient 207 to obtain a quantized transform coefficient 209, for example, by applying scalar quantization or vector quantization. The quantized transform coefficient 209 may also be referred to as a quantized residual coefficient 209.
[0136] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, if n is greater than m, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization. By adjusting the quantization parameter (QP), the degree of quantization may be changed. For example, in the case of scalar quantization, different scales may be used to implement finer or coarser quantization. A smaller quantization step corresponds to finer quantization, and a larger quantization step corresponds to coarser quantization. The appropriate quantization step may be indicated by the quantization parameter (QP). For example, the quantization parameter may be an index to a predetermined set of appropriate quantization steps. For example, a smaller quantization parameter may correspond to finer quantization (a smaller quantization step), and a larger quantization parameter may correspond to coarser quantization (a larger quantization step), or vice versa. Quantization can include division by the quantization step, and for example, the corresponding inverse quantization and / or inverse-inverse quantization by the inverse quantization unit 210 can include multiplication by the quantization step. In embodiments according to some standards such as HEVC, the quantization parameter may be used to determine the quantization step size. Generally, the quantization step can be calculated by a fixed-point approximation of an equation involving division, based on the quantization parameter. Additional scale factors may be introduced for quantization and inverse quantization to restore the norm of the residual block, and the norm of the residual block may be changed due to the scale used in the equation for the quantization step and the fixed-point approximation of the quantization parameter. In an exemplary embodiment, the scale of the inverse transform can be combined with the scale of the inverse quantization. Alternatively, a customized quantization table may be used and signaled, for example, in a bitstream from the encoder to the decoder. Quantization is a lossy operation, and a larger quantization step indicates a larger loss.
[0137] In one embodiment, the video coder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP). For example, after the quantization parameter is encoded or compressed by the entropy coder 270, the quantization parameter may be directly output, or the quantization parameter may be output. As a result, for example, the video decoder 30 may receive the quantization parameter and use it for decoding.
[0138] Inverse quantization The inverse quantization unit 210 is configured to apply an inverse quantization of the quantization unit 208 with respect to the quantization coefficients, for example, by applying an inverse scheme of the quantization scheme applied by the quantization unit 208 based on or using the same quantization step as the quantization unit 208, to obtain inverse quantization coefficients 211. The inverse quantization coefficients 211 may also be referred to as inverse quantization residual coefficients 211 and correspond to the transform coefficients 207. However, the inverse quantization coefficients 211 are typically not the same as the transform coefficients due to losses caused by quantization.
[0139] Inverse transform The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or the corresponding inverse quantization coefficients 211) within the sample region. The reconstructed residual block 213 may also be referred to as the transform block 213.
[0140] Reconstruction The reconstruction unit 214 (for example, the adder 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the predicted block 265 by adding, for example, the sample values of the reconstructed residual block 213 and the sample values of the predicted block 265, to obtain a reconstructed block 215 in the pixel region.
[0141] Filtering The loop filter unit 220 (also abbreviated as the "loop filter" 220) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or is typically configured to filter the reconstructed samples to obtain filtered sample values. For example, the loop filter unit is configured to perform smooth pixel conversion or improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and then an ALF filter. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop rescaler) is added. This process is performed before deblocking. In another example, the deblocking filter process may be applied to internal sub-block edges such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. The loop filter unit 220 is shown as a loop filter in FIG. 2, but in other forms, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221.
[0142] In one embodiment, the video coder 20 (correspondingly, the loop filter unit 220) may be configured to output loop filter parameters (such as SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly output the loop filter parameters, or output the loop filter parameters after entropy encoding is performed on the loop filter parameters by the entropy coding unit 270. As a result, for example, the decoder 30 can receive and use the same loop filter parameters or different loop filter parameters for decoding.
[0143] Decoded image buffer The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data used by the video coder 20 during video data encoding. The DPB 230 may be formed by any one of various memory devices such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store previously reconstructed and filtered blocks 221, such as other previously filtered blocks, for example, the same current picture or a different picture, a previously reconstructed picture, etc., and may also provide, for example, for inter prediction, a fully previously reconstructed, i.e., decoded picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer 230 can be further configured to store one or more unfiltered reconstructed blocks 215, or generally, unfiltered reconstructed samples, for example, unfiltered reconstructed blocks 215 that have not been filtered by the loop filter unit 220, or other reconstructed blocks or samples for which no other processing has been performed.
[0144] Mode Selection (Partitioning and Prediction) The mode selection unit 260 includes a splitting unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain the original image data, such as the original block 203 (the current block 203 of the current image 17), and the reconstructed image data, such as that of the same (current) image and / or one or more previously decoded images, from, for example, the reconstructed sample or block that has been filtered and / or not filtered from the decoded image buffer 230 or other buffer (e.g., a column buffer not shown in the figure). The reconstructed image data is used as reference image data for prediction, such as inter prediction or intra prediction, to obtain the predicted block 265 or the predicted value 265.
[0145] The mode selection unit 260 may be configured to determine or select a split for the current block prediction mode (including no split) and the prediction mode (e.g., the intra or inter prediction mode), and generate a corresponding predicted block 265 to be used for the calculation of the residual block 205 and the reconstruction of the reconstructed block 215.
[0146] In one embodiment, the mode selection unit 260 may be configured to select split and prediction modes (e.g., supported by or available to the mode selection unit 260) that provide the best match or minimum residual (where the minimum residual refers to better compression for transmission or storage), or that provide the minimum signaling overhead (where the minimum signaling overhead refers to better compression for transmission or storage), or where the minimum residual and minimum signaling overhead are considered or balanced in the prediction mode. The mode selection unit 260 may be configured to determine the split mode and the prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion optimization. Terms such as "best," "lowest," "optimal," etc. as used herein do not necessarily generally mean "the best," "the lowest," "the optimal," but may refer to situations where the criteria for termination or selection are met. For example, values that exceed or fall below a threshold or other limit may be a "second best choice" but reduce complexity and processing time.
[0147] In other words, the splitting unit 262 may be configured to repeatedly use, for example, quad-tree partitioning (QT), binary-tree partitioning (BT), or triple-tree partitioning (TT), or any combination thereof, to split the images of the video sequence into a sequence of image blocks (also referred to as coding tree units (CTUs)), and the CTU 203 may be further split into smaller block partitions or sub-blocks (which again form blocks), and for example, prediction may be performed for each of the block partitions or sub-blocks. Mode selection may include selection of the tree structure of the split blocks 203, and the prediction mode is applied to each of the block partitions or sub-blocks.
[0148] Next, the splitting (e.g., by the splitting unit 262) and prediction (e.g., by the inter prediction unit 244 and the intra prediction unit 254) performed by the video coder 20 will be described in detail.
[0149] Splitting The splitting unit 262 can partition (or split) the encoding tree unit 203 into smaller partitions, such as small square or rectangular blocks. In the case of an image having three sample arrays, a CTU includes a block of N×N luma samples and two corresponding blocks of chroma samples. The maximum allowable size of the luma block in a CTU is specified to be 128×128 in the developing Versatile Video Coding (VVC), but in the future it may be specified to be, for example, 256×256 instead of 128×128. The CTUs of an image may be clustered / grouped as slice / tile groups, tiles, or bricks. A tile covers a rectangular region of the image, and a tile may be divided into one or more bricks. A brick contains multiple CTU rows within a tile. A tile that is not divided into multiple bricks may be referred to as a brick. However, a brick is a true subset of a tile and thus is not referred to as a tile. In VVC, two modes of tile groups are supported: the raster scan slice / tile group mode and the rectangular slice mode. In the raster scan tile group mode, a slice / tile group includes a sequence of tiles in the tile raster scan of the image. In the rectangular slice mode, a slice includes several bricks of the image that collectively form a rectangular region of the image. The bricks within a rectangular slice are arranged in the order of the brick raster scan of the slice. These small blocks (which may also be referred to as sub-blocks) can be further divided into smaller divisions. This is also referred to as tree splitting or hierarchical tree splitting. For example, a root block at root tree level 0 (hierarchical level 0, depth 0) can be recursively divided into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchical level 1, depth 1). These blocks may be further divided into two or more blocks at the next lower level, such as at tree level 2 (hierarchical level 2, depth 2), until the splitting is terminated (for example, because an end criterion is met, such as reaching the maximum tree depth or the minimum block size).Blocks that are not further divided are also referred to as leaf blocks or leaf nodes of the tree. A tree divided into two partitions is called a binary-tree (BT), a tree divided into three partitions is called a ternary-tree (TT), and a tree divided into four partitions is called a quad-tree (QT).
[0150] For example, a Coding Tree Unit (CTU) may be a Coding Tree Block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of an image having three sample arrays, a CTB of samples of a black-and-white image, or a CTB of samples of an image encoded using three separate color planes and syntax structures (used to encode the samples), or may include them. Correspondingly, a Coding Tree Block (CTB) can be a block of N×N samples. By setting N to a specific value, a CTB can be obtained by dividing the component. This is the division. A Coding Unit (CU) may be a coding block of luminance samples, two corresponding coding blocks of chrominance samples of an image having three sample arrays, a coding block of samples of a black-and-white image, or a coding block of samples of an image encoded by using three separate color planes and syntax structures (used to encode the samples), or may include them. Correspondingly, a Coding Block (CB) can be a block of M×N samples. M and N can be set to specific values to divide the CTB into coding blocks. This is the division.
[0151] In an embodiment, for example, according to HEVC, a coding tree unit (CTU) may be divided into a plurality of CUs by using a quadtree structure represented as a coding tree. A decision on whether to encode an image region using inter (temporal) prediction or intra (spatial) prediction is made at the leaf CU level. Each leaf CU can be further divided into one, two, or four PUs based on the PU split type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder in PU units. After a residual block is obtained by applying a prediction process based on the PU split type, the leaf CU can be divided into transform units (TUs) based on another quadtree structure similar to the coding tree of the CU.
[0152] In an embodiment, for example, according to the latest video coding standard currently under development (referred to as Versatile Video Coding (VVC)), a combined quadtree nested multi-type tree (e.g., a binary tree and a ternary tree) divides the segmentation structure used to divide coding tree units. In the coding tree structure of a coding tree unit, the CU can be square or rectangular. For example, a coding tree unit (CTU) is first divided by a quadtree structure. Next, the quadtree leaf nodes are further divided by a multi-type tree structure. The multi-type tree structure has four split types: vertical binary tree split (SPLIT_BT_VER), horizontal binary tree split (SPLIT_BT_HOR), vertical ternary tree split (SPLIT_TT_VER), and horizontal ternary tree split (SPLIT_TT_HOR). The multi-type tree leaf node is referred to as a coding unit (CU), and as long as the CU is not overly large for the maximum transform length, this segmentation is used for prediction and transform processing without further division. This means that in most cases, the CU, PU, and TU have the same block size in a coding block structure where the quadtree is nested with a multi-type tree. An exception occurs when the maximum supported transform length is shorter than the width or height of the color component of the CU. A unique signaling mechanism for dividing or segmenting information in a coding structure where the quadtree is nested with a multi-type tree is formulated in VVC. In the signaling mechanism, the coding tree unit (CTU) is treated as the root of the quadtree and is first divided by the quadtree structure. Each quadtree leaf node (when it is large enough as possible) is then further divided by the multi-type tree structure. In the multi-type tree structure, a first flag (mtt_split_cu_flag) indicates whether the node is further divided. If the node is further divided, a second flag (mtt_split_cu_vertical_flag) indicates the split direction, and then a third flag (mtt_split_cu_binary_flag) indicates whether the split is a binary tree split or a ternary tree split.Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the decoder can derive the multi-type tree splitting mode (MttSplitMode) of the CU based on predefined rules or tables. Note that in certain designs, such as the 64×64 luma block and 32×32 chroma pipeline design in a VVC hardware decoder, as shown in FIG. 6, when the width or height of the luma coding block is greater than 64, TT splitting is not permitted. When the width or height of the chroma coding block is greater than 32, TT splitting is also not permitted. In a pipeline design, an image is divided into a plurality of virtual pipeline data units (VPDUs), and a VPDU is defined as a non-overlapping unit within the image. In a hardware decoder, consecutive VPDUs are processed simultaneously in multiple pipeline stages. The VPDU size is approximately proportional to the buffer size in most pipeline stages. Therefore, it is necessary to maintain a small VPDU size. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) splitting and binary tree (BT) splitting can result in an increase in the VPDU size.
[0153] In addition, note that when a part of the tree node block crosses the lower or right image boundary, the tree node block is forcibly split until all samples of all coded CUs are located inside the image boundary.
[0154] For example, the intra sub-partition (ISP) tool can vertically or horizontally split a luma intra-predicted block into two or four sub-partitions based on the block size.
[0155] In one example, the mode selection unit 260 of the video encoder 20 can be configured to perform any combination of the aforementioned splitting techniques.
[0156] As described above, the video coder 20 is configured to determine or select the best or optimal prediction mode from among a (predetermined) set of prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0157] Intra prediction The set of intra prediction modes may include 35 different intra prediction modes, such as non - directional modes like the DC (or average) mode and the planar mode, or directional modes such as those defined in HEVC, or may include 67 different intra prediction modes, such as non - directional modes like the DC (or average) mode and the planar mode, or directional modes such as those defined in VVC. For example, some of the conventional angular intra prediction modes are adaptively replaced with wide - angle intra prediction modes for non - square blocks as defined in VVC. In another example, only the long side is used to calculate the average of non - square blocks in order to avoid the division operation for DC prediction. Additionally, the result of intra prediction in the planar mode may be further modified by using the position dependent intra prediction combination (PDPC) method.
[0158] The intra prediction unit 254 is configured to use the reconstructed samples of neighboring blocks of the same current picture to generate an intra - predicted block 265 based on the intra prediction mode of the set of intra prediction modes.
[0159] The intra prediction unit 254 (or usually the mode selection unit 260) is further configured to output the intra prediction parameters (or usually the information indicating the intra prediction mode selected for the block) to be sent to the entropy encoding unit 270 in the form of the syntax element 266 included in the encoded image data 21. As a result, the video decoder 30 can perform operations such as receiving and using the prediction parameters for decoding.
[0160] Inter prediction In a possible embodiment, the set of inter prediction modes depends on the available reference images (i.e., images that have been at least partially decoded previously and stored, for example, in the DPB 230) and other inter prediction parameters, such as whether the entire reference image or only a part of it is used to search for the reference block that best matches the search window area around the current block area of the reference image, and / or whether pixel interpolation, such as half / semi pel, quarter pel, and / or 1 / 16 pel interpolation, is applied.
[0161] In addition to the aforementioned prediction modes, the skip mode and / or the direct mode may be further applied.
[0162] For example, the merge candidate list in the extended merge prediction mode sequentially includes five types of candidates: spatial MVPs from spatially neighboring CUs, temporal MVPs from collocated CUs, history-based MVPs from the FIFO table, pairwise average MVPs, and zero MVPs. Bi-directional merge-based decoder side motion vector refinement (DMVR) can be used to improve the accuracy of the MV in the merge mode. The merge mode with MVD (MMVD) is derived from the merge mode with a motion vector difference. The MMVD flag is sent immediately after the skip flag and the merge flag are sent, and specifies whether the MMVD mode is used for the CU. An adaptive motion vector resolution (AMVR) scheme at the CU level may be used. AMVR supports the encoding of the MVD of the CU with different precisions. The current CU's MVD can be adaptively selected based on the prediction mode of the current CU. When the CU is encoded in the merge mode, a combined inter / intra prediction (CIIP) mode may be applied to the current CU. To achieve CIIP prediction, a weighted average is performed on the inter prediction signal and the intra prediction signal. For affine motion compensation prediction, the affine motion field of the block is described using the motion information of the motion vectors of two control points (four parameters) or three control points (six parameters). Subblock-based temporal motion vector prediction (SbTMVP) is similar to the temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vectors of the sub-CUs in the current CU. The bi-directional optical flow (BDOF), previously called BIO, is a simpler version with much less required computation, especially regarding the number of multiplications and the values of the multipliers.In the triangular partitioning mode, the CU is evenly divided into two triangular parts by diagonal partitioning and non - diagonal partitioning. Further, the bi - prediction mode is extended beyond simple averaging to enable weighted averaging of two prediction signals.
[0163] The inter - prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (neither of which is shown in FIG. 2). The motion estimation unit may be configured to receive or obtain, for motion estimation, the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or more previously reconstructed blocks, for example, the reconstructed blocks of one or more other / different previously decoded images 231. For example, the video sequence can include the current image and the previously decoded image 231. In other words, the current image and the previously decoded image 231 are part of or can form a sequence of images that form the video sequence.
[0164] For example, the encoder 20 may be configured to select one reference block from a plurality of reference blocks of the same or different images of a plurality of other images, and provide an offset (spatial offset) between the reference image (or reference image index) and / or the position (x and y coordinates) of the reference block and the position of the current block to the motion estimation unit as an inter - prediction parameter. This offset is also referred to as a motion vector (MV).
[0165] The motion compensation unit is configured to, for example, obtain an inter prediction parameter and perform an inter prediction based on or using the inter prediction parameter to obtain an inter-predicted block 246. The motion compensation performed by the motion compensation unit can include extracting or generating a predicted block based on the motion / block vector determined by motion estimation, and can further include performing interpolation for sub-pixel accuracy. Interpolation filtering generates additional pixel samples from known pixel samples and can thus potentially increase the amount of candidate predicted blocks that may be used for encoding an image block. When receiving the motion vector corresponding to the PU of the current image block, the motion compensation unit may identify the predicted block indicated by the motion vector in one of the reference image lists.
[0166] When decoding an image block of a video slice for use by the video decoder 30, the motion compensation unit may further generate syntax elements regarding the block and the video slice. In addition to or instead of the slice and corresponding syntax elements, tile groups and / or tiles and corresponding syntax elements may be generated or used.
[0167] Entropy coding The entropy encoding unit 270 is configured to apply, for example, an entropy encoding algorithm or method (e.g., variable length coding (VLC) method, context adaptive VLC (CAVLC) method, arithmetic coding method, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method or technique) to the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters, and / or other syntax elements so that, for example, the encoded image data 21 can be output via the output terminal 272 in the form of an encoded bitstream 21 for the video decoder 30 to receive and use the parameters for decoding. The encoded bitstream 21 may be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.
[0168] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal without involving the transform processing unit 206 for some blocks or frames. In other embodiments, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 combined in a single unit.
[0169] Decoder and Decoding Method FIG. 3 is an exemplary diagram of the structure of a video decoder 30 according to the present application. The video decoder 30 is configured to receive encoded image data 21 (e.g., encoded bitstream 21) encoded by an encoder 20 and obtain a decoded image 331. The encoded image data or bitstream includes information for decoding the encoded image data, e.g., data representing image blocks of an encoded video slice (and / or tile group or tile) and associated syntax elements.
[0170] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be a motion compensation unit or may include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding process generally inverse to the encoding process described with respect to the video encoder 20 of FIG. 2.
[0171] As described for the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer DPB 230, the inter prediction unit 244, and the intra prediction unit 254 are further referred to as forming the "built-in decoder" of the video encoder 20. Thus, the function of the inverse quantization unit 310 may be the same as that of the inverse quantization unit 210, the function of the inverse transform processing unit 312 may be the same as that of the inverse transform processing unit 212, the function of the reconstruction unit 314 may be the same as that of the reconstruction unit 214, the function of the loop filter 320 may be the same as that of the loop filter 220, and the function of the decoded picture buffer 330 may be the same as that of the decoded picture buffer 230. Thus, the descriptions provided for the corresponding units and functions of the video encoder 20 are correspondingly applicable to the corresponding units and functions of the video decoder 30.
[0172] Entropy decoding The entropy decoding unit 304 analyzes the bit stream 21 (or, generally, the encoded image data 21), and, for example, performs entropy decoding on the encoded image data 21 to obtain, for example, quantization coefficients 309 and / or decoded encoded parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference image index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or method corresponding to the encoding method of the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360 and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received or used.
[0173] Inverse quantization The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or general information related to inverse quantization) and quantized coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the quantized coefficients 309 decoded based on the quantization parameter to obtain inverse quantized coefficients 311. The inverse quantized coefficients 311 may also be referred to as transform coefficients 311. The inverse quantization process may include the use of quantization parameters calculated by the video coder 20 for each video block within a video slice to determine the degree of inverse quantization that needs to be applied, which is the same as the degree of quantization.
[0174] Inverse transformation The inverse transformation processing unit 312 may be configured to receive the inverse quantized coefficients 311, which are also referred to as transform coefficients 311, and apply a transformation to the inverse quantized coefficients 311 to obtain a reconstructed residual block 313 within the pixel region. The reconstructed residual block 313 may also be referred to as a transform block 313. The transformation may be an inverse transformation, such as an inverse DCT, inverse DST, inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may be further configured to receive transformation parameters or corresponding information from the encoded image data 21 (e.g., by parsing and / or decoding, e.g., by the entropy decoding unit 304) to determine the transformation to be applied to the inverse quantized coefficients 311.
[0175] Reconstruction The reconstruction unit 314 (e.g., adder 314) is configured to add the reconstructed residual block 313 to the predicted block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the predicted block 365, to obtain a reconstructed block 315 in the pixel region.
[0176] Filtering The loop filter unit 320 (either in or after the encoding loop) filters the reconstructed block 315 to obtain a filtered block 321 and is configured to smooth pixel shift or improve video quality. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and then an ALF filter. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is executed before deblocking. In another example, the deblocking filter process may be applied to internal sub-block edges such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. The loop filter unit 320 is shown as a loop filter in FIG. 3, but in other forms, the loop filter unit 320 may be implemented as a post-loop filter.
[0177] Decoded image buffer Next, the buffered block 321 of the image is stored in the decoded image buffer 330, and the decoded image buffer 330 stores the decoded image 331 as a reference image, and the reference image is used for subsequent motion compensation of other images and / or for outputting respective displays.
[0178] The decoder 30 is configured to output the decoded image 331, for example, via the output terminal 332, for presentation or viewing by the user.
[0179] Prediction The inter prediction unit 344 may be functionally identical to the inter prediction unit 244 (particularly the motion compensation unit), and the intra prediction unit 354 may be functionally identical to the intra prediction unit 254. It performs splitting or split determination and prediction based on the splitting parameters and / or prediction parameters, or each piece of information received from the encoded image data 21 (for example, by analysis and / or decoding by the entropy decoder unit 304). The mode application unit 360 may be configured to perform prediction (intra or inter prediction) on each block based on the reconstructed image, or block, or corresponding sample (filtered or unfiltered) in order to obtain the predicted block 365.
[0180] When a video slice is coded as an intra coded (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a predicted block 365 of an image block of the current video slice based on the indicated intra prediction mode and data from previously decoded blocks of the current picture. When the video picture is coded as an inter coded (e.g., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a predicted block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoder unit 304. In the case of inter prediction, the predicted block may be generated from one reference picture within one reference picture list. The video decoder 30 may configure the reference frame lists, i.e., list 0 and list 1, by using a default construction technique based on the reference pictures stored in the DPB 330. The same or similar process may be applied to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices), e.g., the video may be coded using I, P, or B tile groups and / or tiles.
[0181] The mode application unit 360 is configured to determine prediction information regarding a video block of a current video slice by analyzing motion vectors and other syntax elements, and to create a predicted block regarding the currently decoded video block using the prediction information. For example, the mode application unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to encode a video block of a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information regarding one or more of the reference picture lists regarding the slice, a motion vector for each inter-coded video block of the slice, an inter prediction state for each inter-coded video block of the slice, and other information for decoding a video block within the current video slice. The same or a similar process may be applied to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices), e.g., the video may be encoded using I, P, or B tile groups and / or tiles.
[0182] In one embodiment, a video decoder 30 as shown in FIG. 3 may be further configured to divide and / or decode an image using slices (also referred to as video slices), and the image may be divided or decoded using one or more slices (usually non-overlapping). Each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., tiles in the H.265 / HEVC / VVC standard and bricks in the VVC standard).
[0183] In one embodiment, the video decoder 30 shown in FIG. 3 may be further configured to divide and / or decode an image by using slices / tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). The image may be divided or decoded using one or more slices / tile groups (usually non-overlapping), and each slice / tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be rectangular or another shape and may include one or more complete blocks or fragmentary blocks (e.g., CTUs).
[0184] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 can directly inverse quantize the residual signal without the inverse transform processing unit 312 for some blocks or frames. In other embodiments, the video decoder 30 may have an inverse quantization unit 310 and an inverse transform processing unit 312 that are combined into a single unit.
[0185] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step can be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations, such as clip or shift, may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0186] Video encoding and decoding have been mainly described in the foregoing embodiments. However, it should be noted that the embodiments of the encoding system 10, the encoder 20, and the decoder 30 described herein and other embodiments can also be used for still image processing or encoding and decoding, that is, the processing or encoding and decoding of a single image that is independent of the preceding or consecutive images in video encoding and decoding. Generally, when image processing is limited to a single image, the inter prediction unit 244 (encoder) and the inter prediction unit 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30, such as the residual calculation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210 / 310, the (inverse) transform processing unit 212 / 312, the splitting unit 262, the intra prediction unit 254 / 354 and / or the loop filter unit 220 / 320, the entropy encoding unit 270, and the entropy decoding unit 304 can also be used for still image processing.
[0187] FIG. 4 is a diagram of a video encoding device 400 according to the present application. The video encoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video encoding device 400 may be a decoder such as the video decoder 30 in FIG. 1A, or an encoder such as the video encoder 20 in FIG. 1A.
[0188] The video encoding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 configured to receive data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) configured to transmit data, and a memory 460 for storing data. The video encoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the inlet port 410, the receiver unit 420, the transmitter unit 440, and the outlet port 450 for the ingress and egress of optical or electrical signals.
[0189] The processing unit 430 is implemented by using hardware or software. The processing unit 430 may be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processing unit 430 communicates with the inlet port 410, the receiver unit 420, the transmitter unit 440, the outlet port 450, and the memory 460. The processing unit 430 includes an encoding module 470. The encoding module 470 implements the embodiments disclosed above. For example, the encoding module 470 implements, processes, prepares, or provides various encoding operations. Thus, the encoding module 470 provides a significant improvement to the functionality of the video encoding device 400 and results in the switching of the video encoding device 400 to different states. Alternatively, the encoding module 470 is implemented using instructions stored in the memory 460 and executed by the processing unit 430.
[0190] Memory 460 may include one or more disks, tape drives, and solid state drives, and may be used as an overflow data storage device for storing such programs when they are selected for execution and for storing instructions and data read during program execution. Memory 460 may be volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0191] FIG. 5 is a schematic block diagram of an apparatus 500 according to the present application. Apparatus 500 may be used as either or both of the source device 12 and the destination device 14 in FIG. 1A.
[0192] The processor 502 within apparatus 500 may be a central processing unit. Alternatively, processor 502 may be any other type of device or devices capable of manipulating or processing information, whether currently existing or to be developed in the future. The disclosed embodiments may be implemented using a single processor such as processor 502 shown in the figures, but advantages in terms of speed and efficiency can be achieved by using multiple processors.
[0193] In one embodiment, the memory 504 in the apparatus 500 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that are accessed by the processor 502 through the bus 512. The memory 504 may further include an operating system 508 and an application 510, and the application 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application 510 may include applications 1 to N, and further include a video encoding application that executes the methods described in this specification.
[0194] The apparatus 500 may also include one or more output devices such as a display 518. In one example, the display 518 can be a touch-sensitive display that combines a display with a touch-sensitive element configured to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0195] Although the bus 512 of the apparatus 500 is described herein as a single bus, the bus 512 may include multiple buses. Further, the secondary storage device may be directly coupled to the other components of the apparatus 500 or accessed via a network, and may include a single integrated unit such as a memory card, or multiple units such as multiple memory cards. Thus, the apparatus 500 can have various configurations.
[0196] For ease of understanding, the nouns or terms in the embodiments of this application will be first explained and described below. The nouns or terms are also used as part of the content of the present invention.
[0197] 1. Neural Network A neural network (NN) is a machine learning model. A neural network can include neurons. A neuron may be an arithmetic unit that uses x s and an intercept of 1 as inputs, and the output of the arithmetic unit can be as follows: [Number]
[0198] where s = 1, 2, … or n, n is a natural number greater than 1, W s is the weight of x s , b is the bias of the neuron, and f is the activation function of the neuron. The activation function is used to introduce non-linear features into the neural network and convert the input signal of the neuron into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function may be a non-linear function such as ReLU. A neural network is a network formed by connecting a large number of single neurons to each other. Specifically, the output of a neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field may be a region containing multiple neurons.
[0199] 2. Multi-layer perceptron (MLP) An MLP is a simple deep neural network (DNN) (where different layers are fully connected) and is also referred to as a multi-layer neural network. An MLP can be understood as a neural network with many hidden layers. In this specification, there is no special measurement criterion for "many". A DNN is divided based on the position of different layers, and the neural network of a DNN can be divided into three types: an input layer, a hidden layer, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are hidden layers. The layers are fully connected. Specifically, any neuron in the i-th layer is necessarily connected to any neuron in the (i + 1)-th layer. Although a DNN may seem complex, it is not complex in terms of the operations at each layer. Briefly speaking, the operation at each phase is the following linear relational expression.
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0200] 3. Convolutional Neural Network A convolutional neural network (CNN) is a deep neural network with a convolutional structure and is a deep learning architecture. In a deep learning architecture, multi-layer learning is performed at different levels of abstraction according to machine learning algorithms. As a deep learning architecture, a CNN is a feed-forward artificial neural network, and each neuron in the feed-forward artificial neural network can respond to an image input to the feed-forward artificial neural network. A convolutional neural network includes a feature extractor that includes convolutional layers and pooling layers. The feature extractor can be regarded as a filter. The convolution process may be considered to use a trainable filter to perform convolution on an input image or a convolutional feature map.
[0201] A convolutional layer is a neuron layer within a convolutional neural network where a convolution operation is performed on an input signal. The convolutional layer may include a plurality of convolution operators. A convolution operator is also referred to as a kernel. In image processing, a convolution operator functions as a filter that extracts specific information from an input image matrix. A convolution operator may essentially be a weight matrix, and the weight matrix is usually predefined. In the process of performing a convolution operation on an image, the weight matrix typically processes pixels at a granularity level of 1 pixel (or 2 pixels depending on the value of the stride) in the horizontal direction of the input image to extract specific features from the image. The size of the weight matrix needs to be associated with the size of the image. Note that the depth dimension of the weight matrix is the same as the depth dimension of the input image. In the process of performing the convolution operation, the weight matrix spans the entire depth of the input image. Therefore, a single convolution output of a single depth dimension is generated by convolution with a single weight matrix. However, in most cases, a single weight matrix is not used, and multiple weight matrices of the same size (rows × columns), that is, multiple matrices of the same type, are applied. The outputs of the weight matrices are stacked to form the depth dimension of the convolutional image. The dimensions in this specification may be understood to be determined based on the aforementioned "plurality". Different weight matrices may be used to extract different features from an image. For example, one weight matrix is used to extract edge information of an image, another weight matrix is used to extract a specific color of the image, and yet another weight matrix is used to blur unnecessary noise in the image. The sizes (rows × columns) of the multiple weight matrices are the same. The sizes of the feature maps extracted by multiple weight matrices of the same size are also the same, and then, multiple extracted feature maps of the same size are combined to form the output of the convolution operation. The weight values of these weight matrices need to be obtained through a large amount of training during actual application. Each weight matrix containing the weight values obtained through training may be used to extract information from the input image, and as a result, the convolutional neural network performs correct predictions.When a convolutional neural network has multiple convolutional layers, a large number of general features are usually extracted in the initial convolutional layer. The general features may be referred to as low-level features. As the depth of the convolutional neural network increases, the features extracted in subsequent convolutional layers become more complex, for example, high-level semantic features. Features with high-level semantics are more applicable to the problem to be solved.
[0202] The number of training parameters often has to be reduced. Therefore, a pooling layer often has to be introduced periodically after the convolutional layer. One pooling layer may follow one convolutional layer, or one or more pooling layers may follow multiple convolutional layers. During image processing, the pooling layer is only used to reduce the spatial size of the image. The pooling layer may include an average pooling operator and / or a max pooling operator to perform sampling on the input image to obtain a small-sized image. The average pooling operator can be used to perform calculations on the pixel values within a specific range of the image to generate an average value, which is used as the average pooling result. The max pooling operator can be used to select the pixel with the maximum value within a specific range as the max pooling result. In addition, similar to the requirement that the size of the weight matrix in the convolutional layer needs to be associated with the size of the image, the operator in the pooling layer also needs to be associated with the size of the image. The size of the processed image output from the pooling layer can be smaller than the size of the image input to the pooling layer. Each sample in the image output from the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.
[0203] After the processing is executed in the convolutional layer / pooling layer, the convolutional neural network is not ready to output the required output information. The reason is that, as described above, in the convolutional layer / pooling layer, only features are extracted and the parameters obtained from the input image are reduced. However, in order to generate the final output information (required type information or other relevant information), the convolutional neural network needs to use neural network layers to generate the output of one required type or a group of required types. Therefore, the convolutional neural network layer may include multiple hidden layers. The parameters included in the multiple hidden layers can be obtained through pre-training based on the relevant training data of a specific task type. For example, the task type may include image recognition, image classification, and super-resolution image reconstruction.
[0204] Optionally, in the neural network layer, after the multiple hidden layers, an output layer of the entire convolutional neural network follows. The output layer has a loss function similar to categorical cross-entropy, and the loss function is specifically used for calculating the prediction error. When the forward propagation of the entire convolutional neural network is completed, backpropagation is started to update the weight values and biases of each layer described above, and to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network using the output layer and the ideal result.
[0205] 4. Recurrent Neural Network Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, the layers from the input layer to the hidden layer and the output layer are fully connected, and the nodes in each layer are not connected. Such general neural networks solve many problems but still cannot solve many other problems. For example, to predict the next word in a sentence, since the previous and the next words in the sentence are not independent, usually the previous word needs to be used. The reason why RNNs are called recurrent neural networks is that the current output of the sequence is also related to the previous output of the sequence. A specific manifestation is that the network remembers previous information and applies the previous information to the calculation of the current output. Specifically, the nodes in the hidden layer are now connected without being disconnected, and the input to the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNNs can process sequence data of any length. The training of RNNs is similar to that of traditional CNNs or DNNs. The error backpropagation algorithm is also used, but when RNNs are extended, there is a difference in that parameters such as the W of the RNN are shared. This is different from the traditional neural networks described in the above example. In addition, during the use of the gradient descent algorithm, the output at each step depends not only on the network at the current step but also on the network state at some previous steps. The learning algorithm is called the Back propagation Through Time (BPTT) algorithm.
[0206] Why is a recurrent neural network still needed when a convolutional neural network is available? The reason is simple. In a convolutional neural network, elements are independent of each other, and there is a premise that the input and output are also independent, like "dog" and "cat". However, in the real world, many elements are interconnected. For example, stocks change over time. In another example, a person says, "I like traveling, and my favorite place is Yunnan. I will go there if I have the chance." In this specification, people should be able to tell that this person will go to "Yunnan" because people perform inferences from the context. However, how can a machine do this? Here, the RNN comes into play. The RNN is intended to enable a machine to remember like a human. Therefore, the output of the RNN must depend on the current input information and the historical memory information.
[0207] 5. Loss Function In the process of training a deep neural network, the output of the deep neural network is expected to be as close as possible to the actually expected prediction value. Therefore, the predicted value of the current network may be compared with the actually expected target value, and then the weight vectors of each layer of the neural network are updated based on the difference between the predicted value and the target value (certainly, usually there is an initialization process before the first update, specifically, the parameters are pre-configured for all layers of the deep neural network). For example, if the predicted value of the network is large, the weight vector is adjusted to decrease the predicted value, and the adjustment is continuously executed until the deep neural network can predict the actually expected target value, or a value very close to the actually expected target value. Therefore, it is necessary to pre-define "how to obtain the difference between the predicted value and the target value by comparison". This is the loss function or objective function. The loss function and the objective function are important equations for measuring the difference between the predicted value and the target value. The loss function is used as an example. The higher the output value (loss) of the loss function, the greater the difference indicates. Therefore, the learning of the deep neural network is a process of minimizing the loss as much as possible.
[0208] 6. Backpropagation Algorithm The convolutional neural network may correct the value of the parameters of the initial super-resolution model in the training process according to the error backpropagation (BP) algorithm, so that the error loss of reconstructing the super-resolution model becomes smaller. Specifically, the input signal is transferred forward until the error loss occurs at the output, and the parameters of the initial super-resolution model are updated based on the backpropagation error loss information so as to converge the error loss. The backpropagation algorithm is a backpropagation movement centered on the error loss, which is intended to obtain parameters such as the weight matrix of the optimal super-resolution model.
[0209] 7. Adversarial Generative Network A generative adversarial network (GAN) is a deep learning model. The model includes at least two modules: a Generative Model and a Discriminative Model. The two modules learn from each other through a game to generate better outputs. Both the generative model and the discriminative model may be neural networks, specifically, they may be deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Using GAN to generate images is taken as an example. It is assumed that there are two networks, G (Generator) and D (Discriminator). G is a network for generating images. G receives random noise z, uses the noise to generate an image, and the image is denoted as G(z). D is a discriminative network used to identify whether an image is a "real image". The input parameter of D is x, where x represents an image, and the output D(x) represents the probability that x is a real image. If the value of D(x) is 1, it indicates that the image is 100% real. If the value of D(x) is 0, it indicates that the image cannot be a real image. In the process of training a generative adversarial network, the goal of the generative network G is to generate an image as close to a real image as possible to deceive the discriminative network D, and the goal of the discriminative network D is to distinguish the image generated by G from the real image as much as possible. Thus, there is a dynamic "game" process between G and D, that is, the "adversary" in the "generative adversarial network". The final game result, in an ideal state, is that G can generate an image G(z) that is difficult to distinguish from a real image, and for D, it is difficult to identify whether the image generated by G is a real image. Specifically, D(G(z)) = 0.5. In this way, an excellent generative model G is obtained and can be used to generate images.
[0210] 8. Golomb coding In a computer data storage device, a bit is the minimum storage unit, and 1 bit stores a binary code of logical 0 or 1. Generally, numbers are encoded in binary mode. However, in binary mode, different numbers of the same length are recorded, resulting in a large amount of redundant information being generated. For example, the binary codes of the same length for 255 and 1 are 11111111 and 00000001, and the effective bits of the binary codes are 8 and 1. When any data is transmitted in a binary encoding method, a large amount of redundancy is generated, increasing the network load.
[0211] Golomb coding (Exponential-Golomb coding) is a reversible data compression method. This is a variable-length prefix code and is suitable for coding where the occurrence probability of smaller numbers is higher than that of larger numbers. A shorter code length is used to code smaller numbers, and a longer code length is used to code larger numbers. Exponential Golomb coding is a form of Golomb coding and is a variable-length coding method commonly used in audio and video coding standards. All numbers are divided into different groups of equal size. Shorter code lengths are assigned to groups with smaller symbol values, the symbol lengths within the same group are basically equal, and the group size increases exponentially.
[0212] The format of the number num to be encoded obtained by performing k-th order exponential Golomb coding is [MZeros][1][Info], where MZeros is assumed to represent M zeros. The encoding process is as follows.
[0213] i. num is represented in binary mode (if the number of bits is less than k, 0 is added to the most significant bit), and the least significant k bits are removed (if the number of bits is exactly k, num is 0) to obtain the number n.
[0214] ii. M is the number of bits of the binary number obtained by subtracting 1 from the number (n + 1).
[0215] In step iii.i, place the k-bit binary sequence removed in that step in the least significant (n + 1) bits to obtain [1][Info].
[0216] Table 1 shows diagrams of codewords obtained by performing Golomb coding for 0 to 9.
[0217] [Table 1]
[0218] 9. Block Partitioning The HEVC standard introduces a set of basic units based on quadtree recursive partitioning, including Coding Unit (CU), Prediction Unit (PU), and Transform Unit (TU). In the HEVC encoder, the input image is first divided into multiple Coding Tree Units (CTUs) that have the same size and do not overlap with each other. The size of the CTU is determined by the encoder, and the maximum size of the CTU may be 64×64. In the subsequent encoding process, one CTU may be directly used as one CU, or it may be further divided into multiple CUs of smaller sizes in a quadtree recursive partitioning manner. The CU partitioning depth is determined based on the rate-distortion cost obtained by calculation. The encoder compares different partitioning modes to balance and selects the partitioning mode with the lowest rate-distortion cost for encoding. Finally, the CU located at the leaf node of the quadtree is the basic unit for the encoder to perform subsequent prediction, transformation, and encoding. Usually, during block partitioning, shallow partitioning depth and large CUs are used in flat and smooth regions, while deep partitioning depth and small CUs are used in regions with complex colors. In this way, the processing of modules such as the prediction module, transformation module, and quantization module is more convenient and accurate, and the selection of the encoding mode is more adapted to the image features of the video content, thereby effectively improving the encoding efficiency.
[0219] A PU is the basic unit for the coder to perform prediction. It contains all prediction-related information and is obtained by further dividing a CU. For a 2N×2N CU, as shown in FIG. 6, the PU has a total of 8 optional division modes, including 4 symmetric modes of 2N×2N, N×N, 2N×N, and N×2N, and 4 asymmetric modes of 1:3 or 3:1 division. By introducing multiple division modes, the prediction of complex images can be made more accurate.
[0220] A TU is the basic unit for the coder to perform transformation and quantization. The size of a TU depends on the CU to which the TU belongs, and the CU is allowed to be further divided in a quadtree division manner to obtain the TU. The minimum size of a TU is 4×4, and the maximum size can reach 32×32. According to the content characteristics, the coder can flexibly select the best division mode for the TU.
[0221] 10. Intra Block Copy (IBC) IBC is an important tool for screen content coding in the HEVC Screen Content Coding (SCC) standard. This is a block-based prediction technique, and its mechanism is similar to inter prediction or motion compensation. Motion compensation means that for the current prediction unit, the coder finds the optimal matching block in the previously encoded reference image as the predicted value according to the motion search algorithm, and the motion vector (MV) indicates the matching relationship. The difference obtained by subtracting the predicted value from the pixel value of the current prediction unit is used as the prediction residual. The prediction residual is output to the bitstream after being processed by modules such as a transformation module, a quantization module, and an entropy coding module.
[0222] The main difference between IBC and motion compensation is that the reference samples of IBC are obtained from within the current image (the reconstructed part), and there is a block vector (BV) similar to the motion vector to show the block matching relationship. IBC can better process screen content with multiple similar graphics or words within a video frame. That is, the current block can be a block with similar graphics within the current image, and a prediction residual with a pixel value close to 0 can be obtained. This residual occupies a very small bit rate in the bitstream. The encoding process of the syntax structure and related information included in the rest of IBC is almost the same as that of motion compensation.
[0223] This application provides a word-containing image encoding method and a word-containing image decoding method, and solves the problem that the accuracy and efficiency of word content compression may be reduced because pixels of the same word are classified into different CTUs or CUs in the word-containing image encoding and decoding processes.
[0224] Figure 2 illustrates the structure and principle of a standard video encoding module system, and Figure 3 illustrates the structure and principle of a standard video decoding system. In this application, an encoding method for words in the image is added based on this.
[0225] FIG. 7 is an exemplary block diagram of an encoding system 70 according to the present application. As shown in FIG. 7, compared with a standard video encoding and decoding system (e.g., encoding system 10), on the encoder side (e.g., source device 12), in addition to a standard video encoding module (e.g., encoder 20), a word processing module and a word alignment side information encoding module are added, and on the decoder side (e.g., destination device 14), in addition to a standard video decoding module (e.g., decoder 30), a word alignment side information decoding module and a word reconstruction module are added. The added modules are independent of the standard video encoding / decoding process and form an improved encoding system 70 together with the standard video encoding / decoding process.
[0226] The functions of the aforementioned modules are as follows.
[0227] Word processing module 1. Detection: Extract word regions and non-word regions from the original image (i.e., the image to be encoded).
[0228] 2. Segmentation: Obtain complete characters or character lines within the word region by segmentation.
[0229] 3. Adaptive alignment: Align the segmented characters or character lines with a block segmentation grid. In this case, the positions of the aligned characters or character lines are offset with respect to the original positions.
[0230] Word alignment side information encoding module: Encode the word alignment side information (e.g., the original coordinates of the characters, the size of the character box, and the size of the filling box) generated by the word processing module for word reconstruction on the decoder side.
[0231] Standard video encoding module: Perform standard hybrid video encoding separately on the non-word regions and the aligned word regions.
[0232] Standard video decoding module: Performs standard hybrid video decoding on the bitstream to obtain non-word regions and aligned word regions.
[0233] Word alignment side information decoding module: Decodes the bitstream to obtain word alignment side information (e.g., the original coordinates of the characters, the size of the character box, and the size of the filling box).
[0234] Word reconstruction module 1. Word restoration: Based on the word alignment side information, restore the word region at the correct position from the aligned word region.
[0235] 2. Region joining: Merge the word region at the correct position and the non-word region to form a complete reconstructed image.
[0236] Based on the encoding system 70 shown in FIG. 7, the solutions provided in this application will be described below.
[0237] FIG. 8 is a flowchart of a process 800 for a word-containing image encoding method according to an embodiment of the present application. The process 800 can be executed by a video encoder 20. The process 800 is described as a series of steps or operations. It should be understood that the process 800 is not limited to the execution order shown in FIG. 8 and may be executed in various orders and / or simultaneously. Assume that the current image block of the current video frame of a video data stream including a plurality of video frames is encoded by the video encoder 20 using the process 800 including the following steps.
[0238] Step 801: Obtain the word region in the first image, where the word region includes at least one character.
[0239] The first image may be an independent image or any video frame within a video. This is not specifically limited. In the present embodiment of this application, the first image may include a word region. There may be only one or more word regions. The number of word regions may depend on the number of characters, character distribution rules, character sizes, etc. included in the first image. This is not specifically limited. In addition, a word region can include one character, that is, a word region can include only one character, or a word region can include one line of characters, where the line of characters includes multiple characters, or a word region can include one string of characters, where the string of characters includes multiple characters. The aforementioned characters may be Chinese characters, Chinese phonetic characters, English characters, numbers, etc. This is not specifically limited.
[0240] In this embodiment of this application, the word region within the first image can be obtained using the following four methods.
[0241] In the first method, character recognition is performed on the first image to obtain the word region.
[0242] Binarization and dilation can first be performed on the first image. Usually, in a binarized image, the samples corresponding to the word region are white, and the samples corresponding to the non-word region are black. Dilation can be used to handle pixel adhesion between characters. Then, horizontal projection is performed on the image obtained through binarization and dilation, and the number of white samples in the horizontal direction (corresponding to the word region) is calculated. The amount of white samples on the horizontal line with characters is not zero. Therefore, all lines of characters within the image can be obtained based on this.
[0243] For each text line, the following operations are performed: extracting connected components (connected components refer to continuous white sample regions, and one character is a connected component, and correspondingly, one text line can have multiple connected components), and selecting a character box of an appropriate size. When the character box contains only one character, the size of the character box is greater than or equal to the size of the largest connected component within the text line, and all characters within the same text line use character boxes of the same size. When the character box contains one text line or a text string, the size of the character box is greater than or equal to the sum of the sizes of all connected components within the text line or text string. For example, the connected component of character 1 within a text line occupies a pixel space of 6 (width) × 10 (height), and the connected component of character 2 occupies a pixel space of 11 × 8. In this case, at least a character box of 11×10 is selected to cover the two characters, or at least a character box of 17×10 is selected to cover the two characters of the text line. In the present embodiment of the present application, the region covered by the character box in the first image can be referred to as a word region, that is, the word region and the character box refer to the same region range within the first image.
[0244] Therefore, in order to obtain the word region, the first image can be recognized using the method described above.
[0245] In the second method, optical character recognition is performed on the first image to obtain an intermediate word region, the intermediate word region contains multiple text lines, and line splitting is performed on each text line of the intermediate word region to obtain the word region.
[0246] The difference from the first method is that in the second method, the intermediate word region can be obtained by optical character recognition, the intermediate word region can contain multiple text lines, and then line splitting is performed on the intermediate word region to obtain multiple word regions each containing only one text line. This method may be used when the characters are arranged horizontally. For example, multiple characters are arranged from left to right and from top to bottom.
[0247] In the third method, character recognition is performed on the first image to obtain an intermediate word region, the intermediate word region contains a plurality of character strings, and column splitting is performed for each character string in the intermediate word region to obtain word regions.
[0248] The difference from the first method is that in the third method, the intermediate word region can be obtained by character recognition, the intermediate word region can contain a plurality of character lines, and then column splitting is performed on the intermediate word region to obtain a plurality of word regions each containing only one character string. This method may be used when characters are arranged vertically. For example, a plurality of characters are arranged from top to bottom and from left to right.
[0249] In the fourth method, character recognition is performed on the first image to obtain an intermediate word region, the intermediate word region contains a plurality of character lines or a plurality of character strings, and line splitting and column splitting are performed for each character in the intermediate word region to obtain word regions.
[0250] The difference from the first method is that in the fourth method, the intermediate word region can be obtained by character recognition, the intermediate word region can contain a plurality of character lines or a plurality of character strings, and then line splitting and column splitting are performed on the intermediate word region to obtain a plurality of word regions each containing only one character. This method may be used when characters are arranged horizontally or vertically. For example, a plurality of characters are arranged from top to bottom and from left to right. In another example, a plurality of characters are arranged from top to bottom and from left to right.
[0251] It should be noted that in addition to the four methods described above, in this embodiment of the present application, another method may be used to obtain the word region in the first image. This is not specifically limited.
[0252] Step 802: Fill the word region to obtain a word filling region, and the height or width of the word filling region is n times the preset size, where n ≥ 1.
[0253] The preset size may be any of 8, 16, 32, and 64. In this embodiment of the present application, related encoding and decoding methods can be used to encode an image. In the related encoding and decoding methods, the image usually needs to be divided to obtain coding units (CUs), and subsequent processing is performed based on the CUs. The size of the CU may be one of 8, 16, 32, and 64. In this embodiment of the present application, the preset size may be determined based on the size of the CU. For example, when the size of the CU is 8, the preset size may be 8 or 16, or when the size of the CU is 32, the preset size may be 32.
[0254] In this embodiment of the present application, the word filling area can be obtained by filling the word area. Therefore, the height of the word filling area is greater than or equal to the height of the word area, and / or the width of the word filling area is greater than or equal to the width of the word area. In the aforementioned case where the height of the word filling area is equal to the height of the word area and / or the width of the word filling area is equal to the width of the word area, it may mean that the height and / or width of the word area is exactly n times the preset size. Therefore, the word area does not need to be filled, and the subsequent steps are executed by directly using the word area as the word filling area. In addition, as described in step 801, when the word area is obtained, the height and / or width of the word area of the characters in the same row or the same column are kept constant. Correspondingly, when the word filling area is obtained, the height and / or width of the word area in the same row or the same column are also kept constant. In this way, it can be guaranteed that the height of the word filling area coincides with the height of the CU obtained by the division in the related encoding and decoding methods, and / or the width of the word filling area coincides with the width of the CU obtained by the division in the related encoding and decoding methods. In this embodiment of the present application, the aforementioned effect may also be referred to as the word filling area being aligned with the CU obtained by the division in the related encoding and decoding methods.
[0255] In a possible embodiment, the height of the word filling area is n times the preset size. For example, when the height of the word area is 10 and the preset size is 8, the height of the word filling area may be 16 (2 times 8). As another example, when the height of the word area is 10 and the preset size is 16, the height of the word filling area may be 16 (1 time 16).
[0256] In a possible implementation, the width of the word filling area is n times the preset size. For example, when the width of the word area is 17 and the preset size is 8, the width of the word filling area may be 24 (3 times 8). As another example, when the width of the word area is 17 and the preset size is 16, the width of the word filling area may be 32 (2 times 16).
[0257] In a possible implementation, the height of the word filling area is n1 times the preset size, and the width of the word filling area is n2 times the preset size, where n1 and n2 may or may not be equal. For example, when the size of the word area is 6 (width) × 10 (height) and the preset size is 8, the size of the word filling area may be 8 (1 times 8) × 16 (2 times 8). As another example, when the size of the word area is 17 × 10 and the preset size is 16, the size of the word filling area may be 32 (2 times 16) × 16 (1 times 16).
[0258] In this embodiment of the present application, in order to obtain the word filling area, the following three methods can be used to fill the word area.
[0259] In the first method, the height of the word filling area is obtained based on the height of the word area. The height of the word filling area is n times the preset size. The word area is filled vertically in the word area to obtain the word filling area, and the filling height is the difference between the height of the word filling area and the height of the word area.
[0260] As described above, the height of the word filling area is greater than the height of the word area, and the height of the word filling area is n times the preset size. Based on this, the height of the word filling area can be obtained based on the height of the word area. Refer to the above example. Optionally, when the height of the word area is exactly n times the preset size, the word area does not need to be filled. In this case, the height of the word filling area is equal to the height of the word area. After the height of the word filling area is obtained, the word area may be filled in the vertical direction of the word area (for example, the upper side, the middle, or the lower side of the word area, which is not specifically limited), and the filling height is the difference between the height of the word filling area and the height of the word area. For example, when the height of the word filling area is 16 and the height of the word area is 10, the filling height is 6. The pixel value of the filling may be the pixel value of the background pixel other than the character pixel in the word area, or alternatively, the preset pixel value, such as 0 or 255. This is not specifically limited.
[0261] In the second method, the width of the word filling area is obtained based on the width of the word area, the width of the word filling area is n times the preset size, the word area is filled in the horizontal direction of the word area to obtain the word filling area, and the filling width is the difference between the width of the word filling area and the width of the word area.
[0262] As described above, the width of the word filling area is larger than the width of the word area, and the width of the word filling area is n times the preset size. Based on this, the width of the word filling area can be obtained based on the width of the word area. Refer to the above example. Optionally, when the width of the word area is exactly n times the preset size, the word area does not need to be filled. In this case, the width of the word filling area is equal to the width of the word area. After the width of the word filling area is obtained, the word area may be filled in the horizontal direction of the word area (for example, on the left side, in the center, or on the right side of the word area, which is not specifically limited), and the width of the filling is the difference between the width of the word filling area and the width of the word area. For example, when the width of the word filling area is 24 and the width of the word area is 17, the width of the filling is 7. The pixel value of the filling may be the pixel value of the background pixels other than the character pixels in the word area, or alternatively, the preset pixel value, such as 0 or 255. This is not specifically limited.
[0263] The third method may be a combination of the first method and the second method, that is, after the height and width of the word filling area are obtained, in order to obtain the word filling area, the first two methods are respectively used to fill the word area in the vertical and horizontal directions of the word area.
[0264] It can be found that the size of the word filling area obtained by filling has a plurality of relationships with the size of the CU. Therefore, after such a word filling area is divided, the word filling area can cover the complete CU, that is, the case where the word filling area covers a part of the CU does not occur, and the same CU does not cover two word filling areas, thereby implementing the alignment between the word area and the CU during division.
[0265] Step 803: Obtain a second image based on the word filling area.
[0266] In this embodiment of the present application, a new image, that is, a second image, can be obtained based on the word filling area. The second image includes the characters in the original first image. Obtaining the second image includes the following four cases.
[0267] 1. If there is only one word filling area, the word filling area is used as the second image.
[0268] If there is only one word filling area, that word filling area may be directly used as the second image.
[0269] 2. If there are multiple word filling areas and each of the word filling areas contains one line of characters, the multiple word filling areas are joined together from top to bottom to obtain the second image.
[0270] One word filling area contains one line of characters. In this case, the word filling areas to which the multiple lines of characters belong may be joined together from top to bottom based on the distribution of the lines of characters in the first image. Specifically, the word filling area to which the line of characters arranged at the top of the first image belongs is also joined at the top, the word filling area to which the line of characters arranged in the second line of the first image belongs is also joined in the second line, and the rest may be estimated by analogy. It should be noted that in this embodiment of the present application, the joining may alternatively be performed from bottom to top, and the order of joining is not specifically limited.
[0271] 3. If there are multiple word filling areas and each of the word filling areas contains one string of characters, the multiple word filling areas are joined together from left to right to obtain the second image.
[0272] One word filling area contains one string. In this case, the word filling areas to which multiple strings respectively belong may be joined from left to right based on the distribution of the strings in the first image. Specifically, the word filling area to which the string arranged on the leftmost side in the first image belongs is also joined on the leftmost side, the word filling area to which the string arranged in the second column in the first image belongs is also joined in the second column, and the rest may be estimated by analogy. It should be noted that in this embodiment of the present application, the joining may alternatively be performed from right to left, and the joining order is not specifically limited.
[0273] 4. When there are multiple word filling areas and each of the word filling areas contains one character, the multiple word filling areas are joined from left to right and from top to bottom to obtain the second image, and the multiple word filling areas in the same row have the same height or the multiple word filling areas in the same column have the same width.
[0274] One word filling area contains one character. In this case, based on the distribution of the characters in the first image, the word filling areas to which multiple characters respectively belong may be joined from left to right and from top to bottom. The arrangement order of the word filling areas is consistent with the arrangement order of the characters included in the word filling areas in the first image, and the multiple word filling areas in the same row have the same height or the multiple word filling areas in the same column have the same width.
[0275] It should be noted that in this embodiment of the present application, the joining may alternatively be performed in another order, and the joining order is not specifically limited.
[0276] Step 804: Encode the second image and the word alignment side information to obtain the first bit stream, and the word alignment side information includes the height or width of the word filling area.
[0277] In this embodiment of the present application, standard video encoding or IBC encoding can be performed on the second image. This is not specifically limited. When IBC encoding is used, if the characters in the currently encoded CU have already appeared in the encoded CU, the encoded CU can be determined as the predicted block of the currently encoded CU. In this way, the probability that the obtained residual block is 0 becomes very high, thereby improving the encoding efficiency of the current CU. The word alignment side information may be encoded in an exponential Golomb coding scheme or may be encoded in another coding scheme. This is also not specifically limited. In this embodiment of the present application, the height or width of the word filling area may be directly encoded, or n times the preset size of the height or width of the word filling area may be encoded. This is not specifically limited.
[0278] Optionally, the word alignment side information may further include the height and width of the word area, and the horizontal and vertical coordinates of the pixel at the upper left corner of the word area in the first image. This information can assist the decoder side in restoring the word area in the reconstructed image, thereby improving the decoding efficiency.
[0279] In a possible implementation, to obtain the third image, the pixel values in the word area in the first image may be filled with preset pixel values, and the third image is encoded to obtain the second bitstream.
[0280] In the above steps, the word area is recognized from the first image, and then the word area is filled to obtain the word filling area, and the second image is obtained by joining the word filling area. It can be found that the second image contains only the character content in the first image and the second image is encoded separately. However, the first image contains other content in addition to the characters. Therefore, in order to maintain the integrity of the image, other image content other than the characters needs to be further processed and encoded.
[0281] In this embodiment of the present application, the pixel values of the pixels in the recognized word region in the original first image can be filled with a preset pixel value (for example, 0, 1, or 255, but not limited thereto). This corresponds to removing the word region from the first image and replacing the pixel values with the pixel values of non-character pixels. The difference between the third image obtained in this way and the first image lies in the different pixel values of the word region. Standard video encoding or IBC encoding can also be performed on the third image. This is not specifically limited.
[0282] In this embodiment of the present application, the word region obtained by recognition in the original image is filled so that the size of the obtained word filling region matches the size of the CU for encoding, and alignment between the word filling region and the CU is implemented to avoid classifying the pixels of the same character into different CUs as much as possible, thereby improving the accuracy and efficiency of word content compression.
[0283] FIG. 9 is a flowchart of a process 900 of a word-containing image decoding method according to an embodiment of the present application. The process 900 can be executed by the video decoder 30. The process 900 is described as a series of steps or operations. It should be understood that the process 900 is not limited to the execution order shown in FIG. 9 and may be executed in various orders and / or simultaneously. It is assumed that the bitstream is decoded by the video decoder 30 using the process 900 including the following steps to reconstruct the image.
[0284] Step 901: Obtain a bitstream.
[0285] The decoder side can receive the bitstream from the encoder side.
[0286] Step 902: Decode the bitstream to obtain the first image and word alignment side information. The word alignment side information includes the height or width of the word filling area included in the first image. The height or width of the word filling area is n times the preset size, where n ≥ 1.
[0287] The decoder side can obtain the first image and word alignment side information by decoding. The decoding method can correspond to the encoding method used on the encoder side. The encoder side can encode the image using standard video encoding or IBC encoding, and the decoder side can decode the bitstream using standard video decoding or IBC decoding to obtain the reconstructed image. The encoder side can encode the word alignment side information using Golomb coding, and the decoder side can decode the bitstream using Golomb decoding to obtain the word alignment side information.
[0288] The height or width of the word filling area is n times the preset size, where n ≥ 1, and the preset size is any one of 8, 16, 32, and 64. For details, refer to the above description in Step 802. Details will not be described again here.
[0289] Step 903: Obtain the word filling area based on the first image and the height or width of the word filling area.
[0290] After determining the height or width of the word filling area, the decoder side can extract the corresponding pixel values from the first image based on the height or width to obtain the word filling area.
[0291] Step 904: Obtain the word area in the image to be reconstructed based on the word filling area.
[0292] In this embodiment of the present application, the word alignment side information further includes the height and width of the word region. Therefore, based on the height and width of the word region, corresponding pixel values can be extracted from the word filling region to obtain the word region in the image to be reconstructed. The word region corresponds to the word region in step 801 and can include one character, or can include one line of characters, where the line of characters can include multiple characters, or can include one string of characters, and the string of characters can include multiple characters.
[0293] In a possible implementation, the decoder side can decode the bitstream to obtain a second image and obtain the image to be reconstructed based on the word region and the second image.
[0294] The word alignment side information further includes the height and width of the word region, and the horizontal and vertical coordinates of the pixel at the upper left corner of the word region. Based on this, the replacement region in the second image can be determined based on the height and width of the word region and the horizontal and vertical coordinates of the pixel at the upper left corner of the word region, and the pixel values in the replacement region are filled with the pixel values in the word region to obtain the image to be reconstructed. This process may be the reverse of the process of obtaining the third image in the embodiment shown in FIG. 8, that is, the process of filling the word region into the second image (corresponding to the third image in the embodiment shown in FIG. 8) to obtain the reconstructed image (corresponding to the first image in the embodiment shown in FIG. 8). After the horizontal and vertical coordinates of the pixel at the upper left corner of the word region in the second image, and the height and width of the word region are obtained, the replacement region in the second image may be determined, and then the pixel values in the filling region are filled with the pixel values of the word region. This corresponds to removing the replacement region from the second image and replacing the pixel values in the filling region with the pixel values of the word region.
[0295] In this embodiment of the present application, the alignment information of the word region can be obtained by analyzing the bitstream in order to obtain a character filling region including characters. The size of the character filling region matches the size of the CU for decoding in order to improve the accuracy and efficiency of word content compression and to implement the alignment between the word filling region and the CU.
[0296] FIG. 10 is an exemplary block diagram of an encoding system 100 according to the present application. As shown in FIG. 10, the encoding system 100 includes an encoder side and a decoder side.
[0297] Encoder side 1. Word region detection and segmentation The original image is processed using an existing word detection and segmentation method to output a word region. Any word region can contain one or more characters. If a word region contains multiple characters, the multiple characters are characters in the same character line in the original image. In this embodiment of the present application, the word region may be a rectangular region. Therefore, the word region may be referred to as a character box. The positions of one or more characters included in the character box may be represented by the coordinates of the upper left corner and the size (including height, or height and width) of the character box in the original image.
[0298] In this embodiment of the present application, word region detection and segmentation can be implemented using an existing method based on a horizontal projection histogram or connected components. The details are as follows.
[0299] (1) Preprocessing: Perform binarization and dilation on the original image. In the binarized image, the samples corresponding to the word region are white, and the samples corresponding to the non-word region are black. Dilation is used to handle pixel adhesion between characters.
[0300] (2) Perform a horizontal projection on the preprocessed image and calculate the amount of white samples in the horizontal direction.
[0301] The amount of white samples on the horizontal line with text is not zero. Therefore, all text lines in the image can be acquired.
[0302] Next, the following operations are performed on each text line: extracting connected components (a connected component refers to a continuous white sample region, and one character is a connected component, and correspondingly, one text line can have multiple connected components), and selecting a text box of an appropriate size. When the text box contains only one character, the size of the text box is greater than or equal to the size of the largest connected component in the text line, and all characters in the same text line use text boxes of the same size. When the text box contains one text line, the size of the text box is greater than or equal to the sum of the sizes of all connected components in the text line. For example, the connected component of character 1 in the text line occupies a pixel space of 6 (width) × 10 (height), and the connected component of character 2 occupies a pixel space of 11 × 8. In this case, at least a text box of 11×10 should be selected to include the two characters, or at least a text box of 17×10 should be selected to include the two characters of the text line.
[0303] In this embodiment of the present application, the above-mentioned word region detection and segmentation method can be implemented by the word processing module in FIG. 7. It should be understood that the word processing module may also implement word region detection and segmentation by using a method based on a deep neural network. This is not specifically limited.
[0304] After the word region is acquired, the pixel values in the word region in the original image are filled with preset pixel values, so that the word region can be removed from the original image, thereby obtaining a non-word image with all characters removed from the original image.
[0305] 2. Adaptive Word Alignment The text box is filled within a text filling box having a width Kw and a height Kh, and Kw and Kh may or may not be equal. In the HEVC standard, the sizes of CUs include 8, 16, 32, and 64. In this embodiment of the present application, Kw and Kh can each be a multiple of 8, 16, 32, or 64. In the text filling box, a preset pixel value may be filled in the area other than the area occupied by the text box, for example, the pixel value of the background area within the text box.
[0306] A plurality of text filling boxes can be obtained based on a plurality of text boxes, and it can be found that the plurality of text filling boxes are joined. The text filling boxes to which the characters of the same text line in the original image belong are joined to the word area aligned with the block division grid of the text line. When there are a plurality of text lines, when the height of the text filling box of the current text line is the same as that of the previous line, the text filling box of the current text line may be joined to the right side of the word area aligned with the block division grid of the previous line, or the text filling box of the current text line may be joined below the word area aligned with the block division grid of the previous line. When the height of the text filling box of the current text line is different from that of the previous line, the text filling box of the current text line may be joined below the word area aligned with the block division grid of the previous line. After the adaptive alignment of all text lines is completed, a completely aligned word area can be obtained to obtain a frame of the image (which can be referred to as a word image) including all the characters in the original image. Note that in this embodiment of the present application, the width of the aligned word area may not exceed the width of the original image, or may be another width. This is not specifically limited.
[0307] In this embodiment of the present application, the above-mentioned adaptive word alignment method can be implemented by the word processing module in FIG. 7.
[0308] 3. Word Alignment Side Information Encoding In the above two steps, the width and height of the character box, the width and height of the character filling box, and the vertical and horizontal coordinates of the characters in the line can be obtained. This information helps the decoder side restore the word area to the correct position. Therefore, the above information is used as word alignment side information for encoding to obtain the bit stream of the word alignment side information. It should be noted that the method for encoding the word alignment side information is not specifically limited in the embodiments of this application.
[0309] In this embodiment of this application, the above-mentioned method for encoding the word alignment side information can be implemented by the word alignment side information encoding module in FIG. 7.
[0310] 4. Standard Encoding The word image and non-word image obtained in the above steps can be encoded using a standard video encoding method. For example, the standard encoder that can be used may be an encoder compliant with standards such as H.264, H.265, H.266, AVS3, and AV1. This is not specifically limited in the embodiments of this application.
[0311] In this embodiment of this application, the above-mentioned standard encoding method can be implemented by the standard video encoding module in FIG. 7.
[0312] 5. Bit Stream Transmission The encoder side transmits the bit stream to the decoder side.
[0313] In this embodiment of this application, the above-mentioned bit stream transmission method can be implemented by the transmission module in FIG. 7.
[0314] 6. Decoding of Word Alignment Side Information On the decoder side, in order to obtain the word alignment side information, the bitstream is decoded using a decoding method corresponding to the encoder side, and the word alignment side information can include the width and height of the character box, the width and height of the character filling box, and the vertical and horizontal coordinates of the characters in the line.
[0315] In this embodiment of the present application, the foregoing method for decoding the word alignment side information can be implemented by the word alignment side information decoding module of FIG. 7.
[0316] 7. Standard Decoding On the decoder side, in order to obtain the reconstructed word image and the reconstructed non-word image, the bitstream is decoded using a decoding method corresponding to the encoder side.
[0317] In this embodiment of the present application, the foregoing standard decoding method can be implemented by the standard video decoding module of FIG. 7.
[0318] 8. Word Restoration Based on the word alignment side information, the character filling box is first obtained from the word image, and then the character box is extracted from the character filling box.
[0319] 9. Region Joining The corresponding position in the reconstructed non-word image is found based on the coordinate information in the word alignment side information, and the pixel values in the character frame are used to replace the pixel values at the found corresponding positions in the reconstructed non-word image to obtain the reconstructed image.
[0320] In this embodiment of the present application, the foregoing word restoration method can be implemented by the word reconstruction module of FIG. 7.
[0321] Based on the encoding system 100 shown in FIG. 10, hereinafter, some specific embodiments are used to describe in detail the technical solutions of the embodiments of the foregoing method.
[0322] Embodiment 1 On the encoder side Fig. 11a is a diagram of word region detection and segmentation. As shown in Fig. 11a, the original frame includes
Number
[0323]
Number
[0324] The pixel values within the word regions corresponding to the four character boxes in the original frame are obtained, and as a result, four images containing the characters can be obtained. The resolution of the images is lower than that of the original frame, and the size of the images corresponds to the size of the character boxes.
[0325] Next, the pixel values within the word regions corresponding to the four character boxes in the original frame are filled with a preset pixel value. For example, the preset pixel value may be 0 or 1, or the preset pixel value may be the pixel value of the background region within the character box other than the character pixels. This is not specifically limited. Since the image obtained by re-filling does not contain characters, it may be referred to as a non-word image (or non-word region).
[0326] Fig. 11b is a diagram of adaptive word alignment. Taking the original frame of Fig. 11a from top to bottom and from left to right (in a zigzag shape) as an example,
Number
[0327] The character filling boxes to which the characters in the same character line belong are joined to the word area aligned with the block division grid of the character line. For example,
Number
[0328] In the foregoing process, in order to restore the characters in the aligned word area to the positions of the characters in the original frame, word alignment side information needs to be output and sent to the decoder side. The word alignment side information is obtained by simply performing data conversion on the positions of the character boxes and includes the following elements (FIGS. 11a and 11b are used as examples). (1) The height H (i.e., 10) and width W (i.e., 11) of the character box, and the size K (i.e., 16) of the character filling box. (2) The vertical coordinates y of the upper left corner of the character line: 15, 15, 15, and 15. The vertical coordinates of the same character line are the same. Therefore, the vertical coordinates of multiple characters in the same character line can be simplified to represent the same information. For example, the vertical coordinate of the first character of the character line (i.e., 15) and the number of characters (i.e., 4) are recorded. (3) The horizontal coordinates x of the upper left corner of the character line: 2, 14, 25, and 36. The differences in the horizontal coordinates of multiple characters are 12, 11, and 11, and it can be found that the data is close to the width W of the character box. Therefore, the horizontal coordinates of multiple characters in the same character line can be simplified to represent the same information. For example, the horizontal coordinate of the first character of the character line (i.e., 2) and the differences in the horizontal coordinates (i.e., 12, 11, and 11) are recorded, or the result E (i.e., E1, E2, E3 = [12 - 11], [11 - 11], [12 - 11] = 1, 0, 0) obtained by subtracting the differences in the horizontal coordinates (i.e., 12, 11, and 11) from W is recorded. By recording this data, the horizontal coordinate of the upper left corner of the character line can be restored. For example, the horizontal coordinate of the second character = the horizontal coordinate of the first character + W + E1 = 2 + 11 + 1 = 14, and the rest can be inferred by analogy.
[0329] The word alignment side information includes H, W, K, the horizontal and vertical coordinates of the first character of each character line, and the number of characters E in each character line. H, W, K, and the number of characters in each line are encoded in binary mode. The value of E is close to 0, and the exponential Golomb coding method is used. For the horizontal and vertical coordinates of the first character of each character line, binary coding can first be performed on the horizontal and vertical coordinates of the first character of the first character line. For another line, the difference between the horizontal and vertical coordinates of the first character of the character line and the horizontal and vertical coordinates of the first character of the previous line is calculated, and then exponential Golomb coding is performed on the coordinate difference.
[0330] It should be noted that the above-described method of encoding the word alignment side information is merely an example. In this embodiment of the present application, the word alignment side information may alternatively be encoded by another method. This is not specifically limited.
[0331] In this embodiment of the present application, the non-word image and the word image can be respectively encoded and decoded by a standard encoding / decoding method. Details are not described again here.
[0332] Decoder side The decoder side decodes the bitstream by a decoding method corresponding to the encoder side in order to obtain the word alignment side information.
[0333] Figure 11c is a diagram of word restoration. As shown in Figure 11c, based on the word alignment side information, K×K character filling boxes are extracted from the reconfigured aligned word region (reconfigured word image). Then, based on the word alignment side information, 4 W×H character boxes are extracted from the 4 character filling boxes in a zigzag order. Each of the 4 character boxes contains
Number
[0334] In the zigzag order, the pixel values in the four character boxes are filled in the corresponding positions in the reconstructed non-word image based on the coordinates recorded in the alignment side information. For example,
Number
Number
Number
Number
[0335] Embodiment 2 The difference from Embodiment 1 is that in Embodiment 1, adaptive word alignment is performed in two directions, the width direction and the height direction, while in this embodiment, adaptive word alignment is performed only in the height direction.
[0336] Figure 12a is a diagram of word area detection and division. As shown in Figure 12a, the original frame is
Number
[0337] Pixel values within the word region corresponding to the character box are obtained, and as a result, an image containing four characters can be obtained. The resolution of the image is lower than that of the original frame, and the size of the image corresponds to the size of the character box.
[0338] Next, the pixel values within the word region corresponding to the character box in the original frame are filled with a preset pixel value. For example, the preset pixel value may be 0 or 1, or the preset pixel value may be the pixel value of the background region within the character box other than the character pixels. This is not specifically limited. Since the image obtained by refilling does not contain characters, it may be referred to as a non-word image (or non-word region).
[0339] Figure 12b is a diagram of adaptive word alignment. As shown in Figure 12b, word boxes are filled in a character filling box with a height of K in order from top to bottom. Since K is selected as 16 based on the height of the word box in Figure 12a, the size of the character filling box is 45×16. In this embodiment of the present application, filling is performed on the lower side of the word box so as to reach the size of the character filling box. The pixel value of the filling may be the pixel value of the background area within the word box or a preset pixel value (for example, 0 or 1). This is not specifically limited. Note that the height and width of the character filling box may be different and are set based on the size of the CU. For example, when the size of the CU is 16, K may be 1 times 16 (i.e., 16) in this embodiment. When the size of the CU is 8, in this embodiment, K may be 2 times 8 (i.e., 16). When the size of the CU is 32, K may be 1 times 32 (i.e., 32) in this embodiment. Or when the size of the CU is 64, K may be 1 times 64 (i.e., 64) in this embodiment.
[0340] In the above process, in order to restore the characters in the aligned word area to the positions of the characters in the original frame, word alignment side information needs to be output and sent to the decoder side. The word alignment side information is obtained by simply performing data conversion on the positions of the word boxes and includes the following elements (Figures 12a and 12b are used as examples). (1) The height H (i.e., 10) and width W (i.e., 45) of the word box, and the height K (i.e., 16) of the character filling box. (2) The vertical coordinate y of the upper left corner of the character line: 15 (3) The horizontal coordinate x of the upper left corner of the character line: 2
[0341] The word alignment side information includes H, W, K, and the horizontal and vertical coordinates of the first character of the character line. H, W, and K are encoded in binary mode. For the horizontal and vertical coordinates of the first character of the character line, binary encoding can first be performed on the horizontal and vertical coordinates of the first character of the first character line. For another line, the difference between the horizontal and vertical coordinates of the first character of the character line and the horizontal and vertical coordinates of the first character of the previous line is calculated, and then exponential Golomb encoding is performed on the coordinate difference.
[0342] Note that the above-described method of encoding word alignment side information is merely an example. In this embodiment of the present application, the word alignment side information may alternatively be encoded by another method. This is not specifically limited.
[0343] In this embodiment of the present application, the non-word image and the word image can be respectively encoded and decoded by a standard encoding / decoding method. Details are not described again here.
[0344] Decoder side The decoder side decodes the bit stream in a decoding method corresponding to the encoder side in order to obtain the word alignment side information.
[0345] Based on the word alignment side information, 45×K character filling boxes are extracted from the reconstructed aligned word region (reconstructed word image). Then, based on the word alignment side information, W×H character boxes are extracted from the character filling boxes. The character boxes contain
Number
[0346] The pixel values in the character boxes are filled in the corresponding positions in the reconstructed non-word image based on the coordinates recorded in the alignment side information. For example, in the reconstructed non-word video,
Number
[0347] Embodiment 3 The difference from Embodiment 2 is that the original frame includes a plurality of character lines. In word region detection and segmentation, a plurality of character lines are detected in advance. Therefore, during segmentation, a plurality of character lines need to be further segmented to obtain a plurality of word regions, and each word region includes one character line.
[0348] Subsequently, the aforementioned plurality of word regions are separately encoded / decoded by the method of Embodiment 2. Details are not described again here.
[0349] Embodiment 4 The difference from Embodiment 1 is that the original frame includes a plurality of character lines. In word region detection and segmentation, a plurality of character lines are detected in advance. Therefore, during segmentation, a plurality of character lines need to be further segmented to obtain a plurality of character lines, and then each character line is segmented to obtain a plurality of word regions, and each word region includes one character.
[0350] Subsequently, the aforementioned plurality of word regions are separately encoded / decoded by the method of Embodiment 1. Details are not described again here.
[0351] FIG. 13 is an exemplary diagram of the structure of an encoding device 1300 according to an embodiment of the present application. As shown in FIG. 13, the encoding device 1300 of the present embodiment may be used on the encoder side 20. The encoding device 1300 may include an acquisition module 1301, a filling module 1302, and an encoding module 1303.
[0352] The acquisition module 1301 is configured to acquire a word area in a first image, and the word area includes at least one character. The filling module 1302 is configured to fill the word area to obtain a word filling area, and the height or width of the word filling area is n times a preset size, where n≥1, and is configured to obtain a second image based on the word filling area. The encoding module 1303 is configured to encode the second image and word alignment side information to obtain a first bit stream, and the word alignment side information includes the height or width of the word filling area.
[0353] In a possible implementation, the filling module 1302 is specifically configured to obtain the height of the word filling area based on the height of the word area, the height of the word filling area is n times a preset size, fill the word area in the vertical direction of the word area to obtain the word filling area, and the filling height is the difference between the height of the word filling area and the height of the word area.
[0354] In a possible implementation, the filling module 1302 is specifically configured to obtain the width of the word filling area based on the width of the word area, the width of the word filling area is n times a preset size, fill the word area in the horizontal direction of the word area to obtain the word filling area, and the filling width is the difference between the width of the word filling area and the width of the word area.
[0355] In a possible implementation, the word area can include one character, or the word area can include one line of characters, the line of characters includes a plurality of characters, or the word area can include one string of characters, and the string of characters includes a plurality of characters.
[0356] In a possible implementation, the acquisition module 1301 is specifically configured to perform character recognition on the first image to obtain the word area.
[0357] In a possible embodiment, the acquisition module 1301 is specifically configured to perform character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character lines, and perform line splitting for each character line on the intermediate word region to obtain a word region.
[0358] In a possible embodiment, the acquisition module 1301 is specifically configured to perform character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character strings, and perform column splitting for each character string on the intermediate word region to obtain a word region.
[0359] In a possible embodiment, the acquisition module 1301 is specifically configured to perform character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character lines or a plurality of character strings, and perform line splitting and column splitting for each character on the intermediate word region to obtain a word region.
[0360] In a possible embodiment, the filling module 1302 is specifically configured to use the word filling region as the second image when there is only one word filling region.
[0361] In a possible embodiment, when there are a plurality of word filling regions and each of the word filling regions includes one character line, the filling module 1302 is specifically configured to join the plurality of word filling regions from top to bottom to obtain the second image.
[0362] In a possible embodiment, when there are a plurality of word filling regions and each of the word filling regions includes one character string, the filling module 1302 is specifically configured to join the plurality of word filling regions from left to right to obtain the second image.
[0363] In a possible embodiment, when there are a plurality of word filling areas in the filling module 1302 and each of the word filling areas contains one character, in order to obtain a second image, the plurality of word filling areas are joined from left to right and from top to bottom, and specifically configured such that the plurality of word filling areas in the same row have the same height or the plurality of word filling areas in the same column have the same width.
[0364] In a possible embodiment, the word alignment side information further includes the height and width of the word area, and the horizontal and vertical coordinates of the pixel at the upper left corner of the word area in the first image.
[0365] In a possible embodiment, the preset size is one of 8, 16, 32, and 64.
[0366] In a possible embodiment, the filling module 1302 is further configured to fill the pixel values in the word area in the first image with preset pixel values to obtain a third image, and the encoding module 1303 is further configured to encode the third image to obtain a second bitstream.
[0367] The device in this embodiment may be used to execute the technical solution of the method embodiment shown in FIG. 8. The principle and technical effect of the device embodiment are the same as those of the technical solution of the method embodiment shown in FIG. 8. Details are not described again here.
[0368] FIG. 14 is an exemplary diagram of the structure of a decoding device 1400 according to an embodiment of the present application. As shown in FIG. 14, the decoding device 1400 of this embodiment may be used on the decoder side 30. The decoding device 1400 may include an acquisition module 1401, a decoding module 1402, and a reconstruction module 1403.
[0369] The acquisition module 1401 is configured to acquire a bitstream. The decoding module 1402 is configured to decode the bitstream to obtain a first image and word alignment side information, where the word alignment side information includes the height or width of a word filling area included in the first image, and the height or width of the word filling area is n times a preset size, where n ≥ 1. The reconstruction module 1403 is configured to obtain a word filling area based on the first image and the height or width of the word filling area, and obtain a word area in the image to be reconstructed based on the word filling area.
[0370] In a possible implementation, the word alignment side information further includes the height and width of the word area, and the reconstruction module 1403 is specifically configured to extract corresponding pixel values from the word filling area based on the height and width of the word area to obtain the word area.
[0371] In a possible implementation, the decoding module 1402 is further configured to decode the bitstream to obtain a second image, and the reconstruction module 1403 is further configured to obtain the image to be reconstructed based on the word area and the second image.
[0372] In a possible implementation, the word alignment side information further includes the horizontal and vertical coordinates of the pixel at the upper left corner of the word area. The reconstruction module 1403 determines a replacement area in the second image based on the height and width of the word area and the horizontal and vertical coordinates of the pixel at the upper left corner of the word area, and is specifically configured to fill the pixel values in the replacement area with the pixel values in the word area to obtain the image to be reconstructed.
[0373] In a possible implementation, the word area can include one character, or the word area can include one line of characters, where the line of characters includes a plurality of characters, or the word area can include one string of characters, where the string of characters includes a plurality of characters.
[0374] In a possible embodiment, the preset size is one of 8, 16, 32, and 64.
[0375] The apparatus in this embodiment may be used to execute the technical solution of the method embodiment shown in FIG. 9. The principle and technical effect of the apparatus embodiment are the same as those of the technical solution of the method embodiment shown in FIG. 9. Details are not described again here.
[0376] In an embodiment process, the steps in the above method embodiment can be implemented by using the hardware integrated logic circuit in the processor or by using instructions in the form of software. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or another programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor or the like. The steps of the method disclosed in the embodiments of the present application may be directly provided so as to be executed and completed by a hardware-encoded processor, or executed and completed by a combination of the hardware of the encoded processor and a software module. The software module may be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or a register. The storage medium is located in the memory, and the processor reads the information in the memory and performs the steps in the above method together with the hardware of the processor.
[0377] The memory in the foregoing embodiments may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. The non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM) used as an external cache. By way of example and not limitation, many forms of RAM may be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein includes, but is not limited to, these memories and any other suitable type of memory.
[0378] Those skilled in the art can recognize that the units and algorithm steps in the examples described with reference to the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the function is implemented by hardware or software depends on the specific application of the technical solution and design constraints. Those skilled in the art may use different methods to implement the described functions for each specific application, but the embodiments should not be considered to exceed the scope of this application.
[0379] For the sake of simplicity, for the detailed operation processes of the aforementioned systems, devices, and units, reference may be made to the corresponding processes in the aforementioned method embodiments, which can be clearly understood by those skilled in the art. Details are not described again here.
[0380] It should be understood that in some embodiments provided in this application, the disclosed systems, devices, and methods can also be implemented in other ways. For example, the above-described device embodiments are merely examples. For example, the division into units is merely a logical function division, and other divisions may be possible in actual embodiments. For example, multiple units or components may be combined or integrated into another system, and some features may be ignored or not executed. In addition, the presented or described mutual coupling or direct coupling or communication connection may be implemented via some interfaces. The indirect coupling or communication connection between devices or units may be implemented in electrical form, mechanical form, or another form.
[0381] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units. Specifically, the components may be located in one place or distributed over multiple network units. To achieve the purpose of the solution of the embodiment, some or all of the units may be selected based on actual requirements.
[0382] In addition, the functional units of the embodiments of the present application may be integrated into one processing unit, each unit may physically exist independently, or two or more units may be integrated into one unit.
[0383] When the function is implemented in the form of a software functional unit and sold or used as an independent product, the function may be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of this application, in essence, or the parts contributing to the prior art, or part of the technical solutions, may be implemented in the form of a software product. The software product is stored in a storage medium and includes several instructions for instructing a computer device (such as a personal computer, a server, or a network device) to execute all or part of the steps of the method in the embodiments of this application. The above storage medium includes any medium that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0384] The foregoing description is merely a specific embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification or substitution that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application shall be within the protection scope of the present application. Therefore, the protection scope of the present application shall follow the protection scope of the claims.
Description of Reference Numerals
[0385] 2 Output third layer index 4 Input second index 10 Video coding system 12 Source device 13 Communication channel 14 Destination device 16 Image source 17 Image, image data 18 Preprocessor, preprocessing unit 19 Preprocessed image data 20 Video encoder 21 Encoded image data 22 Communication interface 25 Training engine 28 Communication interface 30 Video decoder 31 Decoded image data 32 Postprocessor, postprocessing unit 33 Postprocessed image data 34 Display device 40 Video encoding system 41 Imaging device 42 Antenna 43 Processor 44 Memory 45 Display device 46 Processing circuit 70 Encoding system 100 Encoding system 201 Input terminal, input interface 203 Image block 204 Residual calculation unit 205 Residual block 206 Transformation processing unit 207 Transformation coefficient 208 Quantization unit 209 Quantization coefficient, quantized residual coefficient 210 Inverse quantization unit 211 Inverse quantization count, inverse quantized residual coefficient 212 Inverse transformation processing unit 213 Reconstructed residual block, transformation block 214 Reconstruction unit, adder 215 Reconstructed block 220 Loop filter unit 221 Reconstructed block 230 Decoded image buffer Decoded image Inter prediction unit Inter prediction block Intra prediction unit Mode selection unit Partition unit Predicted block Syntax element Entropy encoding unit Output terminal, output interface Entropy decoding unit Quantization coefficient Inverse quantization unit Inverse quantization coefficient Inverse transform processing unit Reconstructed residual block, transform block Reconstruction unit, adder Reconstructed block Loop filter unit Filtered block Decoded image buffer Decoded image Output terminal Inter prediction unit Intra prediction unit Mode application unit Predicted block Video encoding device Input port Receiver unit Transmitter unit Output port Memory Encoding module Device Processor Memory Code and data Operating system 510 Application 512 Bus 518 Display 800 Process 1300 Encoding Device 1301 Acquisition Module 1302 Filling Module 1303 Encoding Module 1400 Decoding Device 1401 Acquisition Module 1402 Decoding Module 1403 Reconstruction Module
Claims
1. A method for encoding an image containing words, comprising: obtaining a word region in a first image, wherein the word region contains at least one character; filling the word region to obtain a word filling region, wherein a height or width of the word filling region is n times a preset size, and n≥1; obtaining a second image based on the word filling region; encoding the second image and word alignment side information to obtain a first bit stream, wherein the word alignment side information includes the height or width of the word filling region. The method includes the above steps.
2. The step of filling the word region to obtain a word filling region includes: obtaining the height of the word filling region based on the height of the word region, wherein the height of the word filling region is n times the preset size; filling the word region vertically in the word region to obtain the word filling region, wherein a filling height is a difference between the height of the word filling region and the height of the word region. The method according to claim 1 includes the above steps.
3. The step of filling the word region to obtain a word filling region includes: obtaining the width of the word filling region based on the width of the word region, wherein the width of the word filling region is n times the preset size; filling the word region horizontally in the word region to obtain the word filling region, wherein a filling width is a difference between the width of the word filling region and the width of the word region. The method according to claim 1 or 2 includes the above steps.
4. The word region contains one character, or the word region contains one line of characters, and the line of characters contains a plurality of characters, or the word region contains one string of characters, and the string of characters contains a plurality of characters. The method according to any one of claims 1 to 3.
5. The step of obtaining a word region in a first image includes: performing character recognition on the first image to obtain the word region. The method according to any one of claims 1 to 4, comprising
6. The step of obtaining the word region in the first image is The step of performing character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character lines, and The step of performing line splitting for each character line on the intermediate word region to obtain the word region The method according to any one of claims 1 to 4, comprising
7. The step of obtaining the word region in the first image is The step of performing character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character strings, and The step of performing column splitting for each character string on the intermediate word region to obtain the word region The method according to any one of claims 1 to 4, comprising
8. The step of obtaining the word region in the first image is The step of performing character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character lines or a plurality of character strings, and The step of performing line splitting and column splitting for each character on the intermediate word region to obtain the word region The method according to any one of claims 1 to 4, comprising
9. The step of obtaining the second image based on the word filling region is When only one word filling region exists, using the word filling region as the second image The method according to any one of claims 1 to 5, comprising
10. The step of obtaining the second image based on the word filling region is When a plurality of word filling regions exist and each of the plurality of word filling regions includes one character line, joining the plurality of word filling regions from top to bottom to obtain the second image The method according to any one of claims 1 to 4 and 6, comprising
11. The step of obtaining the second image based on the word filling region is When a plurality of word filling regions exist and each of the plurality of word filling regions includes one character string, joining the plurality of word filling regions from left to right to obtain the second image The method according to any one of claims 1 to 4 and 7, comprising
12. The step of obtaining a second image based on the word filling area is as follows: When there are a plurality of word filling areas and each of the plurality of word filling areas contains one character, the step of joining the plurality of word filling areas from left to right and from top to bottom to obtain the second image, wherein the plurality of word filling areas in the same row have the same height or the plurality of word filling areas in the same column have the same width The method according to any one of claims 1 to 4 and 8, including the above steps.
13. The word alignment side information further includes the height and width of the word area, and the horizontal and vertical coordinates of the pixel at the upper left corner of the word area in the first image. The method according to any one of claims 1 to 12.
14. The preset size is one of 8, 16, 32, and 64. The method according to any one of claims 1 to 13.
15. After the step of obtaining the word area in the first image, the method further includes: Filling the pixel values in the word area in the first image with preset pixel values to obtain a third image; Encoding the third image to obtain a second bit stream. The method according to any one of claims 1 to 14, further including the above steps.
16. A method for decoding a word-containing image, including: Obtaining a bit stream; Decoding the bit stream to obtain a first image and word alignment side information, wherein the word alignment side information includes the height or width of a word filling area included in the first image, and the height or width of the word filling area is n times the preset size, where n≥1; Obtaining the word filling area based on the first image and the height or width of the word filling area; Obtaining a word area in the image to be reconstructed based on the word filling area. A method including the above steps.
17. The word alignment side information further includes the height and width of the word area. The step of obtaining a word area in the image to be reconstructed based on the word filling area is as follows: A step of extracting corresponding pixel values from the word filling area based on the height and width of the word area to obtain the word area The method according to claim 16, including this step.
18. A step of decoding the bit stream to obtain a second image, and A step of obtaining the image to be reconstructed based on the word area and the second image The method according to claim 17, further including these steps.
19. The word alignment side information further includes the horizontal and vertical coordinates of the pixel at the upper left corner of the word area, The step of obtaining the image to be reconstructed based on the word area and the second image is A step of determining a replacement area in the second image based on the height and width of the word area and the horizontal and vertical coordinates of the pixel at the upper left corner of the word area, and A step of filling the pixel values in the replacement area with the pixel values in the word area to obtain the image to be reconstructed The method according to claim 18, including these steps.
20. The word area includes one character, or The word area includes one line of characters, and the line of characters includes a plurality of characters, or The word area includes one string of characters, and the string of characters includes a plurality of characters The method according to any one of claims 16 to 19.
21. The method according to any one of claims 16 to 20, wherein the preset size is one of 8, 16, 32, and 64.
22. An encoding device, comprising An acquisition module configured to acquire a word area in a first image, wherein the word area includes at least one character, and the acquisition module; A filling module configured to fill the word area to obtain a word filling area and obtain a second image based on the word filling area, wherein the height or width of the word filling area is n times a preset size, and n≧1, and the filling module; An encoding module configured to encode the second image and word alignment side information to obtain a first bit stream, wherein the word alignment side information includes the height or width of the word filling area, and the encoding module The device comprising these components.
23. The filling module is Obtaining the height of the word filling area based on the height of the word area, wherein the height of the word filling area is n times the preset size; Filling the word area in the vertical direction of the word area to obtain the word filling area, wherein the filling height is the difference between the height of the word filling area and the height of the word area; The apparatus according to claim 22, further configured to perform the above.
24. The filling module: Obtaining the width of the word filling area based on the width of the word area, wherein the width of the word filling area is n times the preset size; Filling the word area in the horizontal direction of the word area to obtain the word filling area, wherein the filling width is the difference between the width of the word filling area and the width of the word area; The apparatus according to claim 22 or 23, further configured to perform the above.
25. The word area includes one character, or The word area includes one line of characters, the line of characters includes a plurality of characters, or The word area includes one string of characters, the string of characters includes a plurality of characters. The apparatus according to any one of claims 22 to 24.
26. The obtaining module is further configured to perform character recognition on the first image to obtain the word area, the apparatus according to any one of claims 22 to 25.
27. The obtaining module: Performing character recognition on the first image to obtain an intermediate word area, the intermediate word area includes a plurality of lines of characters; Performing line division on the intermediate word area for each line of characters to obtain the word area. The apparatus according to any one of claims 22 to 25, further configured to perform the above.
28. The obtaining module: Performing character recognition on the first image to obtain an intermediate word area, the intermediate word area includes a plurality of strings of characters; Performing column division on the intermediate word area for each string of characters to obtain the word area. The apparatus according to any one of claims 22 to 25, further configured to perform the above.
29. The obtaining module: Performing character recognition on the first image to obtain an intermediate word region, where the intermediate word region includes a plurality of character lines or a plurality of character strings, and Performing row division and column division for each character in the intermediate word region to obtain the word region The apparatus according to any one of claims 22 to 25, further configured to perform.
30. The filling module is further configured to use the word filling region as the second image when only one word filling region exists, the apparatus according to any one of claims 22 to 26.
31. The filling module is further configured to join the plurality of word filling regions from top to bottom to obtain the second image when a plurality of word filling regions exist and each of the plurality of word filling regions includes one character line, the apparatus according to any one of claims 22 to 25 and 27.
32. The filling module is further configured to join the plurality of word filling regions from left to right to obtain the second image when a plurality of word filling regions exist and each of the plurality of word filling regions includes one character string, the apparatus according to any one of claims 22 to 25 and 28.
33. The filling module is further configured to join the plurality of word filling regions from left to right and from top to bottom to obtain the second image when a plurality of word filling regions exist and each of the plurality of word filling regions includes one character, and a plurality of word filling regions in the same row have the same height or a plurality of word filling regions in the same column have the same width, the apparatus according to any one of claims 22 to 25 and 29.
34. The word alignment side information further includes the height and width of the word region, and the horizontal coordinate and vertical coordinate of the pixel at the upper left corner of the word region in the first image, the apparatus according to any one of claims 22 to 33.
35. The preset size is one of 8, 16, 32, and 64, the apparatus according to any one of claims 22 to 34.
36. The filling module is further configured to fill pixel values within the word region in the first image with preset pixel values to obtain a third image. The encoding module is further configured to encode the third image to obtain a second bitstream. The apparatus according to any one of claims 22 to 35.
37. A decoding apparatus, an acquisition module configured to acquire a bitstream, a decoding module configured to decode the bitstream to obtain a first image and word alignment side information, where the word alignment side information includes the height or width of a word filling region included in the first image, and the height or width of the word filling region is n times a preset size, where n ≥ 1; a reconstruction module configured to obtain the word filling region based on the first image and the height or width of the word filling region, and obtain a word region in a reconstruction target image based on the word filling region The apparatus comprising.
38. The word alignment side information further includes the height and width of the word region. The apparatus according to claim 37, wherein the reconstruction module is further configured to extract corresponding pixel values from the word filling region based on the height and width of the word region to obtain the word region.
39. The decoding module is further configured to decode the bitstream to obtain a second image. The reconstruction module is further configured to obtain the reconstruction target image based on the word region and the second image. The apparatus according to claim 38.
40. The word alignment side information further includes the horizontal coordinate and vertical coordinate of a pixel at the upper left corner of the word region. The apparatus according to claim 39, wherein the reconstruction module is further configured to determine a replacement region in the second image based on the height and width of the word region and the horizontal coordinate and vertical coordinate of the pixel at the upper left corner of the word region, and fill pixel values in the replacement region with pixel values in the word region to obtain the reconstruction target image.
41. The word region includes one character, or The word area includes one line of characters, and the line of characters includes a plurality of characters, or The word area includes one character string, and the character string includes a plurality of characters. The apparatus according to any one of claims 37 to 40.
42. The apparatus according to any one of claims 37 to 41, wherein the preset size is one of 8, 16, 32, and 64.
43. One or more processors, A memory configured to store one or more programs Comprising When the one or more programs are executed by the one or more processors, the one or more processors can implement the method according to any one of claims 1 to 15. Encoder.
44. One or more processors, A memory configured to store one or more programs Comprising When the one or more programs are executed by the one or more processors, the one or more processors can implement the method according to any one of claims 16 to 21. Decoder.
45. A computer-readable storage medium storing a computer program, wherein when the computer program is executed on a computer, the computer can implement the method according to any one of claims 1 to 21.
46. A computer program, wherein the computer program includes instructions, and when the instructions are executed on a computer or a processor, the computer or the processor can implement the method according to any one of claims 1 to 21.
Citation Information
Patent Citations
Image processing apparatus, image processing method and image processing program
JP2002369011A
Image processing apparatus, image processing program and image processing method
JP2005159663A
Image processing system, image processing method, and computer program
JP2007025814A
Image compression device, image compression method, and computer program
JP2013125994A
Image processing apparatus, image processing method and program
JP2015198385A