Image coding method, decoding method, device, electronic equipment and storage medium

CN118368429BActive Publication Date: 2026-09-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310080247.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2026-09-08
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

[0002]相关技术中,编码器的量化参数QP(Quantization Parameter)无法根据图像的细节要求进行自定义的调整,图像中细节少的区域和细节多的区域均使用同一组QP值,从而导致渲染出的图像中局部细节区域画质较低

Benefits of technology

[0024] The image encoding method provided in this disclosure obtains encoding unit information of the image to be encoded and determines at least one target region in the image to be encoded based on the encoding unit information, thereby encoding the image of the target region with a first preset quantization parameter value. It can be seen that this disclosure improves the image quality of local details without affecting the overall image bitrate by adjusting only the quantization parameter value of the target region in the image to be encoded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118368429B_ABST
    Figure CN118368429B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image encoding method, a decoding method, an apparatus, an electronic device and a storage medium. The image encoding method comprises: obtaining encoding unit information of an image to be encoded; determining at least one target region in the image to be encoded according to the encoding unit information; and encoding an image of the target region with a first preset quantization parameter value. The method of the present disclosure adjusts only the quantization parameter value of the target region in the image to be encoded, so as to improve the local detail quality without affecting the code rate of the whole image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an image encoding method, decoding method, apparatus, electronic device, and storage medium. Background Technology

[0002] In related technologies, the encoder's quantization parameter (QP) cannot be customized according to the detail requirements of the image. The same set of QP values ​​is used for both areas with little detail and areas with much detail in the image, resulting in low image quality in local detail areas of the rendered image. Summary of the Invention

[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] This disclosure provides an image encoding method, a decoding method, an apparatus, an electronic device, and a storage medium.

[0005] The following technical solution is adopted in this disclosure.

[0006] In some embodiments, this disclosure provides an image encoding method, including:

[0007] Obtain the coding unit information of the image to be encoded;

[0008] Based on the encoding unit information, at least one target region in the image to be encoded is determined;

[0009] The image of the target region is encoded using a first preset quantization parameter value.

[0010] In some embodiments, this disclosure provides an image decoding method, including:

[0011] Receive the image to be encoded and at least one target region image in the image to be encoded sent by the encoding end;

[0012] The image to be encoded and the target region are decoded, and the image of the target region is combined with the image to be encoded and then displayed.

[0013] The target region is determined by the encoding end based on the encoding unit information of the image to be encoded.

[0014] In some embodiments, this disclosure provides an image encoding apparatus, comprising:

[0015] The acquisition module is used to acquire the coding unit information of the image to be encoded;

[0016] The first processing module is used to determine at least one target region in the image to be encoded based on the encoding unit information;

[0017] The second processing module is used to encode the image of the target region using a first preset quantization parameter value.

[0018] In some embodiments, this disclosure provides an image decoding apparatus, including:

[0019] A receiving module is used to receive an image to be encoded and an image of at least one target region in the image to be encoded, sent by an encoding end; the target region is determined by the encoding end based on the encoding unit information of the image to be encoded.

[0020] The third processing module is used to decode the image to be encoded and the target region, and to combine the image of the target region with the image to be encoded for display.

[0021] In some embodiments, this disclosure provides an electronic device, including: at least one memory and at least one processor;

[0022] The memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above method.

[0023] In some embodiments, this disclosure provides a computer-readable storage medium for storing program code that, when run by a processor, causes the processor to perform the methods described above.

[0024] The image encoding method provided in this disclosure obtains encoding unit information of the image to be encoded and determines at least one target region in the image to be encoded based on the encoding unit information, thereby encoding the image of the target region with a first preset quantization parameter value. It can be seen that this disclosure improves the image quality of local details without affecting the overall image bitrate by adjusting only the quantization parameter value of the target region in the image to be encoded. Attached Figure Description

[0025] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0026] Figure 1 This is one of the flowcharts of the image encoding method according to an embodiment of this disclosure.

[0027] Figure 2 This is one of the schematic diagrams of image encoding according to an embodiment of the present disclosure.

[0028] Figure 3 This is a second schematic diagram of image encoding according to an embodiment of this disclosure.

[0029] Figure 4 This is the third schematic diagram of image encoding according to an embodiment of this disclosure.

[0030] Figure 5 This is the fourth schematic diagram of image encoding according to an embodiment of this disclosure.

[0031] Figure 6 This is the second flowchart of the image encoding method according to an embodiment of the present disclosure.

[0032] Figure 7 This is the fifth schematic diagram of image encoding according to an embodiment of the present disclosure.

[0033] Figure 8 This is the sixth schematic diagram of image encoding according to an embodiment of this disclosure.

[0034] Figure 9 This is a flowchart of an image decoding method according to an embodiment of the present disclosure.

[0035] Figure 10 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0036] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0037] It should be understood that the various steps described in the method embodiments of this disclosure can be performed in sequence and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0038] The term "comprising" and its variations as used herein are open-ended inclusion, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. The term "in response to" and related terms refer to a signal or event being affected to some extent by another signal or event, but not necessarily completely or directly. If event x occurs "in response to" event y, then x may be directly or indirectly responsive to y. For example, the occurrence of y may ultimately lead to the occurrence of x, but there may be other intermediate events and / or conditions. In other cases, y may not necessarily lead to the occurrence of x, and x may occur even if y has not yet occurred. Furthermore, the term "in response to" can also mean "at least partially responsive to".

[0039] The term "determine" broadly encompasses a wide variety of actions, including acquisition, calculation, computation, processing, derivation, investigation, search (e.g., searching in a table, database, or other data structure), discovery, and similar actions; it may also include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), parsing, selecting, choosing, building, and similar actions, etc. Definitions for other terms will be provided below.

[0040] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0041] It should be noted that the use of the word "a" in this disclosure is illustrative rather than restrictive, and those skilled in the art should understand that it should be understood as "one or more" unless otherwise expressly indicated in the context.

[0042] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0043] It's important to note that the coding structure is divided into Slice layers and Tile units. An image can be divided into one or more slices, each encoded independently. For example, an image can be divided into N slices, arranged in a strip shape. During encoding, the Coding Tree Units (CTUs) within each slice are encoded in raster scan order. In the High Efficiency Video Coding (HEVC) standard, an image can be divided into several tiles, that is, the image is divided into several rectangular regions horizontally and vertically, and these rectangular regions are called tiles. The tiles are not required to be uniformly distributed; the entire image is divided into N tiles, each tile being rectangular. Typically, each tile contains approximately the same number of CTUs. Each tile contains an integer number of CTUs, which can be encoded independently, and during encoding, the CTUs within each tile are encoded in scan order. The main purpose of dividing into tiles is to enhance parallel processing capabilities without introducing new error propagation.

[0044] It should be noted that the following section introduces the division of macroblocks and coding tree units. H.264, H.265, and H.266 video coding all divide the video into smaller units based on blocks: a single frame of video (i.e., an image) is divided into different blocks, and then each block is encoded separately. From H.264 to H.265, and then from H.265 to H.266, the block division precision increases progressively.

[0045] 1. In H.264 (AVC): the block name is MB

[0046] Each frame of a video is composed of several blocks, which are arranged in the form of slices. In H.264, a frame is first divided into 16x16 blocks of the same size, called macroblocks (MBs), and the size of the macroblocks in each frame is fixed at 16x16 squares.

[0047] 2. In H.265 (HEVC): the block name is CTU

[0048] Each frame of a video is composed of several blocks, which are arranged in the form of slices and tiles. In H.265, each frame is first divided into Coding Tree Units (CTUs). The size of a CTU can be 64x64, 32x32, or 16x16, and all CTUs are square (the size of the CTU for each frame is fixed, but the size of the CTU for different frames can vary). For example, a 64x64 CTU can be divided into four 32x32 Coding Units (CUs). Each 32x32 CTU can be further divided into four 16x16 CUs, and each 16x16 CTU can be further divided into four 8x8 CUs. At this point, the frame cannot be further divided into CUs.

[0049] 3. In H.266 (VVC): the block name is CTU

[0050] In H.266 (VVC), the concepts of CU, prediction unit (PU), and transformation unit (TU) will no longer be distinguished (CU can be further subdivided, but the subdivided CU will still be a CU). Subdivision will only be performed when the CU exceeds the transform block size limit; otherwise, no further subdivision will be performed, and prediction and transformation operations will be performed on the current subdivision.

[0051] It's important to note that the quantization parameter (QP) reflects the degree of spatial detail compression. Decreasing QP preserves most of the detail; increasing QP results in some detail loss, lower bitrate, but also increased image distortion and reduced quality. In other words, QP and bitrate are inversely proportional, and this inverse relationship becomes more pronounced as the complexity of the video source increases. For H.264 and H.265 encoding, the QP value ranges from 0 to 51.

[0052] In H.264 encoding, a smaller QP value results in a smaller quantization step size and higher quantization precision, meaning that for the same image quality, the amount of data generated may be larger. For every 6 increases in the QP value, the quantization step size doubles. The correspondence between H.264 and H.265 differs slightly from QP=31 onwards, but is essentially the same, and will not be elaborated upon here.

[0053] Understandably, in some scenarios, the encoded video bitrate may exceed the set value. Analysis reveals that these scenarios typically involve complex image quality, and the encoder quantization parameter QP_max is set too low, causing bitrate control to fail. While increasing the QP_max value usually resolves the issue, this reduces image quality. Therefore, more complex images require a smaller OQ (smaller stride, higher precision), and a smaller encoder QP results in higher image quality and a higher encoding bitrate.

[0054] Therefore, in order to better adjust the encoder QP value and meet the needs of high-quality user experience and low bitrate transmission, the following issues need to be addressed:

[0055] 1. How to reduce the impact of encoder QP value on bit rate.

[0056] The QP value of an encoder is often inversely proportional to its bit rate. Generally, QP value adjustments are constrained by the bit rate. An encoder's QP value is typically set within a range (min, max), and the encoder adjusts the QP value within this range based on changes in the bit rate.

[0057] However, the above adjustments are based on the encoding of the entire image. In most cases, the proportion of details in a single frame is very limited, as shown below. Figure 2 As shown, the details that are most noticeable to the human eye in the image are not in the sky, but in the details of the two trees (branches, leaves, etc.) on the left and right. Imagine if only the trees in the image were encoded with a low QP value (high image quality), while the other parts of the image still used their original QP values, then the impact on the encoder's bitrate would be minimized as much as possible.

[0058] 2. How to adjust the encoder QP value based on the dimensions of the coding units in the encoded image.

[0059] Understandably, for the detailed parts of an image, the encoder often performs dimensional segmentation encoding on the coding units. The more details there are, the deeper the segmentation. That is, the dimension of the coding units of the encoded image can reflect the complexity of the details of the encoded image. Therefore, how to adjust the QP value of the encoder according to the dimension of the coding units of the encoded image is also a technical problem that urgently needs to be solved.

[0060] However, the relevant technologies generally set the encoder QP value based on the following situations:

[0061] 1. Standardize the encoder QP value.

[0062] In this case, the encoder encodes the current original frame image (without distinguishing IPB frames), and the encoder QP value is either a constant or passively adjusted according to the encoder's bit rate within a given range of (min, max).

[0063] 2. Adjust the encoder QP value according to the different encoded frames of I / P / B frames.

[0064] It is evident that the encoder QP value proposed in the relevant technologies is an average concept, that is, it takes into account both the detailed and non-detailed parts of the image and makes a uniform adjustment within a given QP value range (min, max), and cannot achieve custom QP value adjustment for detailed areas in the image.

[0065] The solutions provided by the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0066] like Figure 1 As shown, Figure 1 This is a flowchart of an image encoding method according to an embodiment of the present disclosure, which includes the following steps.

[0067] Step S01: Obtain the coding unit information of the image to be encoded;

[0068] In some embodiments, the encoder obtains coding unit information after segmenting the image to be encoded into coding units. It is understood that the encoder segments the image to be encoded into coding units, a process that can be quickly implemented and the results fed back within the encoder hardware. The embodiments of this disclosure proceed with further data processing based on the processing results (i.e., coding unit information) from the encoder hardware itself.

[0069] Step S02: Determine at least one target region in the image to be encoded based on the encoding unit information;

[0070] In some embodiments, based on the coding unit information, the distribution of coding units with preset dimensional information values ​​in the image to be encoded is first determined, and then, based on the distribution, at least one target region is selected in the image to be encoded. The target region can be understood as a detail region in the image to be encoded.

[0071] Step S03: Encode the image of the target region using a first preset quantization parameter value.

[0072] In some embodiments, only the quantization parameter values ​​of the image in the target region are adjusted, thereby enhancing local image quality while minimizing negative impact on the bitrate. Therefore, this disclosure embodiment can support custom image quality configuration encoding for different regions within the same frame of the image to be encoded. Furthermore, this disclosure embodiment uses the encoding unit information (the depth of the dimension in which the encoding unit is divided) directly output by the encoder to determine image detail and complexity. From an encoding feasibility perspective, this is simpler and more reliable than introducing other image complexity algorithms to obtain image complexity results.

[0073] The image encoding method provided in this disclosure obtains encoding unit information of the image to be encoded and determines at least one target region in the image to be encoded based on the encoding unit information, thereby encoding the image of the target region with a first preset quantization parameter value. It can be seen that this disclosure improves the image quality of local details without affecting the overall image bitrate by adjusting only the quantization parameter value of the target region in the image to be encoded.

[0074] In some embodiments, it also includes:

[0075] The image to be encoded is encoded with a second preset quantization parameter value, wherein the second preset quantization parameter value is greater than the first preset quantization parameter value;

[0076] The encoded image of the target region and the image to be encoded are sent to the decoding end, so that the decoding end can combine the image of the target region and the image to be encoded and display them.

[0077] In some embodiments, the image to be encoded is encoded using a second preset quantization parameter value. For example, the second preset quantization parameter value can be the encoder's default quantization parameter value. It should be noted that, in this embodiment, the original image to be encoded is encoded so that it can be used as a base image for synthesis by the decoding end. The QP value of the original image to be encoded can be a default value or higher. For example, the average QP value of the original image to be encoded, used as the "base image," can be increased, for example, avg = 30, with a range set to (28, 32). Then, the original image is encoded while maintaining the same resolution, thereby reducing the bitrate required for the original image. It is understood that this embodiment focuses on highlighting the detailed areas in the image to be encoded without increasing the overall bitrate. Therefore, the QP value of the detailed areas will be lower than the QP value of the image to be encoded, used as the base image.

[0078] In some embodiments, the image to be encoded, used as a base map, and the target region selected as a detail enhancement area are separately packetized and sent to the decoding end via network or wired transmission.

[0079] In some embodiments, the decoding end includes an extended reality device;

[0080] Correspondingly, the decoding end combines the image of the target region with the image to be encoded and then displays it, including:

[0081] The decoding end combines the image of the target area with the image to be encoded and then displays it on the extended reality device.

[0082] In some embodiments, the decoding end can be a head-mounted display of an extended reality device. After generating the composite image, the decoding end needs to render the composite image on the head-mounted display of the extended reality device and finally display it on the head-mounted display.

[0083] In some embodiments, it also includes:

[0084] The coordinates of the target region in the image to be encoded are sent to the decoding end.

[0085] In some embodiments of this disclosure, the image corresponding to the target region needs to be mapped onto the image to be encoded, which is used as a base image, at the decoding end. Therefore, when sending the image of the target region, the coordinate position of the target region in the image to be encoded is also sent, including the diagonal coordinates of the target region.

[0086] In some embodiments, the decoding end combines the image of the target region with the image to be encoded and then displays the result, including:

[0087] The decoding end overlays the image of the target region onto the image to be encoded based on the coordinate position of the target region in the image to be encoded.

[0088] In some embodiments, the decoding end pastes the decoded target region image onto the corresponding position in the image to be encoded based on the diagonal coordinates of the parsed target region in the image to be encoded. The pasting code can be implemented as follows:

[0089] m_d3dContext->CopySubresourceRegion(sharedResource,0,x,y,z,winResource,0,&pRegion);

[0090] in,

[0091] Microsoft::WRL::ComPtr <id3d11texture2d>winResource_;

[0092] Microsoft::WRL::ComPtr <id3d11resource>sharedResource_;

[0093] sharedResource — for composite graphs

[0094] winResource - A target region image segmented from a given rectangle.

[0095] pRegion specifies the size of the target region image to be pasted.

[0096] x, y, z — Specifies the location of the texture point.

[0097] In some embodiments, obtaining the coding unit information of the image to be encoded includes:

[0098] Obtain the coding unit dimension information value of the image to be encoded, wherein the coding unit dimension information value includes at least one of coding unit type, coding unit size and coding unit segmentation mode.

[0099] In some embodiments, it should be noted that hardware encoder manufacturers provide an encoder software development kit (SDK) interface. After the user calls the SDK, the encoder unit dimension information value of the current encoded frame image can be obtained. The encoder unit dimension information value includes the encoder unit type, encoder unit size and encoder unit segmentation mode.

[0100] For example, if the dimension information value of a coding unit is (980, 0, 1, 0), the corresponding header data means: this coding unit is the 980th coding unit; its type cuType is 0, indicating that this CU is not a prediction unit; its size cuSize is 1, indicating that the size of this coding unit is 16*16; its partition mode partitionMode is 0, indicating that its partition mode is 2N*2N; where, the coding unit partition mode is as follows... Figure 3 As shown.

[0101] As can be seen from the above, the 980th coding unit (CU) is a 16*16 unit, and it is not further divided into smaller coding units.

[0102] As can be seen, by using the coding unit dimension information value, the number of coding units in a frame of image can be counted (as can be known from the serial number printed by CTB), the number of coding units of each size (in this embodiment of the disclosure, the number of small coding units is preferred, the smaller the size of the coding unit, the more image details in that part), and the number of coding units of different dimensions of "depth" (in this disclosure, the number of coding units in the N*N segmentation pattern is preferred, the granularity of this segmentation pattern is smaller, thus indicating that the image details in that part are more).

[0103] In some embodiments, determining at least one target region in the image to be encoded based on the encoding unit information includes:

[0104] Based on the coding unit information, the distribution of coding units with preset coding unit dimension information values ​​in the image to be encoded is determined;

[0105] Based on the distribution, at least one target region is selected in the image to be encoded.

[0106] In some embodiments, selecting at least one target region in the image to be encoded according to the distribution includes:

[0107] Based on the distribution, at least one coding unit concentration region with a preset value for the dimension information of the coding unit is selected in the image to be encoded.

[0108] In some embodiments, at least one target region in the image to be encoded can be determined based on the coding unit dimension information value. Specifically, such as Figure 4 As shown, the coding unit division of the image to be encoded can be seen using HEVC's bitrate analysis tool: First, the size of the coding unit is 16*16, and the current position of the mouse is the 16*16 coding unit position in the lower right corner; Second, the coding unit segmentation mode is N*N; Third, the smallest coding unit is 8*8; Fourth, the coordinate position of the mouse cursor in the screenshot is 1360x784, which is the coordinate value of the upper left vertex of the selected coding unit.

[0109] In some embodiments, the granularity can be chosen to be 8*8 (in real-world scenarios, different granularity standards such as 4*4 or smaller may be used for statistical analysis; no specific limitation is made here). That is, it is necessary to count how many N*N segmentation pattern coding units are in a 16*16 cuSize. For example, in a 1920*1080 image frame, if 2000 16*16 coding units are found, of which 1000 have an N*N segmentation pattern, then the final number of "smallest coding units of 8*8" is 1000*4 = 4000.

[0110] Then we can statistically determine the distribution of "the smallest coding unit is 8*8" in the image to be encoded (for example, we can use the row and column matrix statistical method to first count which rows have the most distribution, and then count which columns have the most distribution).

[0111] Finally, using as many small rectangles as possible, select the vast majority of granular coding units that meet the requirements into non-overlapping rectangles, such as... Figure 5 As shown.

[0112] In some embodiments, it also includes:

[0113] For each target region, the corresponding quantization parameter value is determined based on the coding unit dimension information value within the target region;

[0114] The image of the target region is encoded according to the quantization parameter value.

[0115] In some embodiments, for Figure 5 The image portion selected by the rectangular block can be executed as multiple encoding tasks. That is, on the encoder side, each identified rectangular block is used to create its own encoder thread, and different resolutions (corresponding to the size of the rectangular block) and QP values ​​are set respectively.

[0116] In some embodiments, the setting of the encoder QP value is related to image quality. The finer the encoder divides the coding units (8*8, 4*4, or even smaller), the more details there are in the image, and therefore the lower the QP value should be (the lower the QP value, the higher the image quality).

[0117] For example, in the original input image, eight rectangular blocks of different sizes are identified based on the granularity of the coding unit. Then, eight coding threads need to be created at the encoder to encode the image in each of these eight rectangular blocks. The encoder QP value configured for each rectangular block can be set according to its own details (granularity of coding unit segmentation) (the typical range of QP value is 1 to 51).

[0118] Therefore, the embodiments of this disclosure support arbitrary rectangular segmentation of the current frame image based on the dimensions of its coding units, adjusting only the coding quality of the detail regions.

[0119] In some embodiments, such as Figure 6 As shown, the flowchart provided in this embodiment includes the following at the encoding end:

[0120] The image to be encoded is divided into coding units, and the result after division is as follows: Figure 7 As shown, the image regions of the coding unit with a "deep" dimension are then selected using multiple non-overlapping rectangles, such as... Figure 8 As shown, the encoder QP value is then adjusted for the selected rectangular areas.

[0121] In some embodiments, such as Figure 6 As shown, the flowchart provided in this embodiment includes the following at the decoding end:

[0122] After decoding the image to be encoded and the selected area image, the images are stitched together. The selected area image is stitched onto the image to be encoded, and the stitched image is rendered and displayed on the head-mounted display of the extended reality device.

[0123] As can be seen, in this embodiment, the image to be encoded is arbitrarily divided into rectangles according to the depth of the coding unit segmentation dimension. The QP value of the encoder is adjusted in the selected area. That is, the higher the granularity of the coding unit, the more details the image has, and the smaller the QP value set during encoding, thus improving accuracy.

[0124] Therefore, the details selected in the frame will be preserved as much as possible, thus improving the image quality of that part. It should be noted that the rectangular area divided in this embodiment is not equally divided, nor is it a complete segmentation of a frame, but only a segmentation of the local area where details need to be improved.

[0125] Furthermore, while improving the image quality of local details, this embodiment minimizes the negative impact on bitrate, and may even reduce the overall bitrate of the original image (base image). This is because while the image quality of detailed areas is improved, the image quality of non-detailed areas can be preserved at the original level or appropriately reduced. Therefore, this embodiment solves the technical problem of how to encode detailed areas with high quality by dividing the dimension of the image's coding unit, thereby improving the image quality of local details.

[0126] It is understood that the technical effects of the image encoding method provided in this disclosure include the following:

[0127] 1. It can maximize the improvement of image quality in the details of the picture and avoid the loss of raw data at the beginning of media data stream processing (encoding stage);

[0128] 2. Preserve as much detail as possible where possible, providing high-quality encoding;

[0129] 3. For areas that are not sensitive to the human eye and lack detail, the encoder directly provides the conclusion (the depth of the segmentation dimension of the encoding unit), rather than relying on external algorithms to provide a "not so reliable" judgment.

[0130] 4. The application of QP value for "averaging" a frame of image has been changed. Instead, targeted QP adjustment is applied to the details in the image. That is, it supports custom processing of QP value for different regions in a frame of image.

[0131] In some embodiments, such as Figure 9 As shown, this disclosure also provides an image decoding method, including:

[0132] Step S11: Receive the image to be encoded and at least one target region image from the image to be encoded sent by the encoding end; wherein the target region is determined by the encoding end based on the encoding unit information of the image to be encoded;

[0133] Step S12: Decode the image to be encoded and the target region, and then combine the image of the target region with the image to be encoded for display.

[0134] In some embodiments, the step of combining the image of the target region with the image to be encoded and then displaying it includes:

[0135] The image of the target region is stitched onto the image to be encoded to obtain a stitched composite image;

[0136] The composite image is sent to an extended reality device for display.

[0137] In some embodiments, the image encoding method provided in this disclosure is applied to an encoder, and the method includes:

[0138] Obtain the image data to be encoded;

[0139] Perform dimensional partitioning analysis on the image data to be encoded into coding units;

[0140] Based on the results of the dimensional analysis of the encoding unit, the areas where the detail quality needs to be improved are extracted;

[0141] Encode the extracted region with a specified QP value;

[0142] The original image to be encoded is encoded using the default QP value and used as the base image;

[0143] The base map data and the selected enhancement areas are packaged and sent separately to the decoder via network or wired transmission.

[0144] In some embodiments, the encoder first acquires the image data of the current frame to be encoded. Then, it obtains the dimensional information value of the encoding unit and performs dimensional segmentation analysis of the encoding unit in the image data to be encoded based on the dimensional information value. A standard for the granularity of the encoding unit is determined through a custom method, thereby confirming the distribution of encoding units at the selected granularity. A plurality of rectangles, as small as possible, are used to select most of the granularity encoding units that meet the requirements into non-overlapping rectangles. The diagonal coordinates of each selected rectangle are output for texture copying of the corresponding coordinate points after image segmentation and encoding / decoding. Finally, each rectangle is processed independently, i.e., each rectangle is encoded independently. When sending the packet, the packet header needs to include the diagonal coordinate position of each rectangular region corresponding to the original image, for use in texture copying after decoding at the decoding end.

[0145] In some embodiments, the image decoding method provided in this disclosure is applied at the decoder end, and the method includes:

[0146] The decoder performs packet reception and depacketization.

[0147] The decoder decodes the base image and the enhanced detail area image separately;

[0148] The decoded detail-enhanced region image is stitched onto the base image;

[0149] The stitched composite image is sent to the head-mounted display of the extended reality device for rendering and display.

[0150] In some embodiments, the decoding end requires a "base map" that can have a high QP value while maintaining the same resolution.

[0151] This disclosure also provides an image encoding apparatus, including:

[0152] The acquisition module is used to acquire the coding unit information of the image to be encoded;

[0153] The first processing module is used to determine at least one target region in the image to be encoded based on the encoding unit information;

[0154] The second processing module is used to encode the image of the target region using a first preset quantization parameter value.

[0155] In some embodiments, the second processing module is further specifically used for:

[0156] The image to be encoded is encoded with a second preset quantization parameter value, wherein the second preset quantization parameter value is greater than the first preset quantization parameter value;

[0157] The encoded image of the target region and the image to be encoded are sent to the decoding end, so that the decoding end can combine the image of the target region and the image to be encoded and display them.

[0158] In some embodiments, the decoding end includes an extended reality device; the second processing module is further specifically used for:

[0159] The decoding end combines the image of the target area with the image to be encoded and then displays it on the extended reality device.

[0160] In some embodiments, the second processing module is further specifically used for:

[0161] The coordinates of the target region in the image to be encoded are sent to the decoding end.

[0162] In some embodiments, the second processing module is further specifically used for:

[0163] The decoding end overlays the image of the target region onto the image to be encoded based on the coordinate position of the target region in the image to be encoded.

[0164] In some embodiments, the acquisition module is specifically used for:

[0165] Obtain the coding unit dimension information value of the image to be encoded, wherein the coding unit dimension information value includes at least one of coding unit type, coding unit size and coding unit segmentation mode.

[0166] In some embodiments, the first processing module is specifically used for:

[0167] Based on the coding unit information, the distribution of coding units with preset coding unit dimension information values ​​in the image to be encoded is determined;

[0168] Based on the distribution, at least one target region is selected in the image to be encoded.

[0169] In some embodiments, the first processing module is further specifically used for:

[0170] Based on the distribution, at least one coding unit concentration region with a preset value for the dimension information of the coding unit is selected in the image to be encoded.

[0171] In some embodiments, the second processing module is further specifically used for:

[0172] For each target region, the corresponding quantization parameter value is determined based on the coding unit dimension information value within the target region;

[0173] The image of the target region is encoded according to the quantization parameter value.

[0174] This disclosure also provides an image decoding apparatus, including:

[0175] A receiving module is used to receive an image to be encoded and an image of at least one target region in the image to be encoded, sent by an encoding end; the target region is determined by the encoding end based on the encoding unit information of the image to be encoded.

[0176] The third processing module is used to decode the image to be encoded and the target region, and to combine the image of the target region with the image to be encoded for display.

[0177] In some embodiments, the third processing module is specifically used for:

[0178] The image of the target region is stitched onto the image to be encoded to obtain a stitched composite image;

[0179] The composite image is sent to an extended reality device for display.

[0180] For embodiments of the apparatus, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative, and the modules described as separate modules may or may not be separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0181] The methods and apparatus of this disclosure have been described above based on embodiments and application examples. Furthermore, this disclosure also provides an electronic device and a computer-readable storage medium, which are described below.

[0182] The following is for reference. Figure 10 The figure illustrates a structural schematic of an electronic device (e.g., a terminal device or server) 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in the figure is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.

[0183] Electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from storage device 808 into random access memory (RAM) 803. RAM 803 also stores various programs and data required for the operation of electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0184] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although an electronic device 800 with various devices is shown in the figure, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0185] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.

[0186] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0187] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0188] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0189] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods of the present disclosure.

[0190] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0192] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0193] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0194] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0195] According to one or more embodiments of this disclosure, an image encoding method is provided, comprising:

[0196] Obtain the coding unit information of the image to be encoded;

[0197] Based on the encoding unit information, at least one target region in the image to be encoded is determined;

[0198] The image of the target region is encoded using a first preset quantization parameter value.

[0199] According to one or more embodiments of this disclosure, a method is provided, further comprising:

[0200] The image to be encoded is encoded with a second preset quantization parameter value, wherein the second preset quantization parameter value is greater than the first preset quantization parameter value;

[0201] The encoded image of the target region and the image to be encoded are sent to the decoding end, so that the decoding end can combine the image of the target region and the image to be encoded and display them.

[0202] According to one or more embodiments of this disclosure, a method is provided in which the decoding end includes an extended reality device;

[0203] Correspondingly, the decoding end combines the image of the target region with the image to be encoded and then displays it, including:

[0204] The decoding end combines the image of the target area with the image to be encoded and then displays it on the extended reality device.

[0205] According to one or more embodiments of this disclosure, a method is provided, further comprising:

[0206] The coordinates of the target region in the image to be encoded are sent to the decoding end.

[0207] According to one or more embodiments of this disclosure, a method is provided in which the decoding end combines an image of the target region with the image to be encoded and then displays the result, including:

[0208] The decoding end overlays the image of the target region onto the image to be encoded based on the coordinate position of the target region in the image to be encoded.

[0209] According to one or more embodiments of this disclosure, a method is provided in which obtaining coding unit information of an image to be encoded includes:

[0210] Obtain the coding unit dimension information value of the image to be encoded, wherein the coding unit dimension information value includes at least one of coding unit type, coding unit size and coding unit segmentation mode.

[0211] According to one or more embodiments of this disclosure, a method is provided in which determining at least one target region in an image to be encoded based on the encoding unit information includes:

[0212] Based on the coding unit information, the distribution of coding units with preset coding unit dimension information values ​​in the image to be encoded is determined;

[0213] Based on the distribution, at least one target region is selected in the image to be encoded.

[0214] According to one or more embodiments of this disclosure, a method is provided for selecting at least one target region in the image to be encoded based on the distribution, comprising:

[0215] Based on the distribution, at least one coding unit concentration region with a preset value for the dimension information of the coding unit is selected in the image to be encoded.

[0216] According to one or more embodiments of this disclosure, a method is provided, further comprising:

[0217] For each target region, the corresponding quantization parameter value is determined based on the coding unit dimension information value within the target region;

[0218] The image of the target region is encoded according to the quantization parameter value.

[0219] According to one or more embodiments of this disclosure, an image encoding apparatus is provided, comprising:

[0220] The acquisition module is used to acquire the coding unit information of the image to be encoded;

[0221] The first processing module is used to determine at least one target region in the image to be encoded based on the encoding unit information;

[0222] The second processing module is used to encode the image of the target region using a first preset quantization parameter value.

[0223] According to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one memory and at least one processor;

[0224] The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method described in any one of the above.

[0225] According to one or more embodiments of the present disclosure, a computer-readable storage medium is provided for storing program code that, when executed by a processor, causes the processor to perform the methods described above.

[0226] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0227] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0228] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An image encoding method, characterized in that, include: Obtaining coding unit information of the image to be encoded includes: obtaining the coding unit dimension information value of the image to be encoded; Determining at least one target region in the image to be encoded based on the encoding unit information includes: determining the distribution of encoding units with encoding unit dimension information values ​​of preset values ​​in the image to be encoded based on the encoding unit information; and selecting at least one target region in the image to be encoded based on the distribution. The image of the target region is encoded using a first preset quantization parameter value; The step of selecting at least one target region in the image to be encoded according to the distribution includes: selecting at least one encoding unit concentration region in the image to be encoded with an encoding unit dimension information value of a preset value according to the distribution. The coding unit dimension information value includes at least one of coding unit type, coding unit size, and coding unit segmentation mode.

2. The method according to claim 1, characterized in that, Also includes: The image to be encoded is encoded with a second preset quantization parameter value, wherein the second preset quantization parameter value is greater than the first preset quantization parameter value; The encoded image of the target region and the image to be encoded are sent to the decoding end, so that the decoding end can combine the image of the target region and the image to be encoded and display them.

3. The method according to claim 2, characterized in that, The decoding end includes an extended reality device; Correspondingly, the decoding end combines the image of the target region with the image to be encoded and then displays it, including: The decoding end combines the image of the target area with the image to be encoded and then displays it on the extended reality device.

4. The method according to claim 2, characterized in that, Also includes: The coordinates of the target region in the image to be encoded are sent to the decoding end.

5. The method according to claim 4, characterized in that, The decoding end combines the image of the target region with the image to be encoded and then displays the result, including: The decoding end overlays the image of the target region onto the image to be encoded based on the coordinate position of the target region in the image to be encoded.

6. The method according to claim 1, characterized in that, Also includes: For each target region, the corresponding quantization parameter value is determined based on the coding unit dimension information value within the target region; The image of the target region is encoded according to the quantization parameter value.

7. An image decoding method, characterized in that, include: Receive the image to be encoded and at least one target region image in the image to be encoded sent by the encoding end; The image to be encoded and the target region are decoded, and the image of the target region is combined with the image to be encoded and then displayed. The target region is determined by the encoding end based on the encoding unit information of the image to be encoded, including: determining the distribution of encoding units with encoding unit dimension information values ​​of preset values ​​in the image to be encoded based on the encoding unit information; and selecting at least one target region in the image to be encoded based on the distribution. The step of selecting at least one target region in the image to be encoded according to the distribution includes: selecting at least one encoding unit concentration region in the image to be encoded with an encoding unit dimension information value of a preset value according to the distribution. The coding unit dimension information value includes at least one of coding unit type, coding unit size, and coding unit segmentation mode.

8. The method according to claim 7, characterized in that, The step of combining the image of the target region with the image to be encoded and then displaying it includes: The image of the target region is stitched onto the image to be encoded to obtain a stitched composite image; The composite image is sent to an extended reality device for display.

9. An image encoding device, characterized in that, include: The acquisition module is used to acquire coding unit information of the image to be encoded, including: acquiring the coding unit dimension information value of the image to be encoded; The first processing module is configured to determine at least one target region in the image to be encoded based on the encoding unit information, including: determining the distribution of encoding units with encoding unit dimension information values ​​of preset values ​​in the image to be encoded based on the encoding unit information; and selecting at least one target region in the image to be encoded based on the distribution. The second processing module is used to encode the image of the target region with a first preset quantization parameter value; The step of selecting at least one target region in the image to be encoded according to the distribution includes: selecting at least one encoding unit concentration region in the image to be encoded with an encoding unit dimension information value of a preset value according to the distribution. The coding unit dimension information value includes at least one of coding unit type, coding unit size, and coding unit segmentation mode.

10. An image decoding device, characterized in that, include: A receiving module is used to receive an image to be encoded and an image of at least one target region in the image to be encoded, sent by the encoding end. The third processing module is used to decode the image to be encoded and the target region, and to combine the image of the target region with the image to be encoded for display. The target region is determined by the encoding end based on the encoding unit information of the image to be encoded, including: determining the distribution of encoding units with encoding unit dimension information values ​​of preset values ​​in the image to be encoded based on the encoding unit information; and selecting at least one target region in the image to be encoded based on the distribution. The step of selecting at least one target region in the image to be encoded according to the distribution includes: selecting at least one encoding unit concentration region in the image to be encoded with an encoding unit dimension information value of a preset value according to the distribution. The coding unit dimension information value includes at least one of coding unit type, coding unit size, and coding unit segmentation mode.

11. An electronic device, comprising: At least one memory and at least one processor; The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method of any one of claims 1 to 8.

12. A computer-readable storage medium for storing program code that, when executed by a computer device, causes the computer device to perform the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Quantization method of coding block and quantization method for HEVC (High Efficient Video Coding)

    CN109660803A

  • Video decoding method and video decoder

    CN110881129A

  • Image processing method, and training method and device of image processing model

    CN114743121A

  • Method, apparatus and system for encoding and decoding video data

    KR1020100002632A