Moving image encoding device and moving image decoding device

JP2025081720AActive Publication Date: 2025-05-27KK TOSHIBA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025031669
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27
Estimated Expiration
2029-06-18

AI Technical Summary

Technical Problem

In the direct mode of motion compensation prediction, the degree of freedom in calculating motion vectors is low, and the amount of code for motion vector selection information is increased, leading to inefficiencies in encoding and decoding moving images.

Method used

A moving image encoding and decoding apparatus that selects one encoded block from adjacent blocks with motion vectors, encodes the selection information using a code table corresponding to the number of available blocks, and performs motion compensation prediction using the selected block's motion vector.

Benefits of technology

This approach increases the degree of freedom in calculating motion vectors while reducing the additional information required for motion vector selection, thereby improving the efficiency of motion compensation prediction encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081720000001_ABST
    Figure 2025081720000001_ABST
Patent Text Reader

Abstract

To provide a moving image encoding device and a moving image decoding device for reducing additional information of motion vector selection information while selecting one of encoded blocks to increase a degree of freedom of motion vector calculation.SOLUTION: A moving image encoding device for performing motion compensative prediction encoding of a moving image includes an acquisition part / selection part 110 for acquiring available blocks having a motion vector and the number of the available blocks from encoded blocks adjacent to an encoding object block, and selecting one selection block from the encoded block available blocks, a variable length encoder 111 having a selection information encoding part 112 for encoding selection information for identifying the selection block by using a code table corresponding to the number of the available blocks, and a multiplexer 113 for multiplexing encoded quantization orthogonal transform coefficient information and selection information, and outputting encoded data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a moving image encoding apparatus and a moving image decoding apparatus that obtain motion vectors from encoded and decoded images and perform motion compensation prediction.

Background Art

[0002] One of the techniques used for encoding moving images is motion compensation prediction. In motion compensation prediction, a motion vector is obtained using an image to be encoded that is newly to be encoded and a locally decoded image that has already been obtained in a moving image encoding apparatus, and a predicted image is generated by performing motion compensation using this motion vector.

[0003] As one of the methods for obtaining a motion vector in motion compensation prediction, there is a direct mode in which a predicted image is generated using a motion vector of a block to be encoded derived from a motion vector of an already encoded block (see Japanese Patent No. 4020789 and U.S. Patent No. 7233621). In the direct mode, since the motion vector is not encoded, the amount of code for the motion vector information can be reduced. The direct mode is adopted in H.264 / AVC.

Summary of the Invention

[0004] In the direct mode, when predicting and generating a motion vector of a block to be encoded, the motion vector is generated by a fixed method of calculating the motion vector from the median value of the motion vectors of the already encoded blocks adjacent to the block to be encoded. Therefore, the degree of freedom in calculating the motion vector is low. Further, in order to increase the above-mentioned degree of freedom, when using a method of calculating a motion vector that selects one from a plurality of already encoded blocks, in order to indicate the selected already encoded block, the position of the block always has to be sent as motion vector selection information. This causes an increase in the amount of code.

[0005] An object of the present invention is to provide a moving image encoding apparatus and a moving image decoding apparatus that select one from encoded blocks to increase the degree of freedom in calculating motion vectors while reducing additional information of motion vector selection information.

[0006] One aspect of the present invention is a moving image encoding apparatus that performs motion compensation prediction encoding on a moving image, and includes an acquisition unit that obtains an available block that is an encoded block adjacent to an encoding target block and has a motion vector, and the number of the available blocks, a selection unit that selects one selected block from the available blocks of the encoded block, a selection information encoding unit that encodes selection information for specifying the selected block using a code table corresponding to the number of the available blocks, and an image encoding unit that performs motion compensation prediction encoding on the encoding target block using the motion vector of the selected block. A moving image encoding apparatus is provided.

[0007] Another aspect of the present invention is a moving image decoding apparatus that performs motion compensation prediction decoding on a moving image, and includes a selection information decoding unit that switches a code table according to the number of available blocks that are decoded blocks adjacent to a decoding target block and have motion vectors, and decodes selection information, and a motion vector selection unit that selects one motion vector indicated by the selection information decoded by the selection information decoding unit from the available blocks, and an image decoding unit that performs motion compensation prediction decoding on a decoding target image using the motion vector selected by the motion vector selection unit. A moving image decoding apparatus is provided.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0010] Referring to FIG. 1, a moving image encoding apparatus according to an embodiment will be described. The subtractor 101 is configured to calculate the difference between the input moving image signal 11 and the predicted image signal 15 and output a prediction error signal 12. The output terminal of the subtractor 101 is connected to the variable length encoder 111 via the orthogonal transformer 102 and the quantizer 103. The orthogonal transformer 102 orthogonally transforms the prediction error signal 12 from the subtractor 101 to generate orthogonal transformation coefficients, and the quantizer 103 quantizes the orthogonal transformation coefficients and outputs quantized orthogonal transformation coefficient information 13. The variable length encoder 111 variably encodes the quantized orthogonal transformation coefficient information 13 from the quantizer 103.

[0011] The output terminal of the quantizer 103 is connected to the adder 106 via the inverse quantizer 104 and the inverse orthogonal transformer 105. The inverse quantizer 104 inverse quantizes the quantized orthogonal transform coefficient information 13 and converts it into orthogonal transform coefficients. The inverse orthogonal transformer 105 converts the orthogonal transform coefficients into a prediction error signal. The adder 106 adds the prediction error signal from the inverse orthogonal transformer 105 and the predicted image signal 15 to generate a local decoded image signal 14. The output terminal of the adder 105 is connected to the motion compensation predictor 108 via the frame memory 107.

[0012] The frame memory 107 stores the local decoded image signal 14. The setting unit 114 sets the motion compensation prediction mode (prediction mode) of the block to be encoded. The prediction mode includes unidirectional prediction using one reference image and bidirectional prediction using two reference images. The unidirectional prediction includes L0 prediction and L1 prediction of AVC. The motion compensation predictor 108 includes a predictor 109 and an acquisition / selection unit 110. The acquisition / selection unit 110 obtains an available block having a motion vector and the number of available blocks from the encoded blocks adjacent to the block to be encoded, and selects one selected block from the available blocks. The motion compensation predictor 108 generates a predicted image signal 15 from the local decoded image signal 14 and the input moving image signal 11 in the frame memory 107. The acquisition / selection unit 110 selects one block (selected block) from the adjacent blocks adjacent to the block to be encoded. For example, among the adjacent blocks, a block having an appropriate motion vector is selected as the selected block. The acquisition / selection unit 110 selects the motion vector of the selected block as the motion vector 16 used for motion compensation prediction and sends it to the predictor 109. In addition, the acquisition / selection unit 110 generates selection information 17 of the selected block and sends it to the variable length encoder 111.

[0013] The variable-length encoder 111 has a selection information encoding unit 112. The selection information encoding unit 112 variably encodes the selection information 17 while switching the code table so that the code table has the same number of entries as the number of available blocks in the encoded block. An available block is a block having a motion vector among the encoded blocks adjacent to the block to be encoded. The multiplexer 113 multiplexes the encoded quantized orthogonal transform coefficient information and the selection information, and outputs the encoded data.

[0014] The operation of the moving image encoding apparatus having the above configuration will be described with reference to the flowchart of FIG. 2.

[0015] First, a prediction error signal 12 is generated (S11). In the generation of this prediction error signal 12, a motion vector is selected, and a prediction image is generated using the selected motion vector. The difference between the signal of this prediction image, that is, the prediction image signal 15 and the input moving image signal 11 is calculated by the subtracter 101, whereby the prediction error signal 12 is generated.

[0016] The prediction error signal 12 is subjected to an orthogonal transform by the orthogonal transformer 102 to generate orthogonal transform coefficients (S12). The orthogonal transform coefficients are quantized by the quantizer 103 (S13). The quantized orthogonal transform coefficient information is inverse quantized by the inverse quantizer 104 (S14), and then inverse transformed by the inverse orthogonal transformer 105 to obtain a reproduced prediction error signal (S15). In the adder 106, the reproduced prediction error signal and the prediction image signal 15 are added to generate a local decoded image signal 14 (S16). The local decoded image signal 14 is stored (as a reference image) in the frame memory 107 (S17), and the local decoded image signal read from the frame memory 107 is input to the motion compensation predictor 108.

[0017] The predictor 109 of the motion compensation predictor 108 performs motion compensation prediction on the local decoded image signal (reference image) using the motion vector 16 to generate the predicted image signal 15. The predicted image signal 15 is sent to the subtracter 101 to obtain the difference from the input moving image signal 11, and is also sent to the adder 106 to generate the local decoded image signal 14.

[0018] The acquisition / selection unit 110 selects one motion vector used for motion compensation prediction from adjacent blocks, sends the selected motion vector 16 to the predictor 109, and generates selection information 17. The selection information 17 is sent to the selection information encoding unit 112. When selecting a motion vector from adjacent blocks, an appropriate motion vector with a smaller coding amount can be selected.

[0019] The orthogonal transform coefficient information 13 quantized by the quantizer 103 is also input to the variable length encoder 111 and undergoes variable length coding (S18). The acquisition / selection unit 110 outputs the selection information 16 used for motion compensation prediction and inputs it to the selection information encoding unit 112. In the selection information encoding unit 112, the code table is switched so that the code table has the same number of entries as the number of available blocks that are adjacent to the block to be encoded and are encoded blocks with motion vectors, and the selection information 17 is variable length encoded. The quantized orthogonal transform coefficient information and the selection information from the variable length encoder 111 are multiplexed by the multiplexer 113, and a bit stream of the encoded data 18 is output (S19). The encoded data 18 is sent to an accumulation system or a transmission path (not shown).

[0020] In the flowchart of FIG. 2, the flow of steps S14 to S17 and the flow of steps S18 and S19 can be replaced. That is, following the quantization step S13, the variable length coding step S18 and the multiplexing step S19 are performed, and the inverse quantization step S14 to the storage step S17 can be performed in pairs for the multiplexing step S19.

[0021] Next, the operation of the acquisition / selection unit 110 will be described using the flowchart shown in FIG. 3.

[0022] First, with reference to the frame memory 107, search for candidate available blocks which are encoded blocks with motion vectors adjacent to the block to be encoded (S101). When candidate available blocks are searched, the block sizes of motion compensation prediction for these candidate available blocks are determined (S102). Next, it is determined whether the candidate available blocks are for uni-directional or bi-directional prediction (S103). Based on the determination result and the prediction mode of the block to be encoded, available blocks are extracted from among the candidate available blocks. One selected block is selected from the extracted available blocks, and information for specifying the selected block is obtained as selection information (S104).

[0023] Next, with reference to FIGS. 4A to 4C, the determination of the block size (S102) will be described.

[0024] The adjacent blocks used in this embodiment are the blocks located to the left, upper left, above, and upper right of the block to be encoded. Therefore, when the block to be encoded is located at the upper leftmost of the frame, there are no available blocks adjacent to the block to be encoded, and thus the present invention cannot be applied to this block to be encoded. When the block to be encoded is at the upper end of the screen, the available block is only the leftmost one block, and when the block to be encoded is at the left end of the screen and not at the right end, the available blocks are the two blocks above and upper right of the block to be encoded.

[0025] When the macro block size is 16x16, the block sizes for motion compensation prediction of adjacent blocks are four types: 16x16, 16x8, 8x16, and 8x8, as shown in FIGS. 4A to 4C. When considering these four types, the adjacent blocks that can be available blocks are 20 types as shown in FIGS. 4A to 4C. That is, there are 4 types of 16x16 size shown in FIG. 4A, 10 types of 16x8 and 8x16 sizes shown in FIG. 4B, and 6 types of 8x8 size shown in FIG. 4C. In the block size determination (S102), available blocks are searched according to the block size from among these 20 types of blocks. For example, when the size of the available block is only 16x16, the available blocks determined by this block size are 4 types of 16x16 blocks as shown in FIG. 4A. That is, the available blocks are the block on the upper left side of the block to be encoded, the block on the upper side of the block to be encoded, the block on the left side of the block to be encoded, and the block on the upper right side of the block to be encoded. Also, when the macro block size is extended to 16x16 or more, it can be an available block in the same way as when the macro block size is 16x16. For example, when the macro block size is 32x32, the block sizes for motion compensation prediction of adjacent blocks are four types: 32x32, 32x16, 16x32, and 16x16, and the adjacent blocks that can be available blocks are 20 types.

[0026] Next, with reference to FIG. 5, the determination (S103) of unidirectional or bidirectional prediction performed by the acquisition / selection unit 110 will be described with examples.

[0027] For example, assume that the block size is restricted to 16x16, and for the block to be encoded, the unidirectional or bidirectional prediction of adjacent blocks is as shown in FIG. 5. In the determination of unidirectional or bidirectional prediction (S103), available blocks are searched according to the prediction direction. For example, adjacent blocks including the prediction direction L0 are defined as the available blocks determined by the prediction direction. That is, the blocks above, to the left, and to the upper right of the block to be encoded shown in FIG. 5(a) are the available blocks determined by the prediction direction. In this case, the upper left block of the block to be encoded is not used. If adjacent blocks including the prediction direction L1 are defined as the available blocks determined by the prediction method, the blocks to the upper left and above the block to be encoded shown in FIG. 5(b) are the available blocks determined by the prediction direction. In this case, the left and upper right blocks of the block to be encoded are not used. If adjacent blocks including the prediction direction L0 / L1 are defined as the available blocks determined by the prediction method, only the block above the block to be encoded shown in FIG. 5(c) is the available block determined by the prediction direction. In this case, the left, upper left, and upper right blocks of the block to be encoded are not used. Note that the prediction direction L0 (L1) corresponds to the prediction direction of L0 prediction (L1 prediction) in AVC.

[0028] Next, the selection information encoding unit 112 will be described with reference to the flowchart shown in FIG. 6.

[0029] An available block, which is an encoded block with a motion vector, is searched from among adjacent blocks adjacent to the block to be encoded, and available block information determined by the block size and unidirectional or bidirectional prediction is obtained (S201). Using this available block information, a switching of the code table according to the number of available blocks as shown in FIG. 8 is performed (S202). Using the switched code table, the selection information 17 sent from the acquisition / selection unit 110 is variable-length encoded (S203).

[0030] Next, an example of the index of the selection information will be described with reference to FIG. 7.

[0031] When there is no available block as shown in Fig. 7(a), the present invention cannot be applied to this block, so no selection information is sent. When there is one available block as shown in Fig. 7(b), since the motion vector of the available block used for motion compensation of the block to be coded is uniquely determined, no selection information is sent. When there are two available blocks as shown in Fig. 7(c), selection information with index 0 or 1 is sent. When there are three available blocks as shown in Fig. 7(d), selection information with index 0, 1 or 2 is sent. When there are four available blocks as shown in Fig. 7(e), selection information with index 0, 1, 2 or 3 is sent.

[0032] Also, as an example of how to index available blocks, Fig. 7 shows an example in which available blocks are indexed in the order of left, upper left, upper, and upper right of the block to be coded. That is, indexes are assigned in sequence to the blocks that are used except for the blocks that are not used.

[0033] Next, the code table of selection information 17 will be described with reference to Fig. 8.

[0034] The selection information encoding unit 112 switches the code table according to the number of available blocks (S202). As described above, it is necessary to encode the selection information 17 when there are two or more available blocks.

[0035] First, when there are two available blocks, indexes 0 and 1 are required, and the code table is the table shown on the left side of Fig. 8. When there are three available blocks, indexes 0, 1, and 2 are required, and the code table is the table shown in the center of Fig. 8. When there are four available blocks, indexes 0, 1, 2, and 3 are required, and the code table is the table shown on the right side of Fig. 8. These code tables are switched according to the number of available blocks.

[0036] Next, the encoding method of the selection information will be described.

[0037] Figure 9 shows an overview of the syntax structure used in this embodiment. The syntax mainly consists of three parts. High Level Syntax 801 contains syntax information of upper layers above the slice. In Slice Level Syntax 804, the information required for each slice is specified. In Macroblock Level Syntax 807, the variable-length coded error signals, mode information, etc. required for each macroblock are specified.

[0038] These syntaxes are each composed of more detailed syntaxes. High Level Syntax 801 is composed of sequence and picture level syntaxes such as Sequence parameter set syntax 802 and Picture parameter set syntax 803. Slice Level Syntax 804 consists of Slice header syntax 405, Slice data syntax 406, etc. Furthermore, Macroblock Level Syntax 807 is composed of macroblock layer syntax 808, macroblock prediction syntax 809, etc.

[0039] The syntax information required in this embodiment is macroblock layer syntax 808, and the syntax will be described below.

[0040] The available_block_num shown in FIGS. 10(a) and 10(b) indicates the number of available blocks. When this is 2 or more, encoding of selection information is required. The mvcopy_flag is a flag indicating whether to use the motion vectors of available blocks in motion compensation prediction. When there is 1 or more available blocks and this flag is 1, the motion vectors of available blocks can be used in motion compensation prediction. Furthermore, mv_select_info indicates selection information, and the code table is as described above.

[0041] Figure 10(a) shows the syntax when encoding selection information after mb_type. For example, when the block size is only 16x16, if mb_type is other than 16x16, then mvcopy_flag and mv_select_info do not need to be encoded. If mb_type is 16x16, then mvcopy_flag and mv_select_info are encoded.

[0042] Figure 10(b) shows the syntax when encoding selection information before mb_type. For example, if mvcopy_flag is 1, then there is no need to encode mb_type. If mv_copy_flag is 0, then mb_type is encoded.

[0043] In this embodiment, the encoding scan order can be in any order. For example, the present invention is applicable to line scan, Z scan, etc.

[0044] A moving image decoder according to another embodiment will be described with reference to FIG. 11.

[0045] The encoded data 18 output from the moving image encoder of FIG. 1 enters the demultiplexer 201 of the moving image decoder as the encoded data 21 to be decoded through the storage system or the transmission system. The demultiplexer (demultiplexer) 201 demultiplexes the encoded data 21 and separates the encoded data 21 into quantized orthogonal transform coefficient information and selection information. The output terminal of the demultiplexer 201 is connected to the variable length decoder 202. The variable length decoder 202 decodes the quantized orthogonal transform coefficient information and the selection information. The output terminal of the variable length decoder 202 connects the inverse quantizer 204 and the inverse orthogonal transformer 205 to the adder 206. The inverse quantizer 204 inverse quantizes the quantized orthogonal transform coefficient information and converts it into orthogonal transform coefficients. The inverse orthogonal transformer 205 inverse orthogonal transforms the orthogonal transform coefficients and generates a prediction error signal. The adder 206 adds the prediction error signal to the prediction image signal from the prediction image generator 207 to generate a moving image signal.

[0046] The prediction image generator 207 includes a predictor 208 and a selector 209. The selector 209 selects a motion vector based on the selection information 23 decoded by the selection information decoder 203 of the variable-length decoder 202, and sends the selected motion vector 25 to the predictor 208. The predictor 208 performs motion compensation on the reference image stored in the frame memory 210 using the motion vector 25, and generates a prediction image.

[0047] The operation of the moving image decoder with the above configuration will be described with reference to the flowchart of FIG. 12.

[0048] The encoded data 21 is demultiplexed by the demultiplexer 201 (S31), decoded by the variable-length decoder 202, and quantized orthogonal transform coefficient information 22 is generated (S32). Also, the state of adjacent blocks adjacent to the block to be decoded is investigated by the selection information decoder 203, and the code table is switched and decoded as in the selection information encoding unit 112 of the encoding device according to the number of available blocks that are adjacent encoded blocks having motion vectors. As a result, selection information 23 is output (S33).

[0049] The quantized orthogonal transform coefficient information 22, which is the information output from the variable decoder 202, is sent to the inverse quantizer 204, and the selection information 23, which is the information output from the selection information decoder 203, is sent to the selector 209.

[0050] The quantized orthogonal transform coefficient information 22 is inverse quantized by the inverse quantizer 204 (S34), and then inverse orthogonally transformed by the inverse orthogonal transformer 205 (S35). As a result, a prediction error signal 24 is obtained. In the adder 206, a moving image signal 26 is reproduced by adding a prediction image signal to the prediction error signal 24 (S36). The reproduced moving image signal 27 is stored in the frame memory 210 (S37).

[0051] In the prediction image generator 207, a prediction image 26 is generated using the motion vectors of available blocks, which are decoded blocks adjacent to the block to be decoded and have motion vectors, selected by the decoded selection information 23. In the selection unit 209, the states of adjacent blocks are investigated, and one motion vector for motion compensation prediction is selected from adjacent blocks in the same manner as the acquisition unit / selection unit 110 of the encoding device, based on the available block information of the adjacent blocks and the selection information 23 decoded by the selection information decoding unit 203. Using this selected motion vector 25, the predictor 208 generates a prediction image 26, which is sent to the adder 206 to obtain a moving image signal 27.

[0052] According to the present invention, by encoding selection information corresponding to the number of available blocks, the selection information can be transmitted using an appropriate encoding table, and additional information of the selection information can be reduced.

[0053] Also, by using the motion vectors of available blocks for motion compensation prediction of the block to be encoded, additional information regarding motion vector information can be reduced.

[0054] Furthermore, by selecting an appropriate one from available blocks instead of fixing the motion vector calculation method, the degree of freedom in motion vector calculation becomes higher than in the direct mode.

[0055] The method of the present invention described in the embodiments of the present invention can be executed by a computer, and can also be stored and distributed as a program executable by a computer on a recording medium such as a magnetic disk (flexible disk, hard disk, etc.), an optical disk (CD-ROM, DVD, etc.), or a semiconductor memory.

[0056] Furthermore, the present invention is not limited to the above-described embodiments as they are, and at the implementation stage, the components can be modified and embodied without departing from the gist thereof. Also, various inventions can be formed by appropriately combining a plurality of components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

Industrial Applicability

[0057] The apparatus of the present invention is used for image compression processing in communication, storage, and broadcasting.

Claims

1. A transmission circuit for transmitting the encoded data, the encoded data includes first encoded data and second encoded data; the first encoded data is generated by obtaining available blocks, which are blocks having motion vectors, from a plurality of encoded blocks adjacent to a block to be encoded, selecting one block from the available blocks, selecting one code table from a plurality of code tables according to the number of the available blocks, and encoding identification information for identifying the selected block using the selected code table; the second encoded data is generated by performing motion compensation prediction on the encoding target block using the motion vector of the selected block to generate a predicted image, obtaining a difference between an input image of the encoding target block and the predicted image, converting the difference to obtain a transform coefficient, and encoding the transform coefficient. Transmitting device.

2. The image encoding device according to claim 1 , wherein the encoded data further includes a flag indicating whether or not to use a motion vector of an available block in motion compensation prediction.

Citation Information

Patent Citations

  • Motion vector calculating method

    JP2004208259A

  • Method of deriving direct-mode motion vector

    JP2006191652A

  • Encoder and decoder

    JP2008211697A