Depth Map Encoding Method, Device and Readable Medium Based on 3D-HEVC Depth Map Mode Prediction
The convolutional neural network-based DMM mode prediction model in 3D-HEVC reduces computational redundancy and time by predicting the need for DMM mode calculations, enhancing coding efficiency and quality in depth map encoding.
Patent Information
- Application Number
- CN202310449794.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-04-24
AI Technical Summary
The existing 3D-HEVC standard for depth map coding in 3D video encoding, particularly the DMM mode, results in significant computational redundancy due to excessive full-rate distortion calculations, despite offering limited improvement in coding quality.
A convolutional neural network-based DMM mode prediction model is trained to predict whether the DMM mode should be included in full-rate distortion calculations, reducing unnecessary computations by determining the need for DMM mode inclusion in advance.
This approach significantly reduces the time required for depth map coding while maintaining coding quality by avoiding redundant DMM mode calculations, leveraging shallow convolutional networks to extract features efficiently.
Smart Images

Figure CN116405683B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding, and in particular to a depth map coding method, device and readable medium based on 3D-HEVC depth map mode prediction. Background Art
[0002] In order to present a depth effect, 3D videos need to encode multiple viewpoints to simulate the stereoscopic perception of the human eye, but this greatly increases the coding complexity. To facilitate the transmission and storage of 3D videos, the 3D video extension standard 3D-HEVC based on the High Efficiency Video Coding (HEVC) was proposed. For the newly adopted MVD coding format, this standard proposed many new coding technologies, which can excellently complete the coding of stereoscopic videos, but also brings a large amount of computation. How to accelerate the coding while ensuring the coding quality is an urgent problem to be solved.
[0003] DMM (Depth Modeling Mode) is a newly added mode for depth map coding in 3D-HEVC. Since the depth map mainly records the depth information of objects, which is mainly reflected at the edges of the depth map, that is, high-frequency information, HEVC coding will compress the high-frequency information to improve the video compression rate, which is not desirable for depth maps, otherwise it will seriously affect the quality of the final synthesized viewpoints. Therefore, the proposed DMM mode can protect the edges of the depth map, and in order to ensure the coding quality, the full rate-distortion calculation list of all PUs is forced to be added. However, experimental statistics show that the proportion of the DMM mode being the best mode in the end is less than 5%, but it takes 27.72% of the depth map coding time, resulting in a great deal of computational redundancy. Summary of the Invention
[0004] Aiming at the above-mentioned technical problems, the purpose of the embodiments of the present application is to propose a depth map coding method, device and readable medium based on 3D-HEVC depth map mode prediction, which can pre-judge whether to add the DMM mode to the full rate-distortion cost calculation list during the coding process to solve the technical problems mentioned in the above background art part.
[0005] In a first aspect, the present invention provides a depth map coding method based on 3D-HEVC depth map mode prediction, which is characterized by including the following steps:
[0006] S1, obtain training data, construct a DMM mode prediction model based on a convolutional network, train the DMM mode prediction model with the training data to obtain a trained DMM mode prediction model, where the input of the trained DMM mode prediction model is a block to be encoded, and the output is a label value indicating whether to add the DMM mode to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded;
[0007] S2. Obtain the depth map sequence to be encoded, divide the depth map sequence to be encoded to obtain several current blocks to be encoded at the first-level size, input the current blocks to be encoded into the trained DMM mode prediction model, and the output network prediction value is the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list corresponding to the corresponding size during the encoding of the current block to be encoded;
[0008] S3. Use a 3D-HEVC encoder to encode the current block to be encoded, call the network prediction value during the encoding process, and determine the best mode of the current block to be encoded at the corresponding size;
[0009] S4. Determine whether the size of the current block to be encoded is greater than the second-level size. If so, adjust the size of the current block to be encoded to shrink by one level, and repeat steps S3 - S4. Otherwise, obtain the best modes of the current blocks to be encoded at all sizes.
[0010] Preferably, the DMM mode prediction model includes a first convolutional layer, a first ReLU activation layer, a second convolutional layer, a second ReLU activation layer, a first pooling layer, a third convolutional layer, a third ReLU activation layer, a second pooling layer, a fourth convolutional layer, a fourth ReLU activation layer, a third pooling layer, and a fully connected layer connected in sequence. The convolutional kernel size of the first convolutional layer is 5×5, the stride is 1, and the padding is 2. The convolutional kernel sizes of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer are 3×3, the stride is 1, and the padding is 1. The pooling kernel sizes of the first pooling layer, the second pooling layer, and the third pooling layer are 2×2, and the stride is 1.
[0011] Preferably, obtaining the training data in step S1 specifically includes:
[0012] Obtain the depth map sequence, divide the depth map sequence into several blocks to be encoded according to different size levels, input the depth map sequence into a 3D-HEVC encoder to encode each block to be encoded, determine the label value according to whether the DMM mode is added to the full rate-distortion cost calculation list corresponding to the corresponding size during the encoding process, and use the block to be encoded and its corresponding label value as the training data.
[0013] Preferably, the label value corresponding to adding the DMM mode to the full rate-distortion cost calculation list corresponding to the corresponding size during the encoding of the block to be encoded is 1, and the label value corresponding to not adding the DMM mode to the full rate-distortion cost calculation list corresponding to the corresponding size during the encoding of the block to be encoded is 0.
[0014] Preferably, the size of the first level is 64×64, and the DMM mode is not added to the full rate distortion cost calculation list of the first level size by default. The size of the second level is 4×4. There are also three sizes of 32×32, 16×16, and 8×8 with gradually decreasing levels between the first level size and the second level size.
[0015] Preferably, the cross-entropy loss function is adopted during the training process of the DMM mode prediction model. Assuming that the total number of training data is M, m represents a single sample in the training data, and the loss value L of a single sample m is calculated as follows:
[0016]
[0017] where is the true label value of the sample at the 32×32, 16×16, 8×8, and 4×4 sizes, represents the network prediction value obtained after the sample passes through the DMM mode prediction model, and C(y m , y ′m ) represents the cross-entropy between the true label value and the network prediction value. l i , l j , l k , l h correspond to one of the blocks to be encoded in the 32×32, 16×16, 8×8, and 4×4 sizes respectively. Among them, I, J, K, and H are the quantity ranges of the blocks to be encoded in the 32×32, 16×16, 8×8, and 4×4 sizes respectively. i represents the i-th element within the range of I, j represents the j-th element within the range of J, k represents the k-th element within the range of K, and h represents the h-th element within the range of H. The total loss value of all M samples is represented by L, and the calculation formula is as follows:
[0018]
[0019] The relationship between the network prediction value and the DMM mode selection is as follows:
[0020]
[0021]
[0022]
[0023]
[0024] Preferably, step S3 specifically includes:
[0025] The 3D-HEVC encoder is used to encode the current block to be encoded, and during the encoding process, the tag value indicating whether the DMM mode needs to be added to the rate-distortion cost calculation list of the corresponding size is called during the encoding process of the current block to be encoded;
[0026] Judge whether the tag value is 1. If so, add the DMM mode to the rate-distortion cost calculation list of the corresponding size, and calculate the loss value of the current block to be encoded in the DMM mode and the loss values of the remaining modes in the rate-distortion cost calculation list;
[0027] Otherwise, do not add the DMM mode to the rate-distortion cost calculation list of the corresponding size, skip the calculation of the loss value of the current block to be encoded in the DMM mode, and only calculate the loss values of the remaining modes of the current block to be encoded in the rate-distortion cost calculation list without the DMM mode;
[0028] By comparing the loss values of all modes, select the mode with the smallest loss value as the best mode of the current block to be encoded.
[0029] In a second aspect, the present invention provides a depth map encoding device based on 3D-HEVC depth map mode prediction, including:
[0030] A model construction and training module, configured to obtain training data, construct a DMM mode prediction model based on a convolutional network, train the DMM mode prediction model using the training data, and obtain a trained DMM mode prediction model. The input of the trained DMM mode prediction model is the block to be encoded, and the output is the tag value indicating whether the DMM mode needs to be added to the rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded;
[0031] A tag prediction module, configured to obtain the depth map sequence to be encoded, divide the depth map sequence to be encoded to obtain a plurality of current blocks to be encoded at the first-level size, input the current blocks to be encoded into the trained DMM mode prediction model, and the output network prediction value is the tag value indicating whether the DMM mode needs to be added to the rate-distortion cost calculation list of the corresponding size during the encoding process of the current block to be encoded;
[0032] An encoding module, configured to use the 3D-HEVC encoder to encode the current block to be encoded, call the network prediction value during the encoding process, and determine the best mode of the current block to be encoded at the corresponding size;
[0033] A repetition module, configured to judge whether the size of the current block to be encoded is greater than the second-level size. If so, adjust the size of the current block to be encoded to shrink by one level, and repeatedly execute the encoding module to the repetition module. Otherwise, obtain the best modes of the current blocks to be encoded at all sizes.
[0034] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.
[0035] In a fourth aspect, the present invention provides a computer-readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) By constructing a DMM mode prediction model based on a convolutional network, and through a 3D-HEVC encoder, according to whether to add the DMM mode to the full rate-distortion cost calculation list of the corresponding size during the encoding process to determine the label value, and associating each block to be encoded with its corresponding label value to form training data, and using the training data to train the DMM mode prediction model, a trained DMM mode prediction model is obtained. Through this trained DMM mode prediction model, it can be predicted before encoding whether the current block to be encoded needs to add the DMM mode to the full rate-distortion cost calculation list of the corresponding size, avoiding directly adding the DMM mode to the full rate-distortion cost calculation list, which leads to a redundant rate-distortion calculation process for the DMM mode.
[0038] (2) The DMM mode prediction model adopted by the present invention uses a shallow convolutional network to extract the internal features of small-sized prediction units, which can reduce the network parameter quantity while ensuring learning high-dimensional features, and reduce the time overhead of mode prediction.
[0039] (3) The present invention can significantly save the time required for depth map encoding on the premise of ensuring a certain encoding quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0041] Figure 1 is an exemplary device architecture diagram to which an embodiment of the present application can be applied;
[0042] Figure 2 is a schematic flowchart of a depth map encoding method based on 3D-HEVC depth map mode prediction according to an embodiment of the present application;
[0043] Figure 3 Schematic diagram of the structure of the DMM mode prediction model for the depth map encoding method based on 3D-HEVC depth map mode prediction according to an embodiment of the present application;
[0044] Figure 4 Schematic diagram of the encoding process of the depth map encoding method based on 3D-HEVC depth map mode prediction according to an embodiment of the present application;
[0045] Figure 5 Schematic diagram of the depth map encoding device based on 3D-HEVC depth map mode prediction according to an embodiment of the present application;
[0046] Figure 6 Schematic diagram of the structure of the computer device of the electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Figure 1 Exemplary device architecture 100 to which the depth map encoding method based on 3D-HEVC depth map mode prediction or the depth map encoding device based on 3D-HEVC depth map mode prediction according to the embodiments of the present application can be applied is shown.
[0049] As Figure 1 shown, the device architecture 100 may include terminal devices 101, 102, 103, network 104 and server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or fiber optic cables, etc.
[0050] Users may use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications may be installed on the terminal devices 101, 102, 103, such as data processing applications, file processing applications, etc.
[0051] The terminal devices 101, 102, and 103 can be either hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. They can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or can be implemented as a single software or software module. No specific limitation is made here.
[0052] The server 105 can be a server that provides various services, such as a background data processing server for processing files or data uploaded by the terminal devices 101, 102, and 103. The background data processing server can process the obtained files or data to generate a processing result.
[0053] It should be noted that the depth map encoding method based on 3D-HEVC depth map mode prediction provided by the embodiments of the present application can be executed by the server 105, or can be executed by the terminal devices 101, 102, and 103. Correspondingly, the depth map encoding device based on 3D-HEVC depth map mode prediction can be set in the server 105, or can be set in the terminal devices 101, 102, and 103.
[0054] It should be understood that Figure 1 the numbers of the terminal devices, network, and server in are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, and server. In the case where the data to be processed does not need to be obtained remotely, the above device architecture may not include a network, but only a server or a terminal device.
[0055] Figure 2 The following shows a depth map encoding method based on 3D-HEVC depth map mode prediction provided by the embodiments of the present application, including the following steps:
[0056] S1. Obtain training data, construct a DMM mode prediction model based on a convolutional network, and use the training data to train the DMM mode prediction model to obtain a trained DMM mode prediction model. The input of the trained DMM mode prediction model is the block to be encoded, and the output is the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list corresponding to the size during the encoding process of the block to be encoded.
[0057] In a specific embodiment, obtaining the training data in step S1 specifically includes:
[0058] Obtain a depth map sequence, divide the depth map sequence into several blocks to be encoded according to different size levels, input the depth map sequence into a 3D-HEVC encoder to encode each block to be encoded, determine the tag value according to whether the DMM mode is added to the full rate-distortion cost calculation list of the corresponding size during the encoding process, and use the block to be encoded and its corresponding tag value as training data.
[0059] In a specific embodiment, the tag value corresponding to adding the DMM mode to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded is 1, and the tag value corresponding to not adding the DMM mode to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded is 0.
[0060] Specifically, two depth map sequences with a resolution of 1024×768 and two depth map sequences with a resolution of 1920×1088 are selected, and the sequences include both sequences with rich content transformation and relatively gentle sequences to ensure the diversity of training data. In addition, considering the problem of limited data sets, the embodiments of the present application also perform data enhancement on the depth maps in the depth map sequence. The data enhancement methods include flipping, mirroring, and mirroring after flipping, expanding the total number of sequences to 4 times the original, making the training data more abundant, and enhancing the performance of the network.
[0061] Then, encode through a 3D-HEVC encoder to construct the training data of the blocks to be encoded of all depth map sequences and their corresponding DMM mode selections for training the DMM mode prediction model. Specifically, send all depth map sequences into the 3D-HEVC encoder, encode them with an intra-frame-only profile, and count the best mode selected for each block to be encoded. Determine whether to add the DMM mode to the full rate-distortion candidate list of the corresponding size. If so, the output tag value is 1, otherwise the output tag value is 0. Then, extract frames from all depth map sequences and cut them according to the different sizes of the blocks to be encoded. Finally, associate the obtained tag values with the corresponding blocks to be encoded, and 4 sets of tag values for whether to add the DMM mode to the full rate-distortion candidate list of the corresponding size anchored under the quantization parameters (QP=(25, 34), (30, 39), (35, 42), (40, 45)) can be obtained. Use each block to be encoded and its corresponding tag value for whether to add the DMM mode to the full rate-distortion cost calculation list of the corresponding size as training data to train the DMM mode prediction model.
[0062] In a specific embodiment, the DMM mode prediction model includes a first convolutional layer, a first ReLU activation layer, a second convolutional layer, a second ReLU activation layer, a first pooling layer, a third convolutional layer, a third ReLU activation layer, a second pooling layer, a fourth convolutional layer, a fourth ReLU activation layer, a third pooling layer, and a fully connected layer connected in sequence. The convolutional kernel size of the first convolutional layer is 5×5, the stride is 1, and the padding is 2. The convolutional kernel sizes of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer are 3×3, the stride is 1, and the padding is 1. The pooling kernel sizes of the first pooling layer, the second pooling layer, and the third pooling layer are 2×2, and the stride is 1.
[0063] Specifically, referring to Figure 3 , a DMM mode prediction model based on a convolutional network is constructed to predict whether to add the DMM mode to the full rate-distortion candidate list of the corresponding size for the block to be encoded. The block to be encoded is input into the first convolutional layer, and the output is the feature map after convolution, and the number of output channels is 6. Then, the feature map after convolution is input into the first ReLU activation layer, and then passes through three convolutional structures with the same structure, namely the second convolutional layer, the third convolutional layer, and the fourth convolutional layer, and their numbers of output channels are 12, 20, and 32 respectively. And a ReLU activation layer and a pooling layer are connected behind each convolutional structure. Finally, through the fully connected layer, the output is the label value indicating whether the block to be encoded needs to add the DMM mode to the full rate-distortion candidate list of the corresponding size. The DMM mode prediction model adopts three layers of relatively shallow convolution. An overly complex network will increase the time overhead during prediction and reduce the time-saving performance of the algorithm. Moreover, due to the characteristics of most flat regions and sharp edges in the depth map and the small amount of effective information inside, combined with the size limitation of the CU, the depth of the network structure cannot be further stacked.
[0064] In a specific embodiment, the cross-entropy loss function is used during the training process of the DMM mode prediction model. Assuming that the total number of training data is M, m represents a single sample in the training data, and the loss value L m of a single sample is calculated as follows:
[0065]
[0066] Among them, is the true label value of the sample at the sizes of 32×32, 16×16, 8×8, and 4×4, represents the network prediction value obtained after the sample passes through the DMM mode prediction model, C(y m , y ′m ) represents the cross-entropy between the true label value and the network prediction value, and l i , l j , l k , l hCorrespond to one of the to-be-encoded blocks with sizes of 32×32, 16×16, 8×8, and 4×4 respectively. Among them, I, J, K, and H are the quantity ranges of the to-be-encoded blocks with sizes of 32×32, 16×16, 8×8, and 4×4 respectively. i represents the i-th element within the range of I, j represents the j-th element within the range of J, k represents the k-th element within the range of K, and h represents the h-th element within the range of H. The total loss value of all M samples is represented by L, and the calculation formula is as follows:
[0067]
[0068] The relationship between the network prediction value and the DMM mode selection is as follows:
[0069]
[0070]
[0071]
[0072]
[0073] Specifically, the DMM mode is not defaultly added to the candidate list under the 64×64 size, so it is not considered. In one of the embodiments, I is {1 - 4}, J is {1 - 16}, K is {1 - 64}, and H is {1 - 256}. Therefore, i ∈ {1 - 4}, j ∈ {1 - 16}, k ∈ {1 - 64}, h ∈ {1 - 256}. Among them, i represents the 1st to 4th label values of the network prediction value, indicating whether the DMM mode is added to the candidate list for 4 to-be-encoded blocks with a size of 32×32. j is the 5th to 20th label of the network prediction value, indicating whether the DMM mode is added to the candidate list for 16 to-be-encoded blocks with a size of 16×16. k is the 21st to 84th label of the network prediction value, indicating whether the DMM mode is added to the candidate list for 64 to-be-encoded blocks with a size of 8×8. Finally, h is the 85th to 340th label of the network prediction value, indicating whether the DMM mode is added to the candidate list for 256 to-be-encoded blocks with a size of 4×4. The label values obtained through prediction provide guidance for subsequent encoding.
[0074] S2. Obtain the to-be-encoded depth map sequence, divide the to-be-encoded depth map sequence to obtain several current to-be-encoded blocks at the first-level size, input the current to-be-encoded blocks into the trained DMM mode prediction model, and the output network prediction value is the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list corresponding to the size during the encoding process of the current to-be-encoded blocks.
[0075] In a specific embodiment, the first-level size is 64×64, and the DMM mode is not defaultly added to the full rate-distortion cost calculation list of the first-level size. The second-level size is 4×4. Between the first-level size and the second-level size, there are also three sizes of 32×32, 16×16, and 8×8 with gradually decreasing levels.
[0076] Specifically, first divide the depth map sequence to be encoded into several current blocks to be encoded with a size of 64×64. Input the current block to be encoded into the trained DMM mode prediction model, and the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the current block to be encoded can be predicted. Predict whether the DMM mode is added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the current block to be encoded for rate-distortion calculation, and use the prediction result of the DMM mode prediction model to guide the encoding, thereby avoiding the redundant rate-distortion calculation process of the DMM mode by the 3D-HEVC encoder.
[0077] S3. Use the 3D-HEVC encoder to encode the current block to be encoded. During the encoding process, call the network prediction value and determine the best mode of the current block to be encoded under the corresponding size.
[0078] In a specific embodiment, step S3 specifically includes:
[0079] Use the 3D-HEVC encoder to encode the current block to be encoded. During the encoding process, call the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the current block to be encoded;
[0080] Judge whether the label value is 1. If so, add the DMM mode to the full rate-distortion cost calculation list of the corresponding size, and calculate the loss value of the current block to be encoded in the DMM mode and the loss values of the remaining modes in the full rate-distortion cost calculation list;
[0081] Otherwise, do not add the DMM mode to the full rate-distortion cost calculation list of the corresponding size, skip the calculation of the loss value of the current block to be encoded in the DMM mode, and only calculate the loss values of the remaining modes of the current block to be encoded in the full rate-distortion cost calculation list without the DMM mode;
[0082] By comparing the loss values of all modes, select the mode with the smallest loss value as the best mode of the current block to be encoded.
[0083] Specifically, refer to Figure 4, the prediction result of the trained DMM mode prediction model is used during the encoding process to determine whether the predicted label value is 1. If the label value is 1, it indicates that during the encoding process of the current block to be encoded, the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size. Then, during the encoding, the DMM mode is added to the full rate-distortion cost calculation list of the corresponding size, and operations such as predictive transformation are performed on this mode to calculate the loss value of the current block to be encoded under this mode. If the label value is 0, it indicates that during the encoding process of the current block to be encoded, the DMM mode does not need to be added to the full rate-distortion cost calculation list of the corresponding size. At this time, the block to be encoded before prediction is relatively smooth, the DMM mode is not added to the full rate-distortion cost calculation list, and the cost calculation for it is skipped. Only the loss values of the depth map under 35 intra-frame modes are calculated, which can effectively reduce the depth map encoding time.
[0084] S4. Determine whether the size of the current block to be encoded is greater than the second-level size. If so, adjust the size of the current block to be encoded to shrink by one level, and repeat steps S3 - S4. Otherwise, obtain the best mode of the current block to be encoded under all sizes.
[0085] Specifically, determine whether the current block to be encoded is larger than 4×4. If the current block to be encoded is larger than 4×4, adjust the size of the current block to be encoded to shrink by one level, that is, increase the encoding depth by 1, and then repeat steps S3 - S4. The encoding of the block to be encoded decreases from 32×32 to 4×4 in sequence according to the size. The smaller the size, the higher the encoding depth. If the current block to be encoded is less than or equal to 4×4, it indicates that all sizes of the current block to be encoded have been traversed, and the best mode of all blocks to be encoded has been obtained, and the mode decision process ends.
[0086] Further refer to Figure 5 , as an implementation of the methods shown in the above figures, an embodiment of a depth map encoding device based on 3D-HEVC depth map mode prediction is provided in this application. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0087] An embodiment of a depth map encoding device based on 3D-HEVC depth map mode prediction is provided in this application embodiment, including:
[0088] A model construction and training module 1, configured to obtain training data, construct a DMM mode prediction model based on a convolutional network, and use the training data to train the DMM mode prediction model to obtain a trained DMM mode prediction model. The input of the trained DMM mode prediction model is the block to be encoded, and the output is the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded;
[0089] The label prediction module 2 is configured to obtain a sequence of depth maps to be encoded, divide the sequence of depth maps to be encoded to obtain a plurality of current blocks to be encoded at the first-level size, input the current blocks to be encoded into a trained DMM mode prediction model, and the output network prediction value is the label value indicating whether the DMM mode needs to be added to the rate-distortion cost calculation list corresponding to the size during the encoding process of the current block to be encoded;
[0090] The encoding module 3 is configured to encode the current block to be encoded using a 3D-HEVC encoder, call the network prediction value during the encoding process, and determine the best mode of the current block to be encoded at the corresponding size;
[0091] The repetition module 4 is configured to determine whether the size of the current block to be encoded is greater than the second-level size. If so, adjust the size of the current block to be encoded to shrink by one level, and repeatedly execute the encoding module to the repetition module. Otherwise, obtain the best mode of the current block to be encoded at all sizes.
[0092] Reference is made below to Figure 6 which shows a schematic structural diagram of a computer device 600 suitable for use in implementing the embodiments of the present application (such as Figure 1 the server or terminal device shown). Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0093] As Figure 6 shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 603 or the program loaded from the storage section 609 into the random access memory (RAM) 604. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603, and RAM 604 are connected to each other through a bus 605. The input / output (I / O) interface 606 is also connected to the bus 605.
[0094] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 may also be connected to the I / O interface 606 as needed. A removable medium 612 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 611 as needed so that a computer program read therefrom is installed into the storage section 609 as needed.
[0095] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 610, and / or installed from the removable medium 612. When the computer program is executed by a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, the above functions defined in the methods of the present application are executed.
[0096] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution device, apparatus, or component. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0097] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0099] The modules involved in the embodiments described in the present application can be implemented in software or in hardware. The described modules can also be provided in a processor.
[0100] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain training data, construct a DMM mode prediction model based on a convolutional network, train the DMM mode prediction model using the training data to obtain a trained DMM mode prediction model, where the input of the trained DMM mode prediction model is a block to be encoded, and the output is a label value indicating whether to add the DMM mode to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded; obtain a sequence of depth maps to be encoded, divide the sequence of depth maps to be encoded to obtain a plurality of current blocks to be encoded at a first-level size, input the current blocks to be encoded into the trained DMM mode prediction model, and the output network prediction value is a label value indicating whether to add the DMM mode to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the current blocks to be encoded; encode the current blocks to be encoded using a 3D-HEVC encoder, call the network prediction value during the encoding process, and determine the optimal mode of the current blocks to be encoded at the corresponding size; determine whether the size of the current blocks to be encoded is greater than a second-level size. If so, adjust the size of the current blocks to be encoded to reduce one level, and repeat the above steps. Otherwise, obtain the optimal modes of the current blocks to be encoded at all sizes.
[0101] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.
Claims
1. A depth map encoding method based on 3D-HEVC depth map mode prediction, characterized in that It includes the following steps: S1. Obtain training data, construct a DMM mode prediction model based on a convolutional network, and train the DMM mode prediction model with the training data to obtain a trained DMM mode prediction model. The input of the trained DMM mode prediction model is a block to be encoded, and the output is a label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list corresponding to the size during the encoding process of the block to be encoded; S2. Obtain a depth map sequence to be encoded, divide the depth map sequence to be encoded into several current blocks to be encoded at the first-level size, input the current blocks to be encoded into the trained DMM mode prediction model, and the output network prediction value is the label value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list corresponding to the size during the encoding process of the current block to be encoded; S3. Encode the current block to be encoded using a 3D-HEVC encoder, call the network prediction value during the encoding process, and determine the best mode of the current block to be encoded at the corresponding size; S4. Determine whether the size of the current block to be encoded is greater than the second-level size. If so, adjust the size of the current block to be encoded to shrink by one level, and repeat steps S3 - S4. Otherwise, obtain the best mode of the current block to be encoded at all sizes.
2. The depth map encoding method based on 3D-HEVC depth map mode prediction according to claim 1, characterized in that, The DMM mode prediction model includes a first convolutional layer, a first ReLU activation layer, a second convolutional layer, a second ReLU activation layer, a first pooling layer, a third convolutional layer, a third ReLU activation layer, a second pooling layer, a fourth convolutional layer, a fourth ReLU activation layer, a third pooling layer, and a fully connected layer connected in sequence. The convolutional kernel size of the first convolutional layer is 5×5, the stride is 1, and the padding is 2. The convolutional kernel sizes of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer are 3×3, the stride is 1, and the padding is 1. The pooling kernel sizes of the first pooling layer, the second pooling layer, and the third pooling layer are 2×2, and the stride is 1.
3. The depth map encoding method based on 3D-HEVC depth map mode prediction according to claim 1, wherein In step S1, obtaining the training data specifically includes: Obtain a depth map sequence, divide the depth map sequence into several blocks to be encoded according to different size levels, input the depth map sequence into a 3D-HEVC encoder to encode each block to be encoded, determine the label value according to whether the DMM mode is added to the full rate-distortion cost calculation list corresponding to the size during the encoding process, and use the block to be encoded and its corresponding label value as training data.
4. The depth map encoding method based on 3D-HEVC depth map mode prediction according to claim 3, wherein The label value corresponding to adding the DMM mode to the full rate-distortion cost calculation list corresponding to the size during the encoding process of the block to be encoded is 1, and the label value corresponding to not adding the DMM mode to the full rate-distortion cost calculation list corresponding to the size during the encoding process of the block to be encoded is 0.
5. The depth map encoding method based on 3D-HEVC depth map mode prediction according to claim 1, characterized in that, The first-level size is 64×64. By default, the DMM mode is not added to the full rate-distortion cost calculation list of the first-level size. The second-level size is 4×4. There are also three sizes of 32×32, 16×16, and 8×8 with decreasing levels in sequence between the first-level size and the second-level size.
6. The depth map encoding method based on 3D-HEVC depth map mode prediction according to claim 5, wherein During the training process of the DMM mode prediction model, the cross-entropy loss function is adopted. Assuming that the total number of training data is M, m represents a single sample in the training data, and the loss value L m of a single sample is calculated as follows: Among them, is the true label value of the sample at the sizes of 32×32, 16×16, 8×8, and 4×4, represents the network prediction value obtained after the sample passes through the DMM mode prediction model, C(y m ,y ′m ) represents the cross-entropy between the true label value and the network prediction value, l i , l j , l k , l h correspond to one of the to-be-encoded blocks in the sizes of 32×32, 16×16, 8×8, and 4×4 respectively. Among them, I, J, K, and H are the quantity ranges of the to-be-encoded blocks in the sizes of 32×32, 16×16, 8×8, and 4×4 respectively. i represents the i-th element within the range of I, j represents the j-th element within the range of J, k represents the k-th element within the range of K, h represents the h-th element within the range of H. The total loss value of all M samples is represented by L, and the calculation formula is as follows: The relationship between the network prediction value and the DMM mode selection is as follows:
7. The depth map encoding method based on 3D-HEVC depth map mode prediction according to claim 1, characterized in that The specific steps of step S3 include: Using a 3D-HEVC encoder to encode the current block to be encoded, and calling the tag value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the current block to be encoded; Judging whether the tag value is 1. If so, adding the DMM mode to the full rate-distortion cost calculation list of the corresponding size, and calculating the loss value of the current block to be encoded in the DMM mode and the loss values of the remaining modes in the full rate-distortion cost calculation list; Otherwise, do not add the DMM mode to the full rate-distortion cost calculation list of the corresponding size, skip the calculation of the loss value of the current block to be encoded in the DMM mode, and only calculate the loss values of the remaining modes in the full rate-distortion cost calculation list without the DMM mode for the current encoding; By comparing the loss values of all modes, select the mode with the smallest loss value as the best mode of the current block to be encoded.
8. A depth map encoding device based on depth map mode prediction of 3D-HEVC, characterized in that, It includes: A model construction and training module, configured to obtain training data, construct a DMM mode prediction model based on a convolutional network, and use the training data to train the DMM mode prediction model to obtain a trained DMM mode prediction model. The input of the trained DMM mode prediction model is the block to be encoded, and the output is the tag value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the block to be encoded; A tag prediction module, configured to obtain a depth map sequence to be encoded, divide the depth map sequence to be encoded to obtain a plurality of current blocks to be encoded at the first-level size, input the current blocks to be encoded into the trained DMM mode prediction model, and the output network prediction value is the tag value indicating whether the DMM mode needs to be added to the full rate-distortion cost calculation list of the corresponding size during the encoding process of the current block to be encoded; An encoding module, configured to use a 3D-HEVC encoder to encode the current block to be encoded, call the network prediction value during the encoding process, and determine the best mode of the current block to be encoded at the corresponding size; A repetition module, configured to judge whether the size of the current block to be encoded is greater than the second-level size. If so, adjust the size of the current block to be encoded to shrink by one level, and repeatedly execute the encoding module to the repetition module. Otherwise, obtain the best modes of the current blocks to be encoded at all sizes.
9. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-7.