Video coding and decoding method and device based on tensor rearrangement and storage equipment

By rearranging and position adjustment of multi-channel tensors, the problem of inefficiency in compression of traditional video encoding and decoding methods when processing multi-channel feature tensors is solved, and compatibility between efficient compression and machine intelligence tasks is achieved.

CN120358346APending Publication Date: 2025-07-22ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082756.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-20
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing neural network-based video intelligent encoding and decoding methods cannot effectively eliminate redundant information in the time domain, resulting in the inefficient compression of traditional video encoding and decoding methods when processing multi-channel feature tensors and cannot meet the needs of machine intelligence tasks.

Method used

By rearranging the first tensor of multiple channels, adjusting the position and direction of the tensor map of each channel in the second tensor, making it a second tensor of few channels or single channels, reducing high-frequency data, efficient compression using traditional video encoding methods, and restoring the original tensor at the decoding end to maintain the function of the machine's intelligent tasks.

Benefits of technology

The compression efficiency of traditional video encoding and decoding methods is improved, the amount of compressed data caused by high-frequency information is reduced, and the ability of tensors to be used in machine intelligent tasks is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358346A_ABST
    Figure CN120358346A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding and decoding method and device based on tensor rearrangement and a storage device, and the method comprises the steps: rearranging a multi-channel first tensor into a few-channel or single-channel second tensor; and the arrangement position or direction of the tensor diagram of each channel in the first tensor in the second tensor is adjusted to reduce the high-frequency data in the second tensor, so that the second tensor can be effectively and efficiently compressed by the traditional video coding method without consuming excessive compressed data volume. According to the method, the reconstructed second tensor can be recovered to the reconstructed first tensor at the decoding end according to the rearrangement rule used by the encoding end, and the reconstructed first tensor is used for completing the machine intelligence task which can be completed by the first tensor, so that the function that the first tensor is used for the specific machine intelligence task is effectively maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video encoding and decoding, and more specifically, to a video encoding and decoding method, device, and storage device based on tensor rearrangement. Background Art

[0002] The existing encoding and decoding processing of images and videos includes traditional encoding and decoding methods and intelligent encoding and decoding methods based on neural networks. The traditional encoding and decoding methods eliminate redundant information in images and videos through operations such as prediction, transformation, quantization, entropy encoding, and loop filtering. The intelligent encoding and decoding methods use neural networks to transform images and videos, convert them into feature tensors, and perform operations such as downsampling, quantization, and entropy encoding to achieve compression encoding.

[0003] Many efficient neural network structures have been proposed for feature information extraction of images in the intelligent image encoding and decoding methods. The convolutional neural network CNN is the earliest network structure used for image encoding and decoding. Based on CNN, many improved network structures and probability estimation models have been derived. Taking network structures as an example, including network structures such as generative adversarial network GAN and recurrent neural network RNN, which have greatly improved the end-to-end image compression performance based on neural networks. Among them, the image encoding and decoding method based on generative adversarial network GAN has achieved obvious effects in improving the subjective quality of images.

[0004] The intelligent video encoding and decoding methods mainly focus on three aspects: 1) hybrid neural network encoding, 2) neural network rate distortion optimization encoding, and 3) end-to-end video encoding. Hybrid neural network encoding replaces traditional encoding modules with neural networks and embeds them into the video framework. Generally, the inter-frame prediction module, loop filtering module, and entropy encoding module are more commonly used. Neural network rate distortion optimization encoding uses the highly non-linear characteristics of neural networks to train neural networks into efficient discriminators and classifiers, such as being applied to the video encoding mode decision-making link. End-to-end video encoding generally divides into two types currently: replacing all modules of traditional encoding methods with CNN, or expanding the input dimension of neural networks to perform end-to-end compression on all frames.

[0005] In the image or video encoding and decoding methods using end-to-end neural networks, the common operation is to first extract features from the image or video, then perform encoding and decoding, and then restore it to the image. The process of extracting the feature tensor is E1, and the process of encoding the feature tensor obtained by E1 to obtain the bitstream is E2. The encoding method consists of E1 and E2. Correspondingly, the decoding method consists of D1 and D2, where D1 refers to the process of transforming the feature tensor into the decoded image, and D2 refers to the process of decoding the bitstream into the feature tensor.

[0006] After being encoded and decoded, the image and video are not only used for human viewing in current mainstream application scenarios, but also often used to complete machine intelligence tasks. The intelligent task network analyzes the image or video to complete task objectives such as object detection, object tracking, or action recognition.

[0007] Existing traditional video coding is usually designed for pixel fidelity and cannot effectively complete machine intelligence tasks. In contrast, image intelligent coding and decoding based on neural networks can be trained for the accuracy of machine intelligence tasks, thus achieving better results than traditional image coding methods in machine intelligence tasks. However, the current video intelligent coding and decoding methods based on neural networks are not good enough and cannot surpass the latest traditional video coding methods such as H.266 / VVC. This is mainly because the neural network-based methods cannot fully eliminate temporal redundancy information.

[0008] In the coding and decoding framework of video feature tensors, first, the feature tensors need to be extracted from the image and video. Then, the traditional video codec is used to encode and decode the feature tensors. Finally, the decoded feature tensors are used as the input for machine intelligence tasks for inference to complete the machine intelligence tasks. Since the feature tensors usually contain three dimensions: channels, width, and height, the number of channels in the channel dimension is much larger than the three channels of the image and video processed by traditional video coding, so the feature tensors usually cannot be directly processed by traditional video coding and decoding. Therefore, the current mainstream approach is to expand the feature tensors in the channel dimension, that is, to arrange the tensor maps on all channels in a certain order on a plane to obtain a single-channel spliced tensor, which can be directly processed by traditional video coding and decoding. The problem with this method is that the texture of the image and video processed by traditional video coding and decoding methods is usually continuous in space. In this case, the coding and decoding tools of traditional video coding and decoding methods can effectively compress the image and video. However, the spliced tensor contains many tensor maps from different channels, and there is no obvious continuity in the texture between them, resulting in high-frequency information at the adjacent positions of different tensor maps in the spliced tensor, which will reduce the compression efficiency of traditional video coding and decoding when processing the spliced tensor. Summary of the Invention

[0009] To address the above deficiencies of the prior art, the present invention proposes a video encoding and decoding method, apparatus, and storage device based on tensor rearrangement. This method rearranges a first multi-channel tensor into a second tensor with fewer channels or a single channel, and reduces the high-frequency data in the second tensor by adjusting the position or direction of the tensor maps of each channel in the first tensor when arranging them in the second tensor, enabling the second tensor to be efficiently compressed by traditional video encoding methods without consuming excessive amounts of compressed data. At the decoding end, this method can restore the reconstructed second tensor to the reconstructed first tensor according to the rearrangement rules used at the encoding end and use it to complete the machine intelligence tasks that the first tensor can perform, effectively maintaining the function of the first tensor for specific machine intelligence tasks.

[0010] To this end, the first objective of the present invention is to propose a video decoding method based on tensor rearrangement. For an input video bitstream, the following operations are performed:

[0011] Decode the first tensor data from the bitstream;

[0012] Obtain at least one second tensor data from the first tensor data, where the second tensor data is local data in the first tensor data;

[0013] Obtain the direction information of the at least one second tensor data from the bitstream;

[0014] According to the direction information, adjust the direction of the second tensor data to obtain third tensor data;

[0015] Combine at least two third tensor data in the channel dimension to obtain multi-channel fourth tensor data.

[0016] In some embodiments, the obtaining of at least one second tensor data from the first tensor data includes:

[0017] Obtain the arrangement information of the second tensor data from the bitstream, where the arrangement information includes the position information of the second tensor data in the first tensor data;

[0018] Obtain the second tensor data from the first tensor data according to the position information.

[0019] In some embodiments, the direction information is one or more of flipping up and down or flipping left and right.

[0020] In some embodiments, the combining of at least two third tensor data in the channel dimension to obtain multi-channel fourth tensor data includes:

[0021] Obtain the channel index of the third tensor data from the bitstream;

[0022] According to the channel index, combine the at least two third tensor data in channel order to obtain fourth tensor data.

[0023] In some embodiments, the method further includes:

[0024] Perform an inverse transformation on the fourth tensor data, where the inverse transformation includes feature inverse fusion or channel splitting;

[0025] Use the fourth tensor data after the inverse transformation for an intelligent task network to obtain a task result.

[0026] The second object of the present invention is to propose a video encoding method based on tensor rearrangement. For an input tensor video, perform the following operations:

[0027] Obtain first tensor data of at least one channel from at least one tensor data in the tensor video according to the channel dimension;

[0028] Adjust the direction of the first tensor data to obtain second tensor data;

[0029] Arrange and combine at least one second tensor data spatially to obtain third tensor data;

[0030] Perform encoding and compression on the third tensor data to obtain a bitstream.

[0031] In some embodiments, the method further includes:

[0032] Put the adjusted direction information of the first tensor data into the bitstream, where the direction information is one or more of flipping up and down or flipping left and right.

[0033] The third object of the present invention is to propose a video decoding device based on tensor rearrangement, including:

[0034] A processor;

[0035] A memory for storing the bitstream and decoded tensor data; and

[0036] One or more programs for completing the video decoding method based on tensor rearrangement as described in the first object of the present invention above.

[0037] The fourth object of the present invention is to propose a video encoding device based on tensor rearrangement, including:

[0038] A processor;

[0039] A memory for storing the bitstream and tensor data to be encoded; and

[0040] One or more programs for completing the video encoding method based on tensor rearrangement as described in the second object of the present invention above.

[0041] A fifth object of the present invention is to provide a storage device, which includes a bitstream obtained by using the video encoding method based on tensor rearrangement described in the second object of the present invention as above, and can be decoded by using the video decoding method based on tensor rearrangement described in the first object of the present invention as above.

[0042] The beneficial effects of the present invention are as follows: The method proposed by the present invention can reasonably arrange the positions and select the directions of tensor graphs according to the boundary texture characteristics between tensor graphs during tensor rearrangement at the encoding end, so that the textures of adjacent edges between tensor graphs tend to be continuous, effectively reducing the high-frequency information in the spliced tensor obtained by rearrangement. The spliced tensor obtained in this way can be directly compressed and decompressed efficiently by traditional video encoding and decoding methods, reducing the increase in the amount of compressed data caused by high-frequency information and effectively improving the encoding efficiency. At the decoding end, the method proposed by the present invention can perform inverse rearrangement on the spliced tensor according to the information of position arrangement and direction selection, restoring it to the dimension of the input tensor at the encoding end and maintaining the ability of the tensor to be used for subsequent machine intelligence tasks. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is a schematic flowchart of a video decoding method based on tensor rearrangement according to an embodiment of the present invention;

[0045] Figure 2 It is a schematic flowchart of another video decoding method based on tensor rearrangement according to an embodiment of the present invention. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments.

[0047] Definition of Terms:

[0048] Traditional Video Coding and Decoding: Existing video coding and decoding methods based on hybrid frameworks, such as VVC and HEVC. The full name of VVC is Versatile Video Coding, also known as H.266. The full name of HEVC is High Efficiency Video Coding. Both VVC and HEVC are video coding standards that can compress the input video to obtain a bitstream, which can contain multiple layers of sub-bitstreams, and there can be a reference relationship between the sub-bitstreams.

[0049] Tensor Rearrangement: It is an operation that rearranges the multi-channel tensors extracted from video images by a neural network to obtain a concatenated tensor. The multi-channel tensor is usually composed of a combination of two-dimensional tensor maps of multiple channels. Each channel's tensor map contains a part of the feature information in the video image, so the tensor can also be called a feature. In order to fully reflect the semantic information in the video image, the multi-channel tensor usually contains a very large number of channels, which makes it usually difficult for the tensor to be directly compressed and decompressed by existing traditional video codings. In some neural networks that complete machine tasks such as object detection, instance segmentation, or object tracking, their feature extraction networks usually also contain multi-scale branch networks that can extract multi-channel features of multiple scales from video images to fully capture objects or targets of different scales in the video image, which further increases the data volume of the features. Although existing methods use some tensor fusion methods to reduce the number of channels or the resolution of tensor data, the number of its channels is still much more than that of traditional image videos. Tensor rearrangement can arrange and concatenate multi-channel tensors into few-channel tensors suitable for traditional video coding and decoding, and effectively improve the compression efficiency of tensors processed by traditional video coding and decoding without changing the data of the tensors but only changing the expression form of the tensors.

[0050] Machine Task Network: It is a neural network that processes the input video image and completes a certain machine task. The tasks that can be completed include but are not limited to classification, object detection, instance segmentation, object tracking, or pose estimation, etc. In a video coding and decoding method based on feature tensors, the codec is inserted into a certain node in the machine task network, encodes the feature data extracted by the network before the node to obtain a bitstream, and decodes and reconstructs to obtain the feature data. The reconstructed feature data can be processed by the network after the node to complete the machine task and obtain the task result. For the convenience of description, in the present invention, the network before the node is called the front end of the machine task network, and the network after the node is called the back end of the machine task network.

[0051] This embodiment discloses a video decoding method based on tensor rearrangement. The implementation entity of this method can be a decoder. Specifically, this decoding method mainly includes the following operations: For the input bitstream, the decoder uses traditional video coding and decoding to restore the bitstream to obtain a reconstructed spliced tensor, and parses from the bitstream the arrangement information of the tensor graphs included in the reconstructed spliced tensor. This arrangement information records information such as 1) the channel index of each tensor graph in the reconstructed multi-channel tensor, 2) the width and height of each tensor graph, 3) the spatial position of each tensor graph in the reconstructed spliced tensor, and 4) the orientation of each tensor graph in the reconstructed spliced tensor. The decoder extracts each tensor graph included in the reconstructed spliced tensor according to the arrangement information, and flips or rotates it according to the orientation in which it is placed to obtain a reconstructed tensor graph. All the reconstructed tensor graphs are combined into a reconstructed multi-channel tensor according to the channel index. This reconstructed multi-channel tensor is the restoration of the feature tensor extracted by the front end of the machine task network at the encoding end, and can be processed by the back end of the machine task network to obtain a machine task result. The following describes various implementation manners of this decoding method in detail.

[0052] In one implementation manner, the decoder extracts tensor graphs from the reconstructed spliced tensor according to a fixed spatial arrangement manner. At this time, the channel index and spatial position of each tensor graph can be deduced without parsing from the bitstream.

[0053] In one implementation manner, the above fixed spatial arrangement manner can be an arrangement manner in raster scan order. Specifically, the decoder obtains tensor graphs with increasing channel indices one by one according to the width and height of the tensor graphs from the reconstructed spliced tensor in the order from top to bottom and from left to right. For example, the channel index of the tensor graph in the upper leftmost region of the reconstructed spliced tensor is 0, and the channel index of the tensor graph closest to its right side is 1. In this implementation manner, it is necessary to obtain the total number of channels of the reconstructed multi-channel tensor from the bitstream, because not all regions in the reconstructed spliced tensor contain valid tensor graphs, and there may also be some invalid padding values due to insufficiently compact arrangement. These padding values cannot be put into the reconstructed multi-channel tensor, otherwise it will affect the accuracy of the task result of the back end of the multi-channel tensor completing the machine task network. In another implementation manner, the total number of channels of the reconstructed multi-channel tensor can also be the default value set by the coding and decoding. This method is applicable to the case where the multi-channel tensor is extracted by a fixed method.

[0054] In one embodiment, the above-mentioned fixed spatial arrangement can also be an arrangement in a diagonal scanning order. Specifically, the decoder obtains tensor graphs with increasing channel indices one by one from the reconstructed spliced tensor in a diagonal order from the upper left to the lower right according to the width and height of the tensor graph. For example, the channel index of the tensor graph in the first row and first column of the reconstructed spliced tensor is 0, the channel index of the tensor graph in the second row and first column is 1, the channel index of the tensor graph in the first row and second column is 2, the channel index of the tensor graph in the first row and third column is 3, the channel index of the tensor graph in the second row and second column is 4, the channel index of the tensor graph in the third row and first column is 5, and so on.

[0055] In one embodiment, the directional information of the tensor graph includes left-right flipping and up-down flipping, and simultaneously flipping left-right and up-down flipping is equivalent to rotating 180 degrees. The decoder parses the left-right flipping information and up-down flipping information of each tensor graph from the code stream, and performs reverse flipping according to this information, that is, if the tensor graph is flipped left-right at the encoding end, the decoder must flip the tensor graph left-right again, and the same is true for up-down flipping. In one embodiment, since the purpose of whether the tensor graph is flipped or not is to maintain the texture continuity of the adjacent boundaries between adjacent tensor graphs, usually, different tensor graphs contain texture features of different attributes in the original image, but the spatial distribution of these texture features is consistent with the texture spatial distribution of the original image, that is, the textures contained in the same edges of all tensor graphs are likely to be similar. At this time, the same edges of different tensor graphs should be adjacent to each other as much as possible. In this case, the directional information of the tensor graph can be determined based on the directional information of the surrounding adjacent tensor graphs. For example, in the scanning order from the upper left to the lower right, if the tensor graph adjacent to the left or upper side of the current tensor graph is flipped left to right, then the current tensor graph is not flipped left to right, otherwise, the current tensor graph is flipped left to right; if the tensor graph adjacent to the left or upper side of the current tensor graph is flipped up and down, then the current tensor graph is not flipped up and down, otherwise, the current tensor graph is flipped up and down.

[0056] In another embodiment, the width and height of the tensor graph are not directly parsed from the bitstream, but are derived from the width and height of the reconstructed spliced tensor and the total number of channels of the reconstructed tensor. In this case, the decoder parses the bitstream to obtain the number of tensor graphs arranged horizontally or vertically in the reconstructed spliced tensor. For example, the decoder parses the bitstream to obtain the number of tensor graphs arranged horizontally in the reconstructed spliced tensor as N1, and the number of tensor graphs arranged vertically in the reconstructed spliced tensor can be obtained as N2 according to the ratio between the total number of channels and N1. Then, the width and height of the tensor graph can be obtained by dividing the width and height of the reconstructed spliced tensor by N1 and N2 respectively.

[0057] In one implementation, the specific decoding operation can be completed in the following manner:Figure 2 As shown. Without loss of generality, this embodiment takes the arrangement using the raster scan order as an example. The decoder performs the following operations:

[0058] Operation 101: Parse from the bitstream the total number of channels C of the reconstructed tensor, the width W of the tensor map, and the width xW and height xH of the reconstructed concatenated tensor;

[0059] Operation 102: Obtain the value Ax of the reconstructed concatenated tensor from the bitstream. Without loss of generality, Ax is a two-dimensional tensor with only one channel. In some other embodiments, the reconstructed concatenated tensor may also have more than one channel. In this case, it is necessary to perform derivation or parse from the bitstream the arrangement of each channel in the reconstructed concatenated tensor and the channel index of the starting tensor map;

[0060] Operation 103: According to the total number of channels C of the reconstructed tensor, the width W of the tensor map, and the width xW and height xH of the reconstructed concatenated tensor, calculate 1) the number of tensor maps arranged horizontally in the reconstructed concatenated tensor PW (for example, the value of PW is floor(xW / W)), 2) the number of tensor maps arranged vertically in the reconstructed concatenated tensor PH (for example, the value of PH is floor((C + PW - 1) / PW)), 3) the height H of the tensor map (for example, the value of H is floor(xH / H));

[0061] Operation 104: From i = 0 to i = C, with a step size of 1 for each increase in i, perform the following operations each time: 1) Calculate the position information of the tensor map of the i-th channel in the reconstructed concatenated tensor. For example, the horizontal number idW of the tensor map in the reconstructed concatenated tensor is the remainder of i divided by PW, and the vertical number idH of the tensor map in the reconstructed concatenated tensor is floor((i + PW - 1) / PW). The specific spatial position of the tensor map in the reconstructed concatenated tensor is the area from the (idH*H)-th row to the ((idH + 1)*H - 1)-th row and from the (idW*W)-th column to the ((idW + 1)*W - 1)-th column. 2) Extract the tensor map from the reconstructed concatenated tensor. 3) Flip the tensor map according to the spatial position where the tensor map is located. For example, if idH is odd, vertically flip the tensor map; if idW is odd, horizontally flip the tensor map. 4) Place the tensor map in the i-th channel of the reconstructed tensor. After the loop operation, a reconstructed tensor with C channels, width W, and height H is obtained.

[0062] In one embodiment, the information parsed from the bitstream in the above operations is derived from the syntax elements in the bitstream. Examples of the syntax elements are shown in the following table.

[0063]

[0064] Where h_packed_tensor is the height of the reconstructed concatenated tensor, w_packed_tensor is the width of the reconstructed concatenated tensor, channel_num_id_packed_tensor is the total number of channels of the reconstructed tensor, and w_tensor_map is the width of the tensor map. The descriptor u(n) represents an unsigned number of n bits. In other embodiments, signed numbers or variable-length symbols may also be used, and n may also be other reasonable values determined by the application.

[0065] In one embodiment, the syntax elements in the bitstream may further include information related to the arrangement method. Examples of the syntax elements are shown in the following table.

[0066]

[0067] Where type_pack indicates the sorting method and channel index derivation method of the tensor map in the reconstructed concatenated tensor. For example, when the value of type_pack is 1, it means that the arrangement of the tensor map uses the raster scan order arrangement method, and when the value of type_pack is 2, it means that the arrangement of the tensor map uses the diagonal scan order arrangement method.

[0068] In one embodiment, the syntax elements in the bitstream may also use a more flexible way to describe the information related to the arrangement method. Examples of the syntax elements are shown in the following table.

[0069]

[0070] Where h_in_tensormap_units and w_in_tensormap_units indicate the number in the vertical and horizontal directions in the reconstructed concatenated tensor, channel_id[i][j] represents the channel index of the j-th tensor map in the i-th horizontal direction in the vertical direction, and symmetric_turnover_type[i][j] represents the flipping type of the j-th tensor map in the i-th horizontal direction in the vertical direction. For example, when the value is 0, it means that the tensor map uses horizontal flipping, when the value is 1, it means that the tensor map uses up-and-down flipping, and when the value is 2, it means that the tensor map uses both horizontal flipping and up-and-down flipping.

[0071] In one embodiment, the obtained tensor map is only a part of the data in one channel of the reconstructed tensor. At this time, in addition to the channel index, the tensor map also needs the position information in the channel. The decoder places the tensor map at the position where the data of its channel is located according to the position information, and finally obtains the combined reconstructed tensor.

[0072] In another embodiment, the reconstructed tensor also needs to be restored to the multi-scale and multi-channel reconstructed features through the method of inverse fusion of the feature tensors. This method adopts the order from small scale to large scale. First, the neural network is used to transform the fused features into intermediate features of multiple scales respectively. Then, the intermediate features of the small scale are downsampled or pooled to obtain the downsampled features of the large scale, and they are fused with the intermediate features of the large scale to obtain the reconstructed features of the large scale. Since the intermediate features of the smallest scale cannot be fused with the features of other scales, therefore, by performing the above operations on the intermediate features except the smallest scale in turn, the reconstructed features of all scales can be obtained. In another embodiment, the feature inverse fusion method can also inversely fuse to obtain the reconstructed features of all scales in the order from large scale to small scale, which will not be elaborated here. In one embodiment, the reconstructed features of the smallest scale are further downsampled to obtain the reconstructed features of a smaller scale. In one embodiment, the reconstructed tensor can also be split in the channel dimension to obtain multiple groups of multi-channel tensors, and these multi-channel tensors are tensor data of multiple scales and can be used for the back end of the multi-scale intelligent task network.

[0073] In one embodiment, the reconstructed spliced tensor is decoded by a traditional video decoding method standardized by, for example, H.264 / AVC, H.265 / HEVC, H.266 / VVC, AVS series or AV1 series, or can also be decoded by an intelligent entropy decoding method based on super prior probability estimation.

[0074] This embodiment discloses a video coding method based on tensor rearrangement. The implementation subject of this method can be an encoder. Specifically, this coding method mainly includes the following operations: for the input image or video, a multi-channel tensor is extracted from the image in the image or video using a feature extraction network. In one embodiment, the input to the encoder can also be the already extracted multi-channel tensor data, such as a multi-channel tensor extracted from an image or multiple multi-channel tensors extracted from a video; then the encoder analyzes the tensor map of each channel in the multi-channel tensor to determine the arrangement information of the tensor map; the encoder arranges the tensor maps together according to the arrangement method to obtain a spliced tensor; the encoder compresses the spliced tensor using a traditional video coding method and puts it into the code stream, and at the same time puts the arrangement information into the code stream. This arrangement information records information such as 1) the channel index of each tensor map in the multi-channel tensor, 2) the width and height of each tensor map, 3) the spatial position of each tensor map in the spliced tensor, and 4) the direction of each tensor map in the spliced tensor. The various embodiments of this coding method will be described in detail below.

[0075] In one embodiment, the encoder arranges the tensor maps according to a fixed spatial arrangement to obtain a concatenated tensor. At this time, the channel index and spatial position of each tensor map can be deduced at the decoder end, and the encoder does not need to put this information into the bitstream.

[0076] In one embodiment, the above-mentioned fixed spatial arrangement may be an arrangement in raster scan order. Specifically, the encoder arranges the tensor maps into the concatenated tensor in the order from top to bottom and from left to right with the channel index increasing. For example, the channel index of the tensor map in the upper leftmost region of the concatenated tensor is 0, and the channel index of the tensor map closest to its right side is 1. In this embodiment, the encoder needs to put the total number of channels of the multi-channel tensor into the bitstream because not all regions in the concatenated tensor contain valid tensor maps, and there may also be some invalid padding values due to insufficiently compact arrangement. These padding values cannot be put into the multi-channel tensor reconstructed at the decoder end, otherwise it will affect the task result accuracy of the multi-channel tensor in the back end of the machine task network. In another embodiment, the total number of channels of the multi-channel tensor can also be a default value set by the encoder and decoder. This method is applicable to the case where the multi-channel tensor is extracted in a fixed manner.

[0077] In one embodiment, the above-mentioned fixed spatial arrangement may also be an arrangement in diagonal scan order. Specifically, the encoder arranges the tensor maps into the concatenated tensor in the order from the upper left to the lower right in a diagonal direction with the channel index increasing. For example, the channel index of the tensor map in the first row and the first column of the concatenated tensor is 0, the channel index of the tensor map in the second row and the first column is 1, the channel index of the tensor map in the first row and the second column is 2, the channel index of the tensor map in the first row and the third column is 3, the channel index of the tensor map in the second row and the second column is 4, the channel index of the tensor map in the third row and the first column is 5, and so on.

[0078] In one embodiment, the orientation information of the tensor graph includes left - right flipping and up - down flipping. Performing both left - right flipping and up - down flipping is equivalent to performing a 180 - degree rotation. The encoder puts the left - right flipping information and up - down flipping information of each tensor graph in the bitstream, and flips the tensor graph according to this information. In one embodiment, since the purpose of flipping the tensor graph is to maintain the texture continuity of the adjacent boundaries between adjacent tensor graphs. Generally, different tensor graphs contain texture features of different attributes in the original image, but the spatial distribution of these texture features is the same as the texture spatial distribution of the original image, that is, the textures contained in the same side of all tensor graphs are probably similar. At this time, the same sides of different tensor graphs should be adjacent to each other as much as possible. In this case, the orientation information of the tensor graph can be determined according to the orientation information of the surrounding adjacent tensor graphs. For example, in the scanning order from top - left to bottom - right, if the tensor graph adjacent to the left or top side of the current tensor graph has been left - right flipped, then the current tensor graph is not left - right flipped; conversely, the current tensor graph is left - right flipped. If the tensor graph adjacent to the left or top side of the current tensor graph has been up - down flipped, then the current tensor graph is not up - down flipped; conversely, the current tensor graph is up - down flipped.

[0079] In another embodiment, the encoder does not arrange the tensor graphs in the order of increasing channels by default. Instead, it performs a similarity analysis of the boundaries of all tensor graphs in the multi - channel tensor and arranges the tensor graphs with similar boundaries together. This way is better than the default arrangement because the default arrangement arranges the tensor graphs with adjacent channels together, but the tensor graphs with adjacent channels do not necessarily contain similar feature information, so the difference in their boundaries may be very large.

[0080] In one embodiment, the specific encoding operation can be completed in the following way. Without loss of generality, this embodiment takes the arrangement in raster scan order as an example. The encoder performs the following operations:

[0081] Operation 101: The encoding end knows the total number of channels C, width W, and height H of the multi - channel tensor, and the number of tensor graphs PW arranged horizontally in the stitched tensor. According to this information, it calculates 1) the number of tensor graphs PH arranged vertically in the stitched tensor (for example, the value of PH is floor((C + PW - 1) / PW)), 2) the width xW and height xH of the stitched tensor. For example, the value of xW is PW * W, and the value of xH is PH * H;

[0082] Operation 102: From i = 0 to i = C, with i incremented by 1 each time, perform the following operations each time: 1) Flip the tensor graph of the i-th channel according to the spatial position where the tensor graph is located. For example, if idH is odd, vertically flip the tensor graph; if idW is odd, horizontally flip the tensor graph. 2) Calculate the position information of the tensor graph of the i-th channel in the concatenated tensor. For example, the horizontal number of the tensor graph in the concatenated tensor is the remainder of i divided by PW for the value of idW, and the vertical number of the tensor graph in the concatenated tensor is floor((i + PW - 1) / PW) for the value of idH. The specific spatial position of the tensor graph in the concatenated tensor is the area from the (idH * H)-th row to the ((idH + 1) * H - 1)-th row and from the (idW * W)-th column to the ((idW + 1) * W - 1)-th column, and place the tensor graph of the i-th channel in this area. After the loop operation, a concatenated tensor with 1 channel, width xW, and height xH is obtained;

[0083] Operation 103: Compress the concatenated tensor using a traditional video coding method and put the obtained coded data into the bitstream. Put the total number of channels C of the multi-channel tensor, width W, and the width xW and height xH of the concatenated tensor into the bitstream.

[0084] In another embodiment, the multi-channel tensor also needs to be transformed from single-scale multi-channel features to fused features through a method of feature tensor fusion. This feature fusion method uses a series of neural networks with resolution reduction, eigenvalue pooling or downsampling, and fewer convolutional kernels to transform the fused features with larger resolution and more channels into reconstructed features with smaller resolution and fewer channels.

[0085] In one embodiment, the concatenated tensor can be encoded by a traditional video coding method conforming to standards such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, AVS series, or AV1 series, or can also be encoded by an intelligent entropy coding method based on hyperprior probability estimation.

[0086] This embodiment discloses a decoding device, which includes a processor and a memory, and is used to execute the decoding method disclosed in the present invention.

[0087] This embodiment discloses an encoding device, which includes a processor and a memory, and is used to execute the encoding method disclosed in the present invention.

[0088] This embodiment discloses a storage device, which includes a bitstream. The bitstream is obtained by using the encoding method disclosed in the present invention and can be processed by the decoding method disclosed in the present invention to obtain a reconstructed tensor for an intelligent task network.

[0089] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A video decoding method based on tensor rearrangement, characterized in that For the input video bitstream, perform the following operations: Decode the first tensor data from the bitstream; Obtain at least one second tensor data from the first tensor data, where the second tensor data is local data in the first tensor data; Obtain the direction information of the at least one second tensor data from the bitstream; According to the direction information, perform direction adjustment on the second tensor data to obtain third tensor data; Combine at least two third tensor data in the channel dimension to obtain a multi-channel fourth tensor data.

2. The method according to claim 1, wherein The obtaining of at least one second tensor data from the first tensor data includes: Obtain the arrangement information of the second tensor data from the bitstream, where the arrangement information includes the position information of the second tensor data in the first tensor data; Obtain the second tensor data from the first tensor data according to the position information.

3. The method according to claim 1, wherein The direction information is one or more of flipping up and down or flipping left and right.

4. The method according to claim 1, wherein The combining of at least two third tensor data in the channel dimension to obtain a multi-channel fourth tensor data includes: Obtain the channel index of the third tensor data from the bitstream; According to the channel index, combine the at least two third tensor data in the channel order to obtain the fourth tensor data.

5. The method according to claim 1, characterized in that, It further includes: Perform an inverse transformation on the fourth tensor data, where the inverse transformation includes feature inverse fusion or channel splitting; Use the fourth tensor data after the inverse transformation for an intelligent task network to obtain a task result.

6. A video coding method based on tensor rearrangement, characterized in that, For the input tensor video, perform the following operations: Obtain at least one channel of first tensor data from at least one piece of tensor data in the tensor video according to the channel dimension; Perform direction adjustment on the first tensor data to obtain second tensor data; Arrange and combine at least one second tensor data spatially to obtain third tensor data; Perform encoding and compression on the third tensor data to obtain a bitstream.

7. The method according to claim 6, wherein It further includes: Put the direction information of the adjusted first tensor data into the bitstream, where the direction information is one or more of flipping up and down or flipping left and right.

8. A video decoding device based on tensor rearrangement, comprising: A processor; A memory for storing the bitstream and decoded tensor data; And One or more programs for completing the video decoding method based on tensor rearrangement according to any one of claims 1-5.

9. A video encoding device based on tensor rearrangement, comprising: A processor; A memory for storing the bitstream and tensor data to be encoded; And One or more programs for completing the video encoding method based on tensor rearrangement according to any one of claims 6-7.

10. A storage device, characterized in that, The storage device contains a bitstream, which is obtained by using the video encoding method based on tensor rearrangement according to claim 6 or 7, and can be decoded by using the video decoding method based on tensor rearrangement according to claims 1-5.