Method for encoding / decoding feature maps and recording medium having recorded instructions to execute same

By reordering and truncating channels in feature maps based on activation, the method optimizes image compression for machine-centric tasks, improving encoding/decoding efficiency and reducing data volume.

WO2026014898A1PCT designated stage Publication Date: 2026-01-15ELECTRONICS & TELECOMM RES INST +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/009896
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-06-19
Filing Date
2025-07-08
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Traditional image compression technologies focus on human visual perception, failing to optimize for machine-centric image consumption, leading to inefficiencies in encoding and decoding processes.

Method used

Reorder channels based on activation degree and truncate channels with low activation to increase inter-channel correlation and reduce data volume in feature map encoding/decoding, using metadata to rearrange and restore channels.

Benefits of technology

Enhances encoding/decoding efficiency and reduces data volume by leveraging inter-channel correlation and selective channel removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025009896_15012026_PF_FP_ABST
    Figure KR2025009896_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for encoding feature maps, according to the present disclosure, may comprise the steps of: reordering channels of feature maps; and generating a bitstream including encoded feature maps and metadata. Herein, the metadata may include reordering information of the channels.
Need to check novelty before this filing date? Find Prior Art

Description

Feature map encoding / decoding method and recording medium recording commands for executing the same

[0001] The present disclosure relates to a feature map encoding / decoding method and device, and a recording medium recording a command for performing the method.

[0002] Traditional image compression technologies have evolved to ensure that, when restoring compressed images, the restored image resembles the original as closely as possible, based on human visual perception. In other words, image compression technologies have evolved to minimize bit rate while simultaneously maximizing the image quality of the restored image.

[0003] For example, an encoder receives an image as input and generates a bitstream through a transformation and entropy encoding process for the input image, and a decoder receives the bitstream and restores it into an image similar to the original.

[0004] To measure the similarity between the original and restored images, either objective or subjective quality assessment metrics can be used. Objective quality assessment metrics, such as the Mean Square Error (MSE), which measures the differences in pixel values ​​between the original and restored images, are commonly used. Meanwhile, subjective quality assessment metrics involve a human evaluating the differences between the original and restored images.

[0005] Meanwhile, as machine vision performance improves, more and more machines are viewing and consuming images, rather than people. For example, in areas such as smart cities, self-driving cars, and airport surveillance cameras, machine-centric image consumption is increasing, rather than human-centric.

[0006] Accordingly, in recent years, interest in image compression technology centered on machine vision has been increasing in addition to traditional image compression centered on humans.

[0007] The present disclosure aims to increase encoding / decoding efficiency by increasing inter-channel correlation through channel reordering.

[0008] The present disclosure aims to reduce the amount of data to be encoded / decoded through channel truncation.

[0009] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.

[0010] A feature map encoding method according to the present disclosure may include a step of rearranging channels of a feature map; and a step of generating a bitstream including the encoded feature map and metadata. In this case, the metadata may include information on the rearrangement of the channels.

[0011] In the feature map encoding method according to the present disclosure, the reordering can be performed based on the activation degree of each of the channels.

[0012] In the feature map encoding method according to the present disclosure, the activation degree can be set as a scale vector value of a gain unit for each channel.

[0013] In the feature map encoding method according to the present disclosure, the activation degree can be derived based on at least one of the minimum value, maximum value, average value, standard deviation value, or variance value of the feature values ​​in each channel.

[0014] In the feature map encoding method according to the present disclosure, the reordering information includes reordering order information, the reordering order information is a one-dimensional vector array, and the i-th element of the reordering order information can indicate the original index of a channel whose index is i and is reallocated through the reordering.

[0015] In the feature map encoding method according to the present disclosure, the reordering order information includes section information and position information, the section information may indicate a section to which the original index of the reallocated channel belongs, and the position information may indicate a position of the original index within the section.

[0016] In the feature map encoding method according to the present disclosure, the reordering can be performed only for channels in the active state of the feature map.

[0017] The feature map encoding method according to the present disclosure may further include a step of performing channel truncation to remove at least one channel of the feature map. In this case, the metadata may further include channel truncation information.

[0018] In the feature map encoding method according to the present disclosure, channels with an activation degree lower than a threshold value can be removed through channel pruning.

[0019] In the feature map encoding method according to the present disclosure, the channel cutting information includes channel cutting number information, and the channel cutting number information can indicate the number of channels remaining without being removed after the channel cutting or the number of channels removed through the channel cutting.

[0020] In the feature map encoding method according to the present disclosure, the channel cutting can be performed only for channels that are in an active state of the feature map.

[0021] A feature map decoding method according to the present disclosure may include the steps of: receiving a bitstream; decoding a feature map from the bitstream; and arranging channels of the decoded feature map into their original order based on metadata included in the bitstream. In this case, the metadata may include information on rearranging the channels.

[0022] In the feature map decoding method according to the present disclosure, the reordering information includes reordering order information, the reordering order information is a one-dimensional vector array, and the i-th element of the reordering order information can indicate the original index of a channel whose index is i.

[0023] In the feature map decoding method according to the present disclosure, the reordering order information includes section information and position information, the section information may indicate a section to which the original index of the channel belongs, and the position information may indicate a position of the original index within the section.

[0024] In the feature map decoding method according to the present disclosure, the alignment can be performed only for channels in the active state of the feature map.

[0025] In the feature map decoding method according to the present disclosure, the channels of the decoded feature map are rearranged based on the activation degree, and the original index of the channel can be obtained from a channel index mapping table using the rearranged index of the channel as a key value.

[0026] In the feature map decoding method according to the present disclosure, the activation degree can be set as a scale vector value of an inverse gain unit for each channel.

[0027] The feature map decoding method according to the present disclosure may further include a step of restoring channels removed through channel truncation. In this case, the metadata may further include channel truncation information.

[0028] According to the present disclosure, a computer-readable recording medium having recorded thereon a command for performing a feature map encoding / decoding method may be provided.

[0029] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary skill in the technical field to which the present invention pertains from the description below.

[0030] According to the present disclosure, by reordering channels, inter-channel correlation can be increased, thereby increasing encoding / decoding efficiency.

[0031] According to the present disclosure, the amount of data to be encoded / decoded can be reduced through channel truncation.

[0032] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0033] Figure 1 shows an example in which multiple neural networks are divided by neural network splitting points.

[0034] Figures 2 and 3 illustrate filter kernels that output feature maps for input data.

[0035] Figure 4 is a block diagram of a feature encoding unit and a feature decoding unit.

[0036] Figures 5 and 6 illustrate the order in which channel reordering and feature map pruning are performed, respectively.

[0037] FIG. 7 is a flowchart of a feature map encoding method according to one embodiment of the present disclosure.

[0038] FIG. 8 is a flowchart of a feature map decoding method according to one embodiment of the present disclosure.

[0039] Figure 9 shows an example of how the channels of a feature map are aligned.

[0040] Figures 10 to 12 illustrate examples in which channels are rearranged according to a zigzag scan order.

[0041] Figure 13 shows an example in which only active channels are reordered.

[0042] Figure 14 shows an example in which the values ​​of rearrange_offset and rearrange_value are determined.

[0043] Figure 15 shows an example in which a feature map having the same number of channels as the original feature map is restored.

[0044] Figure 16 shows an example in which reverse rearrangement is performed.

[0045] Figure 17 shows an example of generating a channel index mapping table in a feature encoder.

[0046] Figure 18 shows an example of how a channel index mapping dictionary is created.

[0047] The present disclosure is susceptible to various modifications and embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to include all modifications, equivalents, and substitutes falling within the spirit and scope of the present disclosure. In the drawings, similar reference numerals designate the same or similar functions throughout. The shapes and sizes of elements in the drawings may be exaggerated for clarity. The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments, while different from each other, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present disclosure. Furthermore, it should be understood that the positions or arrangements of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the detailed description set forth below is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments, if properly described, is defined only by the appended claims, along with the full scope equivalents to which such claims are entitled.

[0048] While terms such as "first" and "second" may be used herein to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.

[0049] When a component of the present disclosure is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components present in between. Conversely, when a component is referred to as being "directly connected" or "directly connected" to another component, it should be understood that there are no other components present in between.

[0050] The components shown in the embodiments of the present disclosure are independently depicted to represent different characteristic functions, and do not imply that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure.

[0051] The terminology used in this disclosure is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes plural expressions unless the context clearly indicates otherwise. In this disclosure, it should be understood that terms such as "comprise" or "have" are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. In other words, the description of a specific configuration in this disclosure as "comprising" does not exclude configurations other than the specified configuration, and means that additional configurations may be included in the scope of the implementation or technical idea of ​​the present disclosure.

[0052] Some components of the present disclosure may not be essential components that perform essential functions of the present disclosure, but may be optional components merely for performance enhancement. The present disclosure may be implemented by including only components essential to implementing the essence of the present disclosure, excluding components used solely for performance enhancement. A structure that includes only essential components, excluding optional components used solely for performance enhancement, is also within the scope of the present disclosure.

[0053] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In describing the embodiments of this specification, if a detailed description of a related known configuration or function is judged to obscure the gist of this specification, the detailed description will be omitted. The same reference numerals will be used for identical components in the drawings, and duplicate descriptions of the same components will be omitted.

[0054] The use of machine tasks based on image processing using artificial neural networks (ANNs) is on the rise. For example, machine vision tasks such as object classification, object recognition, object detection, object segmentation, or object tracking are increasingly being utilized, as are image processing tasks such as super-resolution or frame interpolation.

[0055] According to the present disclosure, neural networks can be distinguished from each other by neural network splitting points.

[0056] For example, the first neural network, which is divided by a split point, performs the role of extracting a feature map from input data or restoring an image from a feature map, and the second neural network performs the role of compressing (encoding) or restoring (decoding) the extracted feature map.

[0057] Figure 1 shows an example in which multiple neural networks are divided by neural network splitting points.

[0058] In the illustrated example, the neural network for extracting features from input data (hereinafter referred to as a feature extraction unit) and the neural network for encoding the extracted features (hereinafter referred to as a feature encoding unit) are illustrated as being separated by the first neural network split point. Here, the input data may be data in the form of an image.

[0059] In addition, it was exemplified that a neural network that decodes encoded features (hereinafter referred to as a feature decoding unit) and a neural network that performs machine tasks based on the decoded features (hereinafter referred to as a task execution unit) are distinguished by the second neural network split point.

[0060] Meanwhile, the feature extraction part can be defined as the first neural network (i.e., NN part1), and the task execution part can be defined as the second neural network (i.e., NN part2).

[0061] Features extracted from the first neural network can be used as input to the second neural network.

[0062] Each neural network can be composed of multiple layers. The number of layers constituting the neural network can be an integer, n.

[0063] Each layer can operate based on at least one of a weight multiplication operation, a convolution operation, an activation function, and a pooling operation.

[0064] A neural network (or neural network model) is a computational model composed of multilayer nodes (neurons) that can learn, predict, or classify patterns from given input data.

[0065] Meanwhile, the neural network according to the present disclosure may include at least one of ResNet, Fater R-CNN, Mask R_CNN, or JDE network.

[0066] The neural network split point can be the nth layer of the neural network.

[0067] Alternatively, the neural network split point may be the output of the backbone of the neural network, where the output of the backbone may represent the feature extraction value from an arbitrary layer.

[0068] The feature extraction unit extracts features from input data and classifies the extracted features. Meanwhile, feature extraction can be implemented based on at least one of VGGNet, Inception, ResNet, or FPN.

[0069] Figures 2 and 3 illustrate filter kernels that output feature maps for input data.

[0070] Figure 2 shows an example in which a feature map in the form of a two-dimensional array is output through a convolution operation of multi-channel data, and Figure 3 shows an example in which a feature map in the form of a three-dimensional array is output through a convolution operation of multi-channel data.

[0071] As in the example illustrated in Fig. 2, the feature extraction unit may include one filter kernel with horizontal and vertical sizes h and a channel length k. The feature values ​​output when input data is input to a filter kernel in one layer of a neural network can be defined as a feature map. In the example illustrated in Fig. 2, the feature map is illustrated as a single-channel image with a size of n' x m'.

[0072] As in the example shown, the feature values ​​output from the filter kernel can be defined as a feature map, and the feature map can be expressed as a 1D array, a 2D array, or a 3D array.

[0073] A two-dimensional feature map can be expressed in terms of width and height. That is, a two-dimensional feature map can contain features (i.e., feature values) equal to the product of width and height.

[0074] A 3D feature map can be expressed in terms of width, height, and channel size. That is, a 3D feature map can include features (i.e., feature values) equal to the product of width, height, and channel size.

[0075] Alternatively, as in the example illustrated in Fig. 3, the feature extraction unit may be composed of a plurality of filter kernels (e.g., k') each having a horizontal and vertical size of h and a channel length of k. The feature values ​​output when input data is input to a plurality of filter kernels in one layer of a neural network may be defined as a feature map. In Fig. 3, the feature map is exemplified as a multi-channel (k' channel) image having a size of n' x m'.

[0076] The coefficients of the filter kernel can be referred to as neural network weights. Specifically, the neural network weights may include at least one of a weight for a gain unit and a weight for channel attention.

[0077] Neural network weights can be represented as a 1D array, a 2D array, or a 3D array.

[0078] One-dimensional neural network weights can be expressed in terms of width or height.

[0079] The weights of a two-dimensional neural network can be expressed in terms of width and height.

[0080] The 3D neural network weights can be expressed in terms of width, height, and channel size.

[0081] A neural network feature map can be defined as a multidimensional tensor that is output when input data passes through one or more multidimensional layers.

[0082] Accordingly, the neural network feature map is an intermediate representation calculated on the input data, and reflects the results of the neural network calculations from the first layer to the current layer.

[0083] Meanwhile, each layer of a neural network may include weight multiplication operations, activation functions, pooling, normalization, or nonlinear operations.

[0084] A neural network feature map can be expressed in the form of a one-dimensional vector or a multidimensional tensor. The multidimensionality can be two or three dimensions.

[0085] A neural network task represents a specific problem or goal that a neural network model is designed to accomplish. Depending on the neural network task, input data can be processed and the desired output generated.

[0086] For example, the neural network task may include at least one of image classification, object detection, semantic segmentation, and multi-object tracking.

[0087] Depending on the neural network task, the structure of the neural network model may vary. For example, the structure of the neural network model may be one of ResNet, Faster R-CNN, Mask R-CNN, or JDE-network, depending on the neural network task.

[0088] Figure 4 is a block diagram of a feature encoding unit and a feature decoding unit.

[0089] Referring to FIG. 4, the feature encoding unit may include a feature reduction unit (110), a feature transformation unit (120), and a feature internal encoding unit (130), and the feature decoding unit may include a feature internal decoding unit (210), a feature inverse transformation unit (220), and a feature restoration unit (230).

[0090] At least one of a single-scale feature map or a multi-scale feature map can be input to the feature encoding unit. At least one of a single-scale feature map or a multi-scale feature map having the same dimension as the input image can be output to the feature decoding unit.

[0091] In the feature reduction unit (110), the size of the feature map can be reduced. For example, the feature reduction unit (110) can fuse or reduce an input multi-scale feature map into a single-scale feature map of a smaller size.

[0092] The feature reduction unit (110) can perform at least one of statistical characteristic-based feature nonlinear transformation, feature transformation, or feature channel adjustment.

[0093] For example, in the feature reduction unit (110), multiple feature maps can be fused and reduced into a single feature map through feature transformation.

[0094] For example, in the feature reduction unit (110), some channels of the feature map can be deleted through feature channel adjustment.

[0095]

[0096] In the feature transformation unit (120), the reduced feature map output from the feature reduction unit (110) is transformed into an input form of the feature internal encoding unit (130). To this end, the feature transformation unit may perform at least one of feature packing, feature normalization, or feature quantization.

[0097] Feature packing may be the process of converting an input feature map into a single feature frame.

[0098] Feature normalization can be the process of converting feature values ​​to values ​​between 0 and 1.

[0099] The feature internal encoding unit (130) encodes the transformed feature map output from the feature transformation unit to generate a bitstream. For example, the feature internal encoding unit (130) can encode a feature value (or a quantized feature value) within a single feature frame into a bitstream.

[0100] In the feature internal encoding unit (130), information necessary for decoding a single feature frame can be encoded and signaled in the feature internal decoding unit (210).

[0101] For example, the information may include a syntax element indicating an intra_period.

[0102] Intra-period refers to the interval at which frames encoded / decoded solely through intra prediction (i.e., I-frames) are inserted during video compression and transmission. Since only intra-prediction is used, I-frames can be encoded / decoded without reference to other frames.

[0103] The feature internal decoding unit (210) decodes the encoded feature map. Specifically, the feature internal decoding unit (210) can decode a feature value (or a quantized feature value) based on at least one of the signaled information.

[0104] Meanwhile, feature encoding / decoding can be performed based on a neural network-based codec or a signal processing-based video codec. Examples of signal processing-based video codecs include HEVC, VVC, or AV1.

[0105] The feature inverse transform unit (220) restores the reduced feature map from the decoded feature map. To this end, the feature inverse transform unit (220) may perform at least one of feature inverse quantization, feature inverse normalization, or feature unpacking. Meanwhile, the feature inverse transform unit (220) may operate in the opposite order to the feature transformation unit (120).

[0106] The feature inverse transformation unit (220) can restore the quantized feature value to the feature value before quantization or a value similar thereto through feature inverse quantization.

[0107] The feature inverse transformation unit (220) can restore feature values ​​existing within the range of 0 to 1 to feature values ​​before normalization or similar values ​​through feature inverse normalization.

[0108] The feature inverse transformation unit (220) can restore a single feature frame to a feature map or similar form before feature packing through feature unpacking.

[0109] The feature restoration unit (230) restores a single-scale feature map or a multi-scale feature map based on the restored reduced feature map output from the feature inverse transformation unit (220). Specifically, the feature restoration unit (230) may perform at least one of feature channel inverse adjustment, feature inverse transformation, and feature inverse nonlinear transformation. Meanwhile, the feature restoration unit (230) may operate in the opposite order to the feature reduction unit (110).

[0110] The feature restoration unit (230) can restore the channel-adjusted feature to the feature before adjustment or similar thereto through feature channel inverse adjustment.

[0111] The feature restoration unit (230) can restore the transformed feature to a feature before transformation or similar thereto through feature channel inverse transformation.

[0112] The feature restoration unit (230) can restore the nonlinearly transformed feature to the feature value before the nonlinear transformation or similar thereto through a specific inverse nonlinear transformation.

[0113] For decoding a feature map, at least one of information about the horizontal size of the feature map, information about the vertical size of the feature map, and information about the channel length of the feature map may be encoded and signaled.

[0114] In this disclosure, a method is proposed for sorting channels of a feature map in order of importance during feature map encoding / decoding. To this end, the feature map encoder may further include a channel reordering unit that performs channel reordering, and the feature map decoder may further include a channel de-reordering unit that restores the reordered channels to their original order.

[0115] In addition, the present disclosure proposes a method for removing channels with low importance among the channels of a feature map during feature map encoding / decoding. At least one of the above-described feature map rearrangement and channel pruning is performed by the feature reduction unit (110) of the feature map encoder, and the reverse process thereof can be performed by the feature restoration unit (230) of the feature map decoder.

[0116] Figures 5 and 6 illustrate the order in which channel reordering and feature map pruning are performed, respectively.

[0117] For example, as in the example illustrated in FIG. 5, the feature map reordering and cutting process can receive the feature map output from the feature transformation process as input. Additionally, the output from the feature map reordering and cutting process can be input to the channel adjustment process.

[0118] In the reverse order, the decoder can receive the feature map output from the feature channel detuning process and perform the feature map truncation channel restoration and inverse realignment process. In addition, the output from the feature map truncation channel restoration and inverse realignment process can be input to the feature channel inverse transformation process.

[0119] Alternatively, as in the example illustrated in Fig. 6, the feature map rearrangement and pruning process may receive the feature map output from the feature channel adjustment process as input. Furthermore, the output from the feature map rearrangement and pruning process may be input to the feature channel packing step.

[0120] In the reverse order, the decoder can receive the feature map output from the feature channel refinement process (i.e., the feature channel denormalization process) and perform the feature map truncation, channel restoration, and inverse realignment processes. Furthermore, the output from the feature map truncation, channel restoration, and inverse realignment processes can be input to the feature channel deregulation process.

[0121] Meanwhile, the feature map reordering and pruning process according to the present disclosure can be performed only when the feature map is valid. Here, the validity of the feature map can be determined by channel activation.

[0122] Similarly, the feature map cutting channel restoration and cutting channel restoration processes can be performed only when the feature map is valid.

[0123] Based on the above description, the feature map encoding / decoding method according to the present disclosure will be described in detail.

[0124] FIG. 7 is a flowchart of a feature map encoding method according to an embodiment of the present disclosure, and FIG. 8 is a flowchart of a feature map decoding method according to an embodiment of the present disclosure.

[0125] Referring to Fig. 7, first, in the feature encoder, feature map reordering information can be obtained (S710).

[0126] The feature map reordering information may include at least one of an activation list ch_activation_list for each channel of the feature map, a gain unit scale vector gain_unit_scale_vector, or a channel attention vector channel_attention_vector, which is used to perform feature map reordering.

[0127] Activation can be obtained for each channel. Activation can be expressed as an integer or a real number. The range of real numbers can be defined as -∞ to ∞.

[0128] Activation indicates the importance of a channel. For example, a high activation indicates a high impact of the channel on task performance. Conversely, a low activation indicates a low impact of the channel on task performance.

[0129] The channel-specific activation list ch_activation_list represents a set of channel-specific activations of a feature map. In other words, the channel-specific activation list ch_activation_list can represent a set of activation values ​​equal to the number of channels.

[0130] For example, ch_activation_list[n] can represent the activation level of the nth channel.

[0131] The activation of a channel can be obtained based on at least one of a neural network weight value or a channel value.

[0132] For example, the channel-specific activation list ch_activation_list can be constructed based on at least one of the gain unit scale vector gain_unit_scale_vector or the channel attention vector channel_attention_vector.

[0133] Here, the gain unit scale vector, gain_unit_scale_vector, represents a set of scale factors used to adjust the quality of the output image in a learning-based image compression model. Furthermore, the scale factors can be utilized to adjust the output of a channel. Semantically, the scale vector can indicate the relative importance of a channel.

[0134] Additionally, the channel attention vector channel_attention_vector may represent a neural network weight vector indicating the relative importance of the channel.

[0135] For example, the gain unit scale vector gain_unit_scale_vector corresponding to a channel of a feature map can be set as the activation of the corresponding channel. That is, the activation ch_activation_list[n] for the nth channel can be set as the gain unit scale vector gain_unit_scale_vector[n] for the nth channel.

[0136] For example, the channel attention vector channel_attention_vector corresponding to a channel of a feature map can be set as the activation of the corresponding channel. That is, the activation ch_activation_list[n] for the nth channel can be set as the channel attention vector channel_attention_vector[n] for the nth channel.

[0137]

[0138] Alternatively, the activation may be derived based on at least one of the minimum value, maximum value, mean value, standard deviation value, variance value, dynamic range value, or absolute value of one of the above-listed factors of the feature values ​​within the channel.

[0139] For example, as in the following mathematical expression 1, the dynamic range of the nth channel can be set to ch_activation_list[n], and the activation of the nth channel can be set to ch_activation_list[n]. Here, the dynamic range represents the difference between the maximum and minimum values ​​within the nth channel.

[0140]

[0141] For example, the average value of the nth channel can be set to ch_activation_list[n], as in the following mathematical expression 2.

[0142]

[0143] For example, the absolute value of the average value of the nth channel can be set to ch_activation_list[n], as in the following mathematical expression 3.

[0144]

[0145] For example, the activation index ch_activation_list[n] of the nth channel can be derived by weighting the absolute value of the dynamic range and average value of the nth channel as in the following mathematical expression 4.

[0146]

[0147] In the above mathematical expression 4, it is exemplified that a weight of w1 is assigned to the dynamic range, and a weight of w2 is assigned to the absolute value of the average value.

[0148] For example, as in the following mathematical expression 5, the activation map ch_activation_list[n] of the nth channel can be derived through SSIM calculation between channels. That is, the activation map ch_activation_list[n] of the nth channel can be derived based on the similarity information for the nth channel.

[0149]

[0150] For example, the activation ch_activation_list[n] of the nth channel can be derived by calculating the cosine similarity between channels as in the following mathematical expression 6. That is, the activation ch_activation_list[n] of the nth channel can be derived based on the similarity information for the nth channel.

[0151]

[0152] In the above mathematical expressions 1 to 6, ftensor[n] can represent the nth channel of the feature map.

[0153] Information about the activation level can be encoded and signaled on a channel-by-channel basis.

[0154] For example, the channel-specific activation list ch_activation_list can be encoded and signaled.

[0155] Alternatively, each element constituting the channel-specific activation list ch_activation_list can be encoded and signaled. That is, for the i-th channel, ch_activation_list[i] can be encoded and signaled. Here, i represents a number greater than or equal to 0 and less than n, and n can represent the number of elements in the channel-specific activation list.

[0156] If the gain unit scale vector gain_unit_scale_vector is set to the activation of the channel, the gain unit scale vector gain_unit_scale_vector may be encoded and signaled.

[0157] Alternatively, if the channel attention vector channel_attention_vector is set to the activation of the channel, the channel attention vector channel_attention_vector may be encoded and signaled.

[0158] Alternatively, the characteristic information of the feature map may be encoded and signaled as information on the activation of the feature map. Here, the characteristic information may represent at least one of the mean value of the channel, the variance value, the dynamic range, the absolute value of the mean value, the weighted sum result of the dynamic range and the absolute value of the mean value, the SSIM operation result, the cosine similarity, or a combination of two or more of the above-listed elements.

[0159] Based on the activation level, the channels of the feature map can be rearranged (S720). Specifically, when a feature map ftensor is input, rearrangement can be performed on the input feature map to obtain a rearranged feature map ftensor_rearr. At this time, the channels of the feature map can be rearranged in ascending or descending order of activation level.

[0160] Figure 9 shows an example of how the channels of a feature map are aligned.

[0161] For convenience of explanation, it is assumed that the activation levels of channels are a list of activation levels for each channel, ch_activation_list.

[0162] In the example illustrated in Fig. 9, the activations for channels whose indices before reordering range from 0 to 4 are illustrated as [-0.12, 1.3, 0.5, 2.5, -0.5]. When the activations are sorted in descending order, they are [2.5, 1.3, 0.5, -0.12, -0.5]. Accordingly, the channels can be reordered in the order of [3, 1, 2, 0, 4] based on the channel indices. Following the reordering, the channel indices [0, 1, 2, 3, 4] are changed to the indices [3, 1, 2, 0, 4].

[0163] Information indicating the rearrangement order of channels can be defined as fpps_channel_rearrange_index. fpps_channel_rearrange_index can be a one-dimensional vector that displays the channel indices in the rearranged order. That is, if the number of channels (i.e., channel length) is n, fpps_channel_rearrange_index can be a one-dimensional vector (i.e., a one-dimensional array) with n elements.

[0164] Each element of fpps_channel_rearrange_index represents the index of the corresponding channel. For example, fpps_channel_rearrange_index[i] represents the original index of the i-th channel (i.e., the channel whose reassigned index is i).

[0165] Information regarding channel rearrangement may be encoded and signaled. At this time, the information regarding channel rearrangement may include at least one of a sequence level flag fsps_channel_rearrange_enable_flag indicating whether channels are rearranged, a picture level flag fpps_adaptive_rearrangement_enable_flag indicating whether channels are rearranged, and information fpps_channel_rearrange_index indicating a rearrangement order of channels.

[0166] For example, the syntax fsps_channel_rearrange_enable_flag or the syntax fpps_adaptive_rearrangement_flag being a first value (e.g., 0) indicates that no rearrangement of channels has been performed, and the syntax fsps_channel_rearrange_enable_flag or the syntax fpps_adaptive_rearrangement_flag being a second value (e.g., 1) indicates that rearrangement of channels has been performed.

[0167] Only when the syntax fpps_adaptive_rearrangement_flag is the second value, the information fpps_channel_rearrange_index indicating the rearrangement order of the channels can be encoded and signaled.

[0168] In the feature decoder, the channels can be rearranged to their original order based on the information fpps_channel_rearrange_index indicating the rearrangement order. For example, if the syntax fpps_channel_rearrange_index is [3, 1, 2, 0, 4], the index of the channel whose rearranged index is 0 may be changed to 3, and the index of the channel whose rearranged index is 3 may be changed to 0. Meanwhile, channels whose original indices and rearranged indices are the same may have their indices maintained without being changed. For example, in the example above, the index of the channel whose rearranged index is 1 may be maintained as 1, the index of the channel whose rearranged index is 2 may be maintained as 2, and the index of the channel whose rearranged index is 4 may be maintained as 4.

[0169] Considering the 2D scan order, information fpps_channel_rearrange_index indicating the rearrangement order can also be encoded / decoded. Here, the 2D scan order can be determined based on a zig-zag scan, a horizontal scan, or a vertical scan. Hereinafter, it is assumed that the 2D scan order is determined based on a zig-zag scan.

[0170] After deriving the activation for each channel, the channels can be rearranged in the 2D scan order in descending order of activation.

[0171] Figures 10 to 12 illustrate examples in which channels are rearranged according to a zigzag scan order.

[0172] Figure 10 shows the result of sorting channels in descending order of activation according to the raster scan order, and Figure 11 shows an example of rearranging channels sorted according to the raster scan order of Figure 10 according to the zig-zag scan order.

[0173] As in the example illustrated in Fig. 10, when the channels are arranged in descending order of activation according to the raster scan order, the information fpps_channel_rearrange_index indicating the initial rearrangement order of the channels can be set to [10, 11, 13, 5, 12, 3, 18, 2, 4, 19, 15, 16, 0, 6, 14, 8, 7, 1, 17, 9].

[0174] Meanwhile, when the channels are arranged in a zigzag scan order in descending order of activation, the channels can be arranged as in (a) or (b) of Fig. 11. At this time, when the indices of the channels arranged in the zigzag scan order are read according to the raster scan order, the information fpps_channel_rearrange_index indicating the rearrangement order can be modified as in the example shown in (a) or (b) of Fig. 12. Fig. 12 (a) shows an example when the channels are arranged in the order shown in (a) of Fig. 11, and Fig. 12 (b) shows an example when the channels are arranged in the order shown in (b) of Fig. 11.

[0175] As in the examples shown in FIGS. 10 to 12, the information fpps_channel_rearrange_index, which indicates the initial rearrangement order set in descending order of activation, can be further sorted by considering the zig-zag scan order.

[0176] Meanwhile, a process of converting a 3D feature map into a 2D feature map for video compression may be performed. At this time, the horizontal number of channels (i.e., the number of columns) and the vertical number of channels (i.e., the number of rows) within the feature map may be defined as padd_W and padd_H, respectively.

[0177] When following the zigzag scan order, the index of the kth channel at position (i, j) in the 2D feature map can be expressed as (padd_W * i + j), where i represents the index of the column to which the channel belongs, and j represents the index of the row to which the channel belongs.

[0178] Using this, the index of the channel can be calculated according to the zig-zag scan order, as shown in Table 1 below.

[0179] Diagonal number diag = i+ jk=0for (diag=0 ; diag <padd_w+padd_h-1;diag++){imin, imax = max(0, diag - padd_w + 1), min(padd_h, diag + 1)if (diag % 2 == 0){ / evenfor (i= imin; i<imax; i++){zigzagorder[k] = i * padd_w + (diag-i)k++}}else { / oddfor (i=imax-1; i> =imin; i--)zigzagorder[k] = i * padd_w + (diag-i)k++}}}zigzagged_feature_rearranging_index[k] = feature_rearranging_index[zigzagorder[k] ]Feature_rearranging_index = zigzag_feature_rearranging_index

[0180] The zigzag scan direction can follow one of the methods of Fig. 11 (a) or Fig. 11 (b).

[0181] For example, it is assumed that the number of horizontal channels padd_w is 3, the number of vertical channels padd_h is 3, and feature_rearranging_indx is [0, 1, 2, 3, 4, 5, 6, 7, 8]. At this time, if the channels are rearranged according to the order shown in (a) of Fig. 11, the rearranged indices zigzagged_feature_rearranging_index according to the zigzag scan order can be set to [0, 1, 3, 6, 4, 2, 5, 7, 8].

[0182] For example, it is assumed that the number of horizontal channels padd_w is 3, the number of vertical channels padd_h is 3, and feature_rearranging_indx is [0, 1, 2, 3, 4, 5, 6, 7, 8]. At this time, if the channels are rearranged according to the order shown in (b) of Fig. 11, the rearranged indices zigzagged_feature_rearranging_index according to the zigzag scan order can be set to [0, 3, 1, 2, 4, 6, 7, 5, 8].

[0183] For example, if feature_rearranging_indx is [10, 11, 13, 5, 12, 3, 18, 2, 4, 19, 15, 16, 0, 6, 14, 8, 7, 1, 17, 9] and the above channels are rearranged according to the order shown in (a) of Fig. 11, the rearranged indices zigzagged_feature_rearranging_index according to the zigzag scan order can be set to [10, 11, 3, 18, 6, 13, 12, 2, 0, 14, 5, 4, 16, 8, 17, 19, 15, 7, 1, 9].

[0184] For example, if feature_rearranging_indx is [10, 11, 13, 5, 12, 3, 18, 2, 4, 19, 15, 16, 0, 6, 14, 8, 7, 1, 17, 9] and the above channels are rearranged according to the order shown in (b) of Fig. 11, the rearranged indices zigzagged_feature_rearranging_index according to the zigzag scan order can be set to [10, 13, 5, 19, 15, 11, 12, 4, 16, 7, 3, 2, 0, 8, 1, 18, 6, 14, 17, 9].

[0185] As described above, the feature_rearranging_index can be encoded / decoded, or zigzagged_feature_rearranging can be encoded / decoded instead of the feature_rearranging_index. That is, fpps_channel_rearrange_index can represent the rearrangement order of channels before the additional sorting according to the zig-zag scan order is performed, or it can represent the rearrangement order of channels after the additional sorting according to the zig-zag scan order is performed.

[0186] As another example, reordering of channels may be performed based on at least one of Python's indexing and slicing.

[0187] Here, indexing refers to selecting elements at specific locations, specifically, selecting values ​​based on a specific dimension in a tensor. Accordingly, the syntax fpps_channel_rearrange_index provides an index for a channel dimension, allowing the feature decoder to select channel data corresponding to the syntax fpps_channel_rearrange_index.

[0188] Slicing refers to a method of selecting elements in a continuous range. Specifically, when slicing is used, a partial list can be created by specifying indices corresponding to the start and end points of the range. For example, [1:3] indicates that the second to fourth elements in the list are selected. Meanwhile, selecting all elements in the list can be indicated with [:]. If slicing is supported for two dimensions, it can be expressed as [n:m, x:y]. This indicates that for the first dimension, the (n+1)th element to the (m+1)th element are selected, and for the second dimension, the (x+1)th element to the (y+1)th element are selected. Selecting all elements in two dimensions can be expressed as [:, :].

[0189] If the 3D neural network feature map to be encoded / decoded is ftensor, the rearrangement of channels can be performed by the Python code ftensor [fpps_channel_rearrange_index, ;, ;].

[0190] Additionally, for the feature map ftensor, the rearrangement process to obtain the rearranged feature map ftensor_rearr can be defined as ftensor[fpps_channel_rearrange_index, :, :] ( = ftensor_rearr).

[0191] Reordering of channels may only apply to some of the channels.

[0192] For example, reordering can be performed only on active channels. Alternatively, information indicating the channel's activity status can be explicitly encoded and signaled. This information can indicate whether the channel is active or inactive.

[0193] Alternatively, the activity status of a channel can be determined based on its activation level. Specifically, whether a channel is active or inactive can be determined based on its activation level. For example, if a channel's activation level is less than a threshold, the channel can be determined to be inactive. Conversely, if the channel's activation level is equal to or greater than the threshold, the channel can be determined to be active.

[0194] Here, the threshold value may be predefined in the feature encoder and feature decoder.

[0195] Alternatively, at least one of the minimum, maximum, average, or median of the activation values ​​of the channels may be set as the threshold. For example, the average or scaled average of the activation values ​​of multiple channels may be set as the threshold.

[0196] Randomly generated channels, such as those generated by padding, may be excluded from reordering.

[0197] Figure 13 shows an example in which only active channels are reordered.

[0198] If there are n active channels and m padded channels, reordering is performed only for n active channels, and the padded channels can maintain their original positions (original order).

[0199] For example, in Fig. 13, it is illustrated that there are four active channels (channels 0 to 3) and one padded channel (channel 4). In this case, reordering is performed only for the four active channels, and the position (order) of the one padded channel can be maintained without change.

[0200] For example, in the example illustrated in Fig. 13, the activations of channels corresponding to channels 0 to 3 are illustrated as [-0.12, 1.3, 0.5, 2.5]. Accordingly, the channels [0, 1, 2, 3] can be rearranged into [3, 1, 2, 0] according to their activations. Before and after the rearrangement, the index of the padded channel remains unchanged and can remain at 4.

[0201] Accordingly, the information fpps_channel_rearrange_index indicating the rearrangement order can be derived as [3, 1, 2, 0, 4].

[0202] Meanwhile, the encoder can adaptively determine whether to perform encoding / decoding using the rearranged feature map. Specifically, the encoder can determine whether to rearrange the feature map based on at least one of the similarity between the feature map ftensor before rearrangement and the rearranged feature map ftensor_rearr, the neural network model, or the intra_period.

[0203] For example, whether to reorder feature maps can be adaptively determined based on the similarity between adjacent channels.

[0204] Meanwhile, based on the similarity between the feature map ftensor before reordering and the reordered feature map ftensor_rearr, the similarity between the point channels can be calculated to determine whether to reorder the feature maps. The similarity can be calculated using at least one of SSIM, cosine similarity, or PCA similarity after applying PCA.

[0205] For example, the similarity between channel A and channel B of a feature map can be denoted as SIM (A, B).

[0206] At this time, when calculating similarity based on SSIM, SIM(A, B) can be set to be the same as SSIM(A, B). SSIM(A, B) can be a value measuring the structural similarity between channels A and B. Specifically, SSIM(A, B) can be a scalar value calculated by considering luminance, contrast, and structure. A value of SSIM(A, B) close to 1 can indicate that the two channels A and B are visually (structurally) similar.

[0207] Alternatively, when calculating the similarity based on cosine similarity, SIM(A, B) can be set equal to cosine_similarity(A, B). cosine_similarity(A, B) can be a value measuring the similarity of the directional components between vectors A and vectors B. Specifically, cosine_similarity(A, B) can be derived by dividing the inner product of the two vectors A and B by the magnitude of each vector. cosine_similarity(A, B) can be in the range of -1 to 1, and the closer the value of cosine_similarity(A, B) is to 1, the more directionally similar the two channels A and B are.

[0208] Alternatively, when calculating the similarity based on cosine similarity after applying PCA, SIM(A, B) can be set equal to cosine_similarity(flatten(PCA(A)), flatten(PCA(B))). PCA(X) is an array of eigenvalues ​​obtained by applying PCA to the input value X.

[0209] Principal component analysis (PCA) is a dimensionality reduction technique that maximizes data variance preservation. Specifically, PCA can be defined as the process of reducing data dimensionality while preserving key data characteristics.

[0210] The PCA(A) operation process may include an internal smoothing process for channel A when channel A is input. Furthermore, after calculating a covariance matrix, eigenvectors for the calculated covariance matrix may be obtained. Thereafter, a predetermined number of eigenvectors with large values ​​may be projected to obtain an output value. For example, the predetermined number may be 16.

[0211] The similarity between channels can represent the average similarity between the channel for which similarity is to be calculated (i.e., the target channel) and a predetermined number of channels adjacent to the target channel. Here, the predetermined number can be 1, 2, or a natural number greater than these.

[0212] For example, the average similarity between the i-th channel ft[i] and the n channels adjacent to the i-th channel ft[i] above can be derived according to the following mathematical expression 7.

[0213]

[0214] The similarity between adjacent channels, AVG_SIM, can be derived by averaging the average similarities, AVG_CH_SIM, for multiple target channels. For example, the similarity between adjacent channels, AVG_SIM, for a feature map with a total number of channels, C, can be derived according to the following mathematical expression (8).

[0215]

[0216] The larger the value of AVG_SI, the similarity between adjacent channels, the higher the similarity between adjacent channels of the feature map.

[0217] Whether feature map reordering is performed can be determined based on at least one of the similarity between adjacent channels of the feature map ftensor before reordering is performed, the similarity between adjacent channels of the reordered feature map ftensor_rearr, or the bit overhead offset bitoverhead_offset.

[0218] Here, the bit overhead offset bitoverhead_offset may be determined based on at least one of the occurrence frequency of index information occurring when performing feature reordering and the bit quantity of the index information.

[0219] The occurrence frequency of index information can be determined for each intra period Intra_period of the original image.

[0220] For example, the bit depth of the index information can be set to output_bit_depth x C.

[0221] For example, the bit overhead offset bitoverheadoffset can be determined according to the following mathematical expression (9).

[0222]

[0223] For example, if AVG_SIM(ftensor, n) is less than AVG(ftensor_rearr, n), feature reordering may be applied. Otherwise, feature reordering may not be applied.

[0224] Alternatively, conversely, if AVG_SIM(ftensor, n) is greater than AVG(ftensor_rearr, n), feature reordering may be applied. Otherwise, feature reordering may not be applied.

[0225] For example, if AVG_SIM(ftensor, n) is less than (AVG(ftensor_rearr, n) x w), feature reordering may be applied. Otherwise, feature reordering may not be applied. Here, w is a weight used to determine whether to apply reordering, and can be a real number less than or equal to 1.

[0226] Alternatively, conversely, for example, if AVG_SIM(ftensor, n) is greater than (AVG(ftensor_rearr, n) xw), feature reordering may be applied. Otherwise, feature reordering may not be applied.

[0227] For example, if AVG_SIM(ftensor, n) is less than (AVG_SIM(ftensor_rearr, n) x w1 + bitoverhead_offset x w2), feature reordering may be applied. Otherwise, feature reordering may not be applied. Here, w1 and w2 may be weights having predetermined values. In this way, the bit overhead offset can be used to determine whether feature reordering is applied.

[0228] Alternatively, conversely, if AVG_SIM(ftensor, n) is greater than (AVG_SIM(ftensor_rearr, n) x w1 + bitoverhead_offset x w2), feature reordering may be applied. Otherwise, feature reordering may not be applied.

[0229] Alternatively, depending on the neural network model, it may be determined whether to reorder the feature maps.

[0230] The neural network model may include at least one of Faster R-CNN, Mask R-CNN, or JDE Tracker. The type of the neural network model may be specified by the neural network model identifier vision_model_parameter_set_id.

[0231] For example, when the neural network model identifier vision_model_parameter_set_id has the first value, it indicates that the neural network model is Faster R-CNN.

[0232] For example, when the neural network model identifier vision_model_parameter_set_id has the second value, it indicates that the neural network model is Mask R-CNN.

[0233] For example, when the neural network model identifier vision_model_parameter_set_id has a third value, it indicates that the neural network model is JDE Tracker.

[0234] The neural network model identifier vision_model_parameter_set_id can be explicitly encoded and signaled.

[0235] Meanwhile, whether or not to realign the feature map may be determined depending on the type of the neural network model (i.e., the value of the neural network model identifier vision_model_parameter_set_id). For example, if the neural network model is Faster R-CNN (i.e., if vision_model_parameter_set_id has the first value), feature map realignment may not be performed. Or, conversely, if the neural network model is Faster R-CNN (i.e., if vision_model_parameter_set_id has the first value), feature map realignment may be performed.

[0236] For example, if the neural network model is Mask R-CNN (i.e., vision_model_parameter_set_id has the second value), feature map reordering may not be performed. Or, conversely, if the neural network model is Mask R-CNN (i.e., vision_model_parameter_set_id has the second value), feature map reordering may be performed.

[0237] For example, if the neural network model is a JDE Tracker (i.e., vision_model_parameter_set_id has the third value), feature map reordering may not be performed. Alternatively, conversely, if the neural network model is a JDE Tracker (i.e., vision_model_parameter_set_id has the third value), feature map reordering may be performed.

[0238] Alternatively, whether to reorder feature maps can be determined based on the intra period intra_period.

[0239] The intra period can have a predetermined value. For example, the intra period can be set to 1, 16, 32, 48, or 64.

[0240] For example, if the intra period has the first value, feature map reordering may not be applied. Alternatively, if the intra period is not the first value, feature map reordering may not be applied. Here, the first value may be 1.

[0241] Alternatively, it may be determined whether to apply feature map reordering based on the results of comparing the intra-period and inter-period values.

[0242] For example, if the intra period intra_period is less than the threshold n_threshold, feature map reordering may not be applied. Specifically, if the threshold n_threshold is 16 and the intra period intra_period is less than 16, feature map reordering may not be applied.

[0243] If it is determined that feature map reordering is applied, the reordered feature map ftensor_rearr can be input to the next step.

[0244] On the other hand, if it is determined that feature map reordering is not applied, the same feature map ftensor as the input feature map can be input to the next step.

[0245] Meanwhile, the syntax fpps_adaptive_rearrangement_flag, which indicates whether or not the feature maps are rearranged, can be encoded and signaled. If the syntax fpps_adaptive_rearrangement_flag is True, it indicates that rearrangement has been performed on the feature map ftensor. In this case, information fpps_channel_rearrange_index, which indicates the rearrangement order of the channels, can be additionally encoded and signaled. If it is decided to apply the feature map rearrangement, the encoder can encode the syntax fpps_adaptive_rearrangement_flag with a value of True and input the rearranged feature map ftensor_rearr to the next step.

[0246] On the other hand, if the syntax fpps_adaptivechannel_rearrangement_flag is False, it indicates that no reordering has been performed on the feature map ftensor. In this case, the information fpps_channel_rearrange_index indicating the reordering order of the channels may not be present in the bitstream. If it is decided not to apply the feature map reordering, the encoder can encode the syntax fpps_adaptive_rearrangement_flag as False and input the unrearranged feature map ftensor to the next step.

[0247] Next, channel cutting can be performed (S730).

[0248] Channel pruning means discarding some of the multiple channels that make up the feature map.

[0249] Specifically, among n channels, only K channels can be left, and the remaining channels can be removed (truncated). Here, n can represent the number of channels of the feature map ftensor, and K can be a natural number less than or equal to n. Meanwhile, the syntax fsps_reduced_feature_channel_count, which represents the number of channels of the feature map ftensor, can be explicitly encoded and signaled.

[0250] When performing channel cutting, the number K of channels that are not cut can have a fixed value.

[0251] Alternatively, the activation level of a channel can be compared to a threshold to determine whether to remove that channel.

[0252] The threshold value may be a value predefined in the feature encoder and feature decoder. For example, channels with an activation value less than 0 may be removed.

[0253] Alternatively, a threshold value can be set based on the activation ratio. For example, if the activation ratio is determined as x, the threshold value can be set by multiplying the maximum and minimum values ​​of the activation value by the difference value and then multiplying the ratio.

[0254] Alternatively, the channels to be removed can be determined based on the cutoff ratio. For example, if the cutoff ratio is determined as x, the channels with an activation level in the lower x percent can be removed.

[0255] Alternatively, a predetermined number of channels can be selected in descending order of channel activation and the selected channels can be removed.

[0256] Information related to channel truncation may be encoded and signaled. For example, at least one of a flag fpps_channel_truncation_flag indicating whether channel truncation has been performed, information fpps_active_channel_num indicating the number of channels remaining after channel truncation, information fpps_removed_channel_count_minius1 indicating the number of channels removed by channel truncation, or information truncation_index indicating a channel truncation method may be encoded and signaled.

[0257] For example, in a feature decoder, it can be determined whether channels have been truncated based on the syntax fpps_channel_truncation_flag. For example, a first value (e.g., 0) of the syntax fpps_channel_truncation_flag indicates that channel truncation has not been performed, and a second value (e.g., 1) of the syntax fpps_channel_truncation_flag indicates that channel truncation has been performed.

[0258] If the syntax fpps_channel_truncation_flag indicates that channel truncation has been performed, the syntax fpps_active_channel_num can be used to determine the number of channels remaining after channel truncation.

[0259] For example, if the value of fpps_active_channel_num is n, it indicates that the number of channels remaining after channel reduction is n. Meanwhile, the value n indicated by fpps_active_channel_num can exist in a range greater than or equal to 0 and less than or equal to the number of channels of the feature map (i.e., fsps_reduced_feature_channel_count).

[0260] The syntax fpps_removed_channel_count_minus1 represents the number of channels removed by channel truncation minus 1. That is, the number of channels removed may be equal to the value obtained by adding 1 to the syntax fpps_removed_channel_count_minus1. The decoder can restore the number of channels equal to the value obtained by adding 1 to the syntax fpps_removed_channel_count_minus1.

[0261] The syntax fpps_removed_channel_count_minus1 can be in the range 0 to fsps_reduced_feature_channel_count -1.

[0262] Regarding the relationship between the syntax fpps_active_channel_num and the syntax fpps_removed_channel_count_minus1, first, the syntax fpps_active_channel_num can represent a difference value of the number of channels removed by channel cutting (i.e., fpps_removed_channel_count_minus1 + 1) from the total number of channels (i.e., sps_reduced_feature_channel_count).

[0263] The syntax fpps_removed_channel_count_minus1 can have a value that is 1 less than the difference between the total number of channels (i.e. fsps_reduced_feature_channel_count) and the number of channels remaining after channel cutting (i.e. fpps_active_channel_num).

[0264] Meanwhile, channels removed by channel truncation may not be encoded / decoded and may also not be signaled.

[0265] The syntax truncation_index indicates how to truncate the channel.

[0266] For example, if the syntax truncation_index is the first value, it may mean that only K channels out of n channels are left and the rest are removed.

[0267] For example, the syntax truncation_index being the second value indicates that the channels to be removed are determined based on their activation level. For example, when truncation_index is the second value, channels whose activation level is less than a threshold (e.g., 0) can be removed through channel truncation.

[0268] For example, the syntax truncation_index being the third value indicates that the channels to be removed are determined based on a ratio. For example, if the syntax truncation_index is the third value, channels with an activation rate in the bottom x percent may be removed.

[0269] As described in the following embodiments, information "filling_index" indicating a method for restoring a removed channel may also be encoded and signaled. Alternatively, depending on the channel truncation method, the method for restoring a removed channel may be adaptively determined.

[0270] Meanwhile, channel pruning can be applied only to some channels in a feature map. For example, channel pruning can be performed on active channels. That is, the channels removed by channel pruning can be active channels.

[0271] A bitstream can be generated and the generated bitstream can be signaled (S740).

[0272] The bitstream may include metadata along with the encoded feature map. The metadata may include at least one of information regarding channel reordering and information regarding channel truncation.

[0273] If the information fpps_channel_rearrange_index indicating the rearrangement order of the channels is a one-dimensional vector with n elements, each element can have a value from 0 to (n-1). Here, n can indicate the number of channels in the feature map.

[0274] For example, if the number of channels in the feature map is 320, each element of fpps_channel_rearrange_index can have one of the values ​​from 0 to 319.

[0275] Meanwhile, the minimum number of bits to represent a number n is log2(n). Accordingly, fpps_channel_rearrange_index[i] can be encoded / decoded with log2(n) bits.

[0276] Instead of directly encoding / decoding the original index fpps_channel_rearrange_index of the reallocated channel, the indexes can be divided into multiple sections, and then the information rearrange_offset indicating the section to which the original index belongs and the information rearrange_index indicating the location of the original index within the section can be encoded / decoded.

[0277] Figure 14 shows an example in which the values ​​of rearrange_offset and rearrange_value are determined.

[0278] Figure 14 (a) is an example when the number of channels is 320, and Figure 14 (b) is an example when the number of channels is 600.

[0279] rearrange_offset and rearrange_value can be one-dimensional arrays with n elements.

[0280] We assume that the reallocated indices are defined as an interval of 256 indices.

[0281] In this case, each of rearrange_offset and rearrange_value can be derived by the following mathematical expression 3.

[0282]

[0283] In Equation 10, n represents the original index of the channel whose reallocated index is i. In this case, the feature decoder can derive the original index fpps_channel_rearrange_index[i] of the i-th channel according to the following Equation 11.

[0284]

[0285] By encoding / decoding rearrange_offset and rearrange_value instead of the rearrangement order information fpps_channel_rearrange_index, the amount of encoded bits can be reduced.

[0286] Referring to the example in Fig. 14, the number of bits required to encode / decode fpps_channel_rearrange_index[i] is log2(n). On the other hand, rearrange_offset[i] can be encoded / decoded with (log2(n) - 8) bits.

[0287] For example, if log2(n) is 9, rearrange_offset can be a one-dimensional array consisting of elements represented by 1 bit.

[0288] For example, if log2(n) is 10, rearrange_offset can be a one-dimensional array consisting of elements represented by 2 bits.

[0289] Meanwhile, information about channel rearrangement may include at least one of fsps_channel_rearrange_enable_flag, fpps_adaptive_rearrangement_flag, or fpps_channel_rearrange_index.

[0290] The syntax fsps_channel_rearrange_enable_flag may be a flag indicating whether to perform feature map rearrangement technology. The syntax fsps_channel_rearrange_enable_flag is encoded / decoded through SPS (Sequence Parameter Set), and depending on the value of the syntax fsps_channel_rearrange_enable_flag, it may be determined whether to encode / decode the PPS (Picture Parameter Set) syntax fpps_adaptive_rearrangement_flag.

[0291] The syntax fpps_adaptive_rearrangement_flag may be a flag indicating whether to perform feature map rearrangement technology based on the similarity of the feature maps.

[0292] For example, if the syntax fsps_channel_rearrange_enable_flag is the first value, the syntax fpps_adaptive_rearrangement_flag may be additionally signaled. In this case, each picture referencing the sequence can refer to each fpps_adaptive_rearrangement_flag to determine whether feature map rearrangement has been applied.

[0293] On the other hand, if the syntax fsps_channel_rearrange_enable_flag is the second value, the syntax fpps_adaptive_rearrangement_flag may not be signaled. In this case, it may be determined that feature map rearrangement is not applied to pictures referencing the corresponding sequence.

[0294] Instead of encoding / decoding the syntax fsps_channel_rearrange_enable_flag and the syntax fpps_adaptive_rearrangement_flag step by step, it is also possible to skip encoding / decoding one of them.

[0295] Meanwhile, whether the rearrangement index fpps_channel_rearrange_index for feature map rearrangement is additionally signaled can be determined by at least one of the syntax fsps_channel_rearrange_enable_flag or the syntax fpps_adaptive_rearrangement_flag.

[0296] For example, if the syntax fpps_adaptive_rearrangement_flag indicates that feature map rearrangement is applied, the syntax fpps_channel_rearrange_index can be additionally signaled.

[0297] Alternatively, if the syntax fsps_channel_rearrange_enable_flag indicates that feature map rearrangement is applied, if True, the syntax fpps_channel_rearrange_index may be additionally signaled regardless of whether additional encoding / decoding of the syntax fpps_adaptive_rearrangement_flag is performed.

[0298] Alternatively, instead of the syntax fpps_channel_rearrange_index, the syntax rearrange_offset and syntax rearrange_value can be encoded / decoded. In this case, the decoder can derive the value corresponding to fpps_channel_rearrange_index based on the syntax rearrange_offset and syntax rearrange_value.

[0299] For each intra period, information related to channel rearrangement may be encoded / decoded. That is, at least one of fsps_channel_rearrange_enable_flag, fpps_adaptive_rearrangement_flag, or fpps_channel_rearrange_index may be encoded and signaled for each intra period.

[0300] Alternatively, information related to channel reordering may be encoded / decoded only if the current picture corresponds to a random access point.

[0301] For example, the syntax fcm_nal_rap indicates whether the current picture corresponds to a random access point. Pictures corresponding to random access points should be restored without relying on previously restored pictures. Conversely, pictures not corresponding to random access points may rely on previously restored pictures during encoding / decoding.

[0302] The value of the syntax fcm_nal_rap_flag being the first value (e.g., True or 1) indicates that the current picture corresponds to a random access point. In this case, information related to channel rearrangement (e.g., at least one of fsps_channel_rearrange_enable_flag, fpps_adaptive_rearrangement_flag, or fpps_channel_rearrange_index) may be encoded and signaled.

[0303] On the other hand, if the value of the syntax fcm_nal_rap_flag is the second value (e.g., False or 0), it indicates that the current picture does not correspond to a random access point. In this case, encoding / decoding of information related to channel rearrangement (e.g., at least one of fsps_channel_rearrange_enable_flag, fpps_adaptive_rearrangement_flag, or fpps_channel_rearrange_index) may be omitted.

[0304] Meanwhile, for the current picture, if encoding / decoding of information related to channel reordering is omitted, information related to channel reordering of the picture corresponding to the random access point (i.e., the picture whose value of the previous fcm_nal_rap_flag is the first value) can be applied to the current picture. That is, for the current picture, if information related to channel reordering is not encoded / decoded, information related to channel reordering of the picture whose information related to channel reordering was explicitly encoded / decoded last can be set as information related to channel reordering of the current picture.

[0305] Information about channel truncation may include at least one of fsps_channel_truncation_flag, fpps_active_channel_num, fpps_removed_channel_count_minius1, or truncation_index.

[0306] If the syntax fsps_channel_truncation_flag indicates that channel truncation has been performed, at least one of the syntax fpps_active_channel_num indicating the number of channels remaining after channel truncation or the syntax fpps_removed_channel_count_minius1 indicating the number of channels removed by channel truncation may be additionally signaled. Meanwhile, the feature map to be encoded may have as many channels as the value indicated by fpps_active_channel_num.

[0307] In the encoder, the value of the syntax fsps_channel_truncation_flag can be set depending on whether channel truncation has been performed, and in the decoder, whether to restore the truncated channel can be determined depending on the value of the syntax fsps_channel_truncation_flag.

[0308] Meanwhile, when channel pruning is performed, fpps_channel_index, which indicates the channel reordering order, may indicate the reordered order of the remaining channels after channel pruning. That is, when channel pruning is performed, fpps_channel_index may include a smaller number of elements than the number of channels in the original feature map.

[0309] Accordingly, if the syntax fsps_channel_rearrange_enable_flag indicates that feature map rearrangement has been performed and the syntax fsps_channel_truncation_flag indicates that channel truncation has been performed, the feature map rearrangement index fpps_channel_rearrange_index may be encoded and signaled only for some channels of the feature map. That is, if channel truncation has been performed, the syntax fpps_channel_rearrange_index may consist of fpps_active_channel_num elements.

[0310] Meanwhile, if channel pruning is not performed, the feature map rearrangement index fpps_channel_rearrange_index may be encoded and signaled for all channels of the feature map. That is, if channel pruning is not performed, the syntax fpps_channel_rearrange_index may consist of fsps_channel_num elements.

[0311] In a feature decoder, a bitstream can be received and the received bitstream can be parsed.

[0312] In the feature decoder, an encoded feature map received through a bitstream is decoded, and through parsing, at least one of information related to rearrangement of channels and information related to channel cutting can be obtained (S810).

[0313] Information about channel truncation may include at least one of a flag indicating whether channel truncation was performed, fpps_channel_truncation_flag, information indicating the number of channels remaining after channel truncation, fpps_active_channel_num, information indicating the number of channels removed by channel truncation, fpps_removed_channel_count_minus1, or information indicating a channel truncation method, truncation_index. Meanwhile, in the decoded feature map, there may be as many channels as the value indicated by fpps_active_channel_num.

[0314] If the syntax fpps_channel_truncation_flag indicates that channel truncation has been performed, at least one of the syntax fpps_active_channel_num, the syntax fpps_removed_channel_count_minus1, and the syntax truncation_index may be further parsed.

[0315] Information about channel rearrangement may include at least one of a sequence level flag fsps_channel_rearrange_enable_flag indicating whether channels are rearranged, a picture level flag fpps_adaptive_rearrangement_flag indicating whether channels are rearranged, and information fpps_channel_rearrange_index indicating the rearrangement order of channels.

[0316] For example, if the syntax fsps_channel_rearrange_enable_flag has a first value (e.g., 0), the syntax fpps_adaptive_rearrangement_flag may not be parsed. On the other hand, if the syntax fsps_channel_rearrange_enable_flag has a second value, the syntax fpps_adaptive_rearrangement_flag may be additionally parsed from the bitstream. Here, one of the first value and the second value may be 0 and the other may be 1.

[0317] If the syntax fpps_adaptive_rearrangement_flag indicates that the channels are rearranged, the syntax fpps_channel_rearrange_index can be parsed further.

[0318] The syntax fpps_channel_rearrange_index can consist of as many elements as the number of channels in the feature map. The rearrangement order information for each channel can be obtained according to Table 2 below.

[0319] For( i=0; i <n; i++)fpps_channel_rearrange_index[i]

[0320] Meanwhile, if channel pruning is performed on the feature map, fpps_channel_rearrange_index may be composed of as many elements as the number of channels remaining after channel pruning.

[0321] Instead of the syntax fpps_channel_rearrange_index, the syntax rearrange_offset and the syntax rearrange_value can also be encoded / decoded. If the syntax rearrange_offset and the syntax rearrange_value are decoded, the rearranged index fpps_channel_rearrange_index[i] of the i-th channel can be derived based on Equation 4.

[0322] If channel truncation is performed in the feature encoder (e.g., if the syntax fpps_channel_truncation_flag indicates that channel truncation has been performed), the feature decoder can restore the removed channels (S820). In addition, the feature decoder can obtain a restored feature map by connecting the decoded channels and the restored channels (S830).

[0323] In the feature decoder, the channel truncation method can be determined based on the syntax truncation_index.

[0324] For example, the syntax truncation_index being the first value indicates that the number of channels of the feature map to be currently encoded / decoded has been truncated to a fixed number of K.

[0325] For example, the syntax truncation_index being the second value indicates that the number of channels of the feature map currently to be encoded / decoded has been truncated based on the weight values ​​of the neural network weight information.

[0326] For example, the syntax truncation_index being the third value indicates that the number of channels of the feature map to be currently encoded / decoded has been truncated according to the ratio of weights according to the neural network weight information. For example, the syntax truncation_index being the third value may indicate that the number of channels remaining after channel truncation is performed according to the ratio of weight values ​​is K.

[0327] The syntax fpps_active_channel_num can be used to determine the number of channels remaining after channel pruning. For example, if only K channels remain after channel pruning, the value of the syntax fpps_active_channel_num can be K.

[0328] The number of truncated channels can be derived through the syntax fpps_active_channel_num. For example, the number of truncated channels can be derived by subtracting the number of channels remaining after the channel truncations indicated by the syntax fpps_active_channel_num from the number of channels in the original feature map (i.e., fsps_channel_num). Meanwhile, the information fsps_channel_num indicating the number of channels in the original feature map can also be separately encoded and signaled.

[0329] As another example, instead of the syntax fpps_active_channel_num, the number of channels removed by channel truncation can be determined based on the information fpps_removed_channel_count_minius1, which indicates the number of truncation channels.

[0330] Meanwhile, only one of the syntax fpps_active_channel_num and the syntax fpps_removed_channel_count_minus1 can be set to encode / decode.

[0331] Alternatively, you can set both syntax fpps_active_channel_num and syntax fpps_removed_channel_count_minus1 to encode / decode.

[0332] By restoring the channels removed by channel cutting and concatenating the decoded channels and the restored channels, a feature map having the same number of channels as the original feature map can be restored.

[0333] Figure 15 shows an example in which a feature map having the same number of channels as the original feature map is restored.

[0334] In the example illustrated in Fig. 15, channels removed through channel cutting are illustrated as being restored by filling in a zero vector.

[0335] That is, when the number of cut channels is K, the K channels have a size of [width size of original feature map x height size of original feature map], and the values ​​of each component can be restored as all zero vectors.

[0336] Here, the number of cut channels K can be derived based on at least one of the syntax fpps_active_channel_num or the syntax fpps_removed_channel_count_minius1.

[0337] Meanwhile, information filling_index indicating how to restore channels removed through channel truncation may be encoded and signaled.

[0338] For example, the value of the syntax filling_index being the first value indicates that the removed channels are restored by filling them with a zero vector.

[0339] For example, the value of the syntax filling_index being the second value indicates that the removed channels are restored by filling them with a predefined value.

[0340] For example, a value of the syntax filling_index of the third value indicates that the removed channels are restored based on at least one channel that was not removed. For example, the removed channels can be restored by interpolating multiple channels that were not removed. Alternatively, the removed channels can be restored by copying the channels that were not removed.

[0341] Alternatively, depending on the channel cutting method, the restoration method of the removed channel can be adaptively determined.

[0342] If the channels are rearranged in the feature encoder, the feature decoder can restore the channels to their original order (S840). Restoring the rearranged channels to their original order can be referred to as inverse rearrangement.

[0343] Specifically, the decoder can determine whether to perform channel rearrangement for the current feature map based on at least one of the syntax fsps_channel_rearrange_enable_flag or the syntax fpps_adaptive_rearrangement_flag.

[0344] For example, if the syntax fsps_channel_rearrange_enable_flag has the first value, channel rearrangement may not be performed on the current feature map.

[0345] On the other hand, if the syntax fsps_channel_rearrange_enable_flag is the second value, the syntax fpps_channel_rearrange_enalbe_flag can be additionally parsed, and depending on fpps_channel-rearrange_enable_flag, it can be determined whether to perform channel rearrangement.

[0346] For example, if the syntax fpps_adaptive_rearrangement_flag has the first value, channel reverse rearrangement may not be performed on the current feature map.

[0347] On the other hand, if the syntax fpps_adaptive_rearrangement_flag is the second value, channel reverse rearrangement can be performed on the current feature map.

[0348] Here, one of the first value and the second value can be 0 and the other can be 1.

[0349] Additionally, instead of encoding / decoding the syntax fsps_channel_rearrange_enable_flag and the syntax fpps_adaptive_rearrangement_flag step by step, it is also possible to encode / decode only one of the two syntaxes above. In this case, if the parsed syntax is the first value, channel rearrangement for the feature map is not performed, and if the parsed syntax is the second value, channel rearrangement for the feature map can be performed.

[0350] Figure 16 shows an example in which reverse rearrangement is performed.

[0351] For example, according to the example illustrated in FIG. 16, if fpps_channel_rearrange_index is [3, 1, 2, 0, 4], a channel with a reallocated index of 0 can be changed to index 3, and a channel with a reallocated index of 3 can be changed to index 0.

[0352] Meanwhile, the channels with the same index before and after the reordering, the channels with the same index as 1, 2, and 4, will have the same index even after the reverse reordering.

[0353] Meanwhile, depending on the number of dimensions of the restored feature map, reverse reordering can be performed as follows.

[0354] For example, if the restored feature map is a 3D array, inverse_truncated_feature[fpps_channel_rearrange_index, ;, ;] can be performed for inverse rearrangement. Here, inverse_truncated_feature represents the restored feature map obtained by concatenating the decoded channels and the restored channels.

[0355] If the restored feature map is a 4-dimensional array, inverse_truncated_feature[;, fpps_channel_rearrange_index, ;, ;] can be performed for reverse rearrangement.

[0356] Meanwhile, the activations of channels can be determined in each of the feature encoder and feature decoder. For example, the activation of a specific channel can be derived in the feature encoder based on the scale vector value (Gain_unit_Scale_vector) of the specific channel, and the activation of a specific channel can be derived in the feature decoder based on the inverse scale vector value (Inverse_Gain_Unit_Scale_Vector) of the specific channel.

[0357] The activation map channel_activate_map can be a binary map indicating the effectiveness of a channel within a feature map. The activation map can be in the form of a 1D array or a 2D array and can be obtained through the activation list ch_activation_list. Alternatively, the activation map channel_activate_map can be explicitly encoded and signaled.

[0358] The validity of a channel can be determined by an activation threshold active_threshold, where the threshold can be a given real number.

[0359] For example, if the ith element ch_activation_list[i] in the activation list (i.e., the activation of the ith channel) is less than the threshold active_thres, the ith channel may be an invalid channel. On the other hand, if the ith element ch_activation_list[i] in the activation list (i.e., the activation of the ith channel) is equal to or greater than the threshold active_thres, the ith channel may be a valid channel.

[0360] Invalid channels may not be encoded, i.e., features within the invalid channel may not be transmitted.

[0361] In the decoder, feature values ​​within invalid channels can be set to predefined values. Here, the predefined value can be the average value of valid channels, the average value of invalid channels, the average value of feature values ​​of valid and invalid channels, or the average value of N channels with high activation. N can be 1 or a natural number greater than 1. Here, the average value can be derived from the channel values.

[0362] You can rearrange the feature maps using at least one of the syntax fpps_channel_rearrange_index, syntax fpps_active_channel_num, syntax fpps_removed_channel_count_minus1, or syntax channel_active_map.

[0363] At this time, the syntax fpps_channel_rearrange_index may be composed of a predetermined number of elements. For example, the fpps_channel_rearrange_index may be composed of as many elements as the value indicated by fsps_reduced_feature_channel_count or the value indicated by the syntax fpps_active_channel_num.

[0364] Here, the syntax fsps_reduced_feature_channel_count may indicate the number of channels remaining after channel truncation, and the syntax fpps_active_channel_num may indicate the number of channels in an active state.

[0365] By performing channel inverse reordering, an inverse reordered feature map inverse_truncated_feature can be obtained.

[0366] For example, if the syntax fpps_channel_rearrange_index is composed of as many elements as the value indicated by the syntax fsps_reduced_feature_channel_count, the inverse_truncated_feature map inverse_truncated_feature can be obtained by performing inverse_truncated_feature[fpps_channel_rearrange_index, :, :].

[0367] Meanwhile, the feature map input for reverse reordering can be determined based on at least one of the number of channels C, the vertical size H, and the horizontal size W. For example, the number of channels C of the input feature map corresponds to fsps_reduced_featrue_channel_count, the vertical size H corresponds to fsps_reduced_feature_height, and W corresponds to fsps_reduced_feature_width.

[0368] Alternatively, if the syntax fpps_channel_rearrange_index is composed of as many elements as the value indicated by the syntax fpps_active_channel_num, the inverse rearranged feature map inverse_truncated_feature can be obtained according to Table 3.

[0369] for i=0; i <fpps_active_channel_num, i++)temp_truncated_feature[i] = inverse_truncated_feature[fpps_channel_rearrange_index[i]]

[0370] In Table 3, the temporary feature map temp_truncated_feature may have the same number of channels C, vertical length H, and horizontal length W as the inverse rearranged feature map inverse_truncated_feature.

[0371] That is, for any c, h, w, if temp_truncate_feature[c][h][w] is average(X), then c can have a value greater than or equal to 0 and less than C (i.e., 0 <= c < C), h can have a value greater than or equal to 0 and less than H (i.e., 0 <= h < H), and w can have a value greater than or equal to 0 and less than W (i.e., 0 <= w < W).

[0372] Here, average(X) can represent the average value of all feature values ​​of the input feature map X.

[0373] The input feature map X may be a feature map that has undergone a restoration process for channel cutting.

[0374] When determining the activation of channels in each of the feature encoder and the feature decoder, information fpps_channel_rearrange_index indicating the rearrangement order of channels may be derived based on at least one of the channel importance index mapping table (or channel index mapping table) and the channel importance index mapping dictionary (or channel index mapping dictionary).

[0375] At this time, the channel index mapping table channel_index_mapping_table and the channel index mapping dictionary channel_index_mapping_dictionary may be input from the outside.

[0376] For example, a channel index mapping table channel_index_mapping_table and a channel index mapping dictionary channel_index_mapping_dictionary can be obtained from a specific file.

[0377] Alternatively, a channel index mapping table channel_index_mapping_table and a channel index mapping dictionary channel_index_mapping_dictionary can be derived based on values ​​input from a specific file.

[0378] Alternatively, the channel index mapping table channel_index_mapping_table and the channel index mapping dictionary channel_index_mapping_dictionary can be received from outside.

[0379] Alternatively, a channel index mapping table channel_index_mapping_table and a channel index mapping dictionary channel_index_mapping_dictionary can be derived based on values ​​received from outside.

[0380] As another example, a feature encoder may generate a channel index mapping table, channel_index_mapping_table, and transmit it to a feature decoder. This transmission may be performed as a separate file or as part of compressed data.

[0381] In the feature decoder, a channel index mapping dictionary channel_index_mapping_dictionary can be derived based on the channel index mapping table channel_index_mapping_table.

[0382] The channel index mapping table channel_index_mapping_table and the channel index mapping dictionary channel_index_mapping_dictionary can each be represented as a two-dimensional array. Furthermore, each entry in the channel index mapping dictionary can be represented as a (key, value) pair.

[0383] The channel index mapping table channel_index_mapping_table and the channel index mapping dictionary channel_index_mapping_dictionary can represent correspondences between reordered locations according to predefined criteria. Here, the predefined criteria can be the activation levels of channels.

[0384] Below, we will describe in detail the process of restoring fpps_channel_rearrange_index based on the channel index mapping table channel_index_mapping_table and the channel index mapping dictionary channel_index_mapping_dictionary.

[0385] Figure 17 shows an example of generating a channel index mapping table in a feature encoder.

[0386] In a feature encoder, the activation of each channel can be calculated, and the channels can be rearranged based on the calculated activation. At this time, the activation can be obtained based on the inverse scale vector value (Inverse_Gain_Unit_Scale_Vector) of the channel.

[0387] In the example illustrated in Fig. 17, it is illustrated that the rearranged order [2, 1, 5, 4, 0, 6, 3] in the form of a one-dimensional vector is derived by sorting the channels in descending order of activation.

[0388] Afterwards, the original indices and reallocated indices of the channels can be mapped to create a channel index mapping table, channel_index_mapping_table. That is, the i-th element of the channel index mapping table, channel_index_mapping_table[i], can represent a pair of reallocated indices of the channel with the original index i.

[0389] Figure 18 shows an example of how a channel index mapping dictionary is created.

[0390] In a feature decoder, the activation of each channel can be calculated, and the channels can be rearranged based on the calculated activation. At this time, the activation can be obtained based on the inverse scale vector value (Inverse_Gain_Unit_Scale_Vector) of the channel.

[0391] In the example illustrated in Fig. 18, it is illustrated that the rearranged order [2, 1, 4, 5, 0, 6, 3] in the form of a one-dimensional vector is derived by sorting the channels in descending order of activation.

[0392] Afterwards, the feature decoder can derive the channel index mapping dictionary channel_index_mapping_dictionary based on the channel index mapping table channel_index_mapping_table.

[0393] The channel index mapping dictionary chnnel_index_mapping_dictionary can be composed of a pair of rearrange_index_value, which is a one-dimensional vector, and rearrange_index_key, which is a one-dimensional vector.

[0394] On the feature decoder side, when the channels are sorted in descending order of activation, the index of the rearranged channel can be set to the rearranged index key rearrange_index_key. That is, the i-th element of the rearranged index key, rearrange_index_key[i], indicates the index of the channel with the i-th activation when the channels are rearranged by activation in the feature decoder.

[0395] Meanwhile, the i-th element rearrange_index_value[i] of the rearranged index value indicates the value corresponding to the i-th element rearrange_index_key[i] of the rearranged index key in the channel index mapping table channel_mapping_table (i.e., channel_index_mapping_table[rearrange_index_key[i]]).

[0396] Accordingly, in the example illustrated in FIG. 18, the channel index mapping dictionary channel_index_mapping_dictionary can be expressed as {2:5, 1:1, 4:0, 5:6, 0:2, 6:3, 3:4}.

[0397] The channel rearrangement order fpps_channel_rearrange_index can be derived based on at least one of the channel index mapping table channel_index_mapping_table or the channel index mapping dictionary channel_index_mapping_dictionary.

[0398] For example, given a channel index mapping dictionary channel_index_mapping_dictionary, the i-th element of the channel's rearrange order, fpps_channel_rearrange_index[i], can be set to the rearrange_index_value when rearrange_index_key is i.

[0399] For example, given a channel index mapping table channel_index_mapping_table, the i-th element fpps_channel_rearrange_index[i] of the channel rearrangement order can be set to the i-th element channel_index_mapping_table[i] of the channel index mapping table.

[0400] The names of the syntax elements introduced in the above-described embodiments are merely provisional and assigned for the purpose of describing embodiments according to the present disclosure. The syntax elements may be named with names different from those proposed in the present disclosure.

[0401] The components described in the exemplary embodiments of the present disclosure may be implemented by hardware elements. For example, the hardware elements may include at least one of a digital signal processor (DSP), a processor, a controller, an application-specific integrated circuit (ASIC), a programmable logic element such as an FPGA, a graphics processing unit (GPU), other electronic devices, or a combination thereof. At least some of the functions or processes described in the exemplary embodiments of the present disclosure may be implemented in software, and the software may be recorded on a recording medium. The components, functions, and processes described in the exemplary embodiments may be implemented by a combination of hardware and software.

[0402] A method according to one embodiment of the present disclosure may be implemented as a program that can be executed by a computer, and the computer program may be recorded on various recording media such as a magnetic storage medium, an optical readable medium, a digital storage medium, etc.

[0403] The various technologies described in this disclosure may be implemented as digital electronic circuits or computer hardware, firmware, software, or a combination thereof. The technologies may be implemented as a computer program product, i.e., a computer program tangibly embodied in an information medium (e.g., a machine-readable storage device (e.g., a computer-readable medium) or a data processing device), or a computer program embodied as a signal propagated to be processed by a data processing device or to cause the operation of a data processing device (e.g., a programmable processor, a computer, or multiple computers).

[0404] The computer program(s) may be written in any programming language, including compiled or interpreted languages, and may be distributed in any form, including standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. The computer program(s) may be executed by a single computer, or by multiple computers distributed across one or more sites and interconnected by a communications network.

[0405] Examples of processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, and one or more processors of a digital computer. Typically, a processor receives instructions and data from read-only memory, random-access memory, or both. Components of a computer may include at least one processor for executing instructions, and one or more memory devices for storing instructions and data. Additionally, the computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or may be connected to such mass storage devices to receive and / or transmit data. Examples of information media suitable for implementing computer program instructions and data include semiconductor memory devices (e.g., magnetic media such as hard disks, floppy disks, and magnetic tape), optical media such as compact disc read-only memories (CD-ROMs), digital video discs (DVDs), magneto-optical media such as floptical disks, and read-only memory (ROM), random access memory (RAM), flash memory, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), and other known computer-readable media. The processor and memory may be supplemented by, or integrated with, special purpose logic circuitry.

[0406] A processor can execute an operating system (OS) and one or more software applications running on the OS. The processor device can also access, store, manipulate, process, and generate data in response to the software execution. For simplicity, the processor device is described singularly; however, those skilled in the art will understand that the processor device may include multiple processing elements and / or different types of processing elements. For example, the processor device may include multiple processors or a processor and a controller. Additionally, the processor device may configure different processing structures, such as parallel processors. Furthermore, a computer-readable medium refers to any medium that a computer can access, and may include both computer storage media and transmission media.

[0407] While this disclosure contains detailed descriptions of various detailed implementation examples, it should be understood that such details do not limit the invention or the scope of the claims proposed by this disclosure, but rather illustrate features of specific exemplary embodiments.

[0408] Features described individually in the exemplary embodiments of this disclosure may be implemented by a single exemplary embodiment. Conversely, various features described in the present disclosure with respect to a single exemplary embodiment may also be implemented by combinations or appropriate subcombinations of multiple exemplary embodiments. Furthermore, the present disclosure may disclose that the features operate in a particular combination, and while the combination may initially be described as claimed, in some cases, one or more features may be excluded from the claimed combination, or the claimed combination may be modified into a subcombination or a modification of a subcombination.

[0409] Likewise, even if operations are depicted in a particular order in the drawings, this should not be construed as requiring the execution of the operations in a specific order or sequence, or the performance of all operations, to achieve the desired result. Multitasking and parallel processing may be useful in certain cases. Furthermore, the various device components in the exemplary embodiments of the present invention should not be construed as necessarily being separate, and the program components and devices described above may be packaged into a single software product or multiple software products.

[0410] The exemplary embodiments disclosed herein are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will recognize that various modifications to the exemplary embodiments can be made without departing from the spirit and scope of the claims and their equivalents.

[0411] Accordingly, the present disclosure is intended to include all other replacements, modifications and variations that fall within the scope of the following claims.

[0412] Embodiments according to the present disclosure can be applied to an electronic device capable of encoding / decoding an image.

Claims

A step of rearranging the channels of the feature map; and Comprising the step of generating a bitstream including an encoded feature map and metadata, A feature map encoding method, characterized in that the above metadata includes reordering information of the channels. In the first paragraph, A feature map encoding method, characterized in that the above reordering is performed based on the activation degree of each of the channels. In the second paragraph, A feature map encoding method, characterized in that the above activation degree is set as a scale vector value of a gain unit for each channel. In the second paragraph, A feature map encoding method, characterized in that the above activation level is derived based on at least one of the minimum value, maximum value, average value, standard deviation value, or variance value of feature values ​​within each channel. In the first paragraph, The above reordering information includes reordering order information, The above reordering order information is a one-dimensional vector array, A feature map encoding method, characterized in that the i-th element of the above reordering order information indicates the original index of a channel whose index is i that has been reassigned through the reordering. In the first paragraph, The above reordering order information includes section information and location information, The above section information indicates the section to which the original index of the reallocated channel belongs, A feature map encoding method, characterized in that the above location information indicates the location of the original index within the above section. In the first paragraph, A feature map encoding method, characterized in that the reordering is performed only for channels that are active in the feature map. In the first paragraph, The above feature map encoding method further includes a step of performing channel cutting to remove at least one channel of the feature map, A feature map encoding method, characterized in that the above metadata further includes channel cutting information. In paragraph 8, A feature map encoding method, characterized in that channels having an activation level less than a threshold value are removed through channel pruning. In paragraph 8, The above channel cutting information includes channel cutting number information, A feature map encoding method, characterized in that the above channel cutting number information indicates the number of channels remaining without being removed after the channel cutting or the number of channels removed through the channel cutting. In paragraph 8, A feature map encoding method, characterized in that the channel cutting is performed only for channels that are active in the feature map. Step of receiving a bitstream; A step of decoding a feature map from the above bitstream; and A step of arranging the channels of the decoded feature map in their original order based on metadata included in the bitstream, A feature map decoding method, characterized in that the above metadata includes reordering information of the channels. In Article 12, The above reordering information includes reordering order information, The above reordering order information is a one-dimensional vector array, A feature map decoding method, characterized in that the i-th element of the above reordering order information indicates the original index of the channel whose index is i. In Article 12, The above reordering order information includes section information and location information, The above section information indicates the section to which the original index of the channel belongs, A feature map decoding method, characterized in that the above location information indicates the location of the original index within the above section. In Article 12, A feature map decoding method, characterized in that the alignment is performed only for channels that are active in the feature map. In Article 12, The channels of the decrypted feature map are rearranged based on the activation level, A feature map decoding method characterized in that the original index of the channel is obtained from a channel index mapping table using the rearranged index of the channel as a key value. In Article 16, A feature map decoding method, characterized in that the above activation degree is set to a scale vector value of an inverse gain unit for each channel. In Article 12, The above feature map decoding method further includes a step of restoring the channels removed through channel cutting, A feature map decoding method, characterized in that the above metadata further includes channel cutting information. A step of rearranging the channels of the feature map; and Comprising the step of generating a bitstream including an encoded feature map and metadata, A computer-readable recording medium having recorded thereon a command for executing a feature map encoding method, wherein the above metadata includes reordering information of the channels.

Citation Information

Patent Citations

  • Air conditioner and multi airconditioner having the same

    KR1020240169878A

  • Virtual, augmented, and mixed reality-based livestock data for smart livestock farming: devices and methods

    KR1020250155387A

  • Radiative Cooling Window System

    KR102795347B1

  • Feature encoding / decoding method and device, and recording medium storing bitstream

    WO2023075563A1

  • KR20230046310A