Semantic scalable image coding method, system, device and storage medium
By adopting a semantic scalable image encoding framework based on tree structure in image compression, the problem of inconsistent image compression and feature compression optimization goals in the prior art is solved, and more efficient compression effects and wider applicability are achieved.
Patent Information
- Application Number
- CN202210585322.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-27
AI Technical Summary
The existing semantic scalable image compression technology based on chain structures is inconsistent in the optimization goals of image compression and feature compression during lossy compression, resulting in insufficient compression efficiency and failure to adapt to complex visual tasks.
Using a semantic scalable image encoding framework based on a tree structure, three levels of bit streams are formed through the basic layer and two branches, and some or all bit streams are compressed and transmitted according to the requirements of subsequent tasks.
Improves compression, can be applied to simple and complex visual tasks, while reducing computing resources and transmission volume.
Smart Images

Figure CN117197262B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image and video compression coding, and in particular to a semantic scalable image coding method, system, device and storage medium. Background Art
[0002] In the era of big data, various intelligent applications, such as smart cities and smart monitoring, generate massive amounts of image data every day. In order to process this data more efficiently, machines often perform visual analysis of images instead of humans. In a typical scenario, the image signal is collected by the camera at the front end, compressed and encoded, and transmitted to the cloud server. The cloud server decodes the transmitted signal and then performs visual analysis on the decoded signal. In this scenario, traditional image coding will bring about problems such as the amount of data required to be transmitted is too large, which hinders the widespread implementation of intelligent applications. In order to better serve human vision and machine vision at the same time, researchers have proposed a new image coding method - semantic scalable image coding. Semantic scalable image coding considers compressing the image signal and the features with semantic information extracted from the image at the same time to form a structured and scalable bit stream.
[0003] As deep learning technology continues to mature, deep neural networks are widely used to extract semantic information from images. In order to be applicable to visual analysis tasks at different levels, the bitstream of semantic scalable coding should contain semantic information of different granularities. For example, a three-layer scalable bitstream usually contains coarse-grained semantic information, fine-grained semantic information, and information at the image signal layer. The coarse-grained semantic information can be used for simpler visual tasks such as image classification, the fine-grained semantic information can be used for more complex visual tasks such as target detection, and the information at the image signal layer can be used to reconstruct images. In the context of cloud-based visual analysis, through semantic scalable coding, the sender can first send part of the bitstream containing coarse-grained semantic information to the cloud, and the cloud performs preliminary analysis based on this part of the coarse-grained semantic information. For samples that need further analysis, the sender then sends a bitstream containing fine-grained semantic information or image signal information to the cloud. In this way, the amount of data that needs to be transmitted and the computing resources required by the cloud server can be greatly reduced, thereby improving the efficiency of visual analysis.
[0004] As mentioned before, semantic scalable image coding considers compressing multiple levels of semantic features and input images simultaneously, such as compressing coarse-grained semantic features, fine-grained semantic features, and image signals. Most existing works organize bit streams based on a chain structure when performing compression, such as Figure 1 As shown, it mainly has the following defects:
[0005] 1) Coarse-grained semantic features are used as the base layer of the scalable bitstream, fine-grained semantic features are used as the first-level enhancement layer, and the image signal is used as the second-level enhancement layer. Therefore, when the receiving end (cloud) needs to reconstruct the input image, the front end needs to transmit the entire bitstream, and the receiving end also needs to complete the decoding of coarse-grained semantic features and fine-grained semantic features in turn before decoding the image signal.
[0006] 2) The existing semantically scalable image compression technology based on chain structure is based on the following assumption: the amount of information of high-level representation is always a subset of the amount of information of low-level representation. Specifically, the amount of information of the input image contains the amount of information of fine-grained semantic features, and the amount of information of fine-grained semantic features contains the amount of information of coarse-grained semantic features. However, this assumption is only valid in lossless compression. In lossy compression, image compression compresses information at the signal layer, with the goal of maintaining higher reconstruction accuracy for human vision, while feature compression compresses information at the semantic layer, with the goal of maintaining higher semantic accuracy for machine vision tasks. The inconsistency in the optimization goals of image compression and feature compression results in the amount of information of semantic features after lossy compression no longer being a subset of the amount of information of the compressed input image, which in turn leads to the insufficient efficiency of the existing semantically scalable image compression based on chain structure.
[0007] 3) Existing semantically scalable image compression methods only consider some relatively simple semantic tasks, such as image classification tasks of different granularities, but do not consider some more complex and more practical tasks (for example, image description and object detection tasks), which is also a limitation of existing technologies. Summary of the invention
[0008] The purpose of the present invention is to provide a semantic scalable image coding method, system, device and storage medium, which can improve the compression effect on the one hand and ensure the smooth execution of subsequent tasks on the other hand.
[0009] The objective of the present invention is achieved through the following technical solutions:
[0010] A semantically scalable image coding method, comprising:
[0011] Obtaining an original input image and its corresponding semantic features of two different granularities, the two semantic features of different granularities are respectively referred to as a first semantic feature and a second semantic feature;
[0012] Constructing a semantic scalable image coding framework based on a tree structure, the semantic scalable image coding framework based on the tree structure includes two branches, the front ends of the two branches share the same basic layer, the first branch includes a first-level enhancement layer and a second-level enhancement layer, and the second branch includes the first-level enhancement layer; the basic layer and the two branches constitute a three-level bit stream;
[0013] For the first semantic feature, compress it using a first feature compression model, and use a first bit stream obtained by compression as a base layer;
[0014] In the first branch, the first bit stream is reconstructed, the second semantic feature is compressed by using the first semantic feature obtained by reconstruction in combination with the second feature compression model, and the second bit stream obtained by compression is used as the first level enhancement layer of the first branch; the second bit stream is reconstructed, the original input image is compressed by using the second semantic feature obtained by reconstruction in combination with the image compression model, and the third bit stream obtained by compression is used as the second level enhancement layer of the first branch;
[0015] In the second branch, the first bit stream is reconstructed, and the original input image is compressed using the first semantic feature obtained by reconstruction in combination with an image compression model, and the fourth bit stream obtained by compression is used as the first-level enhancement layer of the second branch.
[0016] A semantically scalable image coding system, the system comprising:
[0017] A data information acquisition unit, used for acquiring an original input image and its corresponding semantic features of two different granularities, wherein the semantic features of the two different granularities are respectively referred to as a first semantic feature and a second semantic feature;
[0018] A semantic scalable image coding framework construction and coding unit, used to construct a semantic scalable image coding framework based on a tree structure, the semantic scalable image coding framework based on the tree structure includes two branches, the front ends of the two branches share the same basic layer, the first branch includes a first-level enhancement layer and a second-level enhancement layer, and the second branch includes the first-level enhancement layer; the basic layer and the two branches constitute three-level bit streams; for the first semantic feature, a first feature compression model is used for compression, and the first bit stream obtained by compression is used as the basic layer; in the first branch, the first bit stream is reconstructed, and the second semantic feature is compressed by combining the reconstructed first semantic feature with the second feature compression model, and the second bit stream obtained by compression is used as the first-level enhancement layer of the first branch; the second bit stream is reconstructed, and the original input image is compressed by combining the reconstructed second semantic feature with the image compression model, and the third bit stream obtained by compression is used as the second-level enhancement layer of the first branch; in the second branch, the first bit stream is reconstructed, and the original input image is compressed by combining the reconstructed first semantic feature with the image compression model, and the fourth bit stream obtained by compression is used as the first-level enhancement layer of the second branch.
[0019] A processing device, comprising: one or more processors; a memory for storing one or more programs;
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0021] A readable storage medium stores a computer program, which implements the above method when the computer program is executed by a processor.
[0022] It can be seen from the technical solution provided by the present invention that three levels of bit streams are formed by the basic layer and the two branches. Part or all of the bit streams can be compressed and transmitted according to the needs of subsequent tasks, thereby reducing computing resources and transmission volume while ensuring the smooth execution of subsequent tasks. It can be applied not only to simple semantic tasks (for example, image classification tasks of different granularities, etc.), but also to more complex tasks (for example, image description, target detection tasks, etc.); at the same time, the compression effect can be improved by jointly compressing images and features. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0024] Figure 1 Schematic diagram of a semantic scalable coding method based on a chain structure provided as the background technology of the present invention
[0025] Figure 2 A schematic diagram of a semantically scalable image coding method provided by an embodiment of the present invention;
[0026] Figure 3 A coding schematic diagram of a tree-structured semantic scalable image coding framework provided by an embodiment of the present invention;
[0027] Figure 4 A schematic diagram of a feature transformation module for feature compression provided by an embodiment of the present invention;
[0028] Figure 5 A schematic diagram of a second feature compression model provided by an embodiment of the present invention;
[0029] Figure 6 A schematic diagram of an image compression model provided by an embodiment of the present invention;
[0030] Figure 7 A first semantic feature compression performance comparison diagram provided by an embodiment of the present invention;
[0031] Figure 8A comparison chart of the compression performance of the second semantic feature provided by an embodiment of the present invention;
[0032] Fig. 9 A comparison chart of image compression performance provided by an embodiment of the present invention;
[0033] Fig.10 A schematic diagram of a semantically scalable image coding system provided by an embodiment of the present invention;
[0034] Fig.11 A schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0036] First, the terms that may be used in this article are explained as follows:
[0037] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.
[0038] The following is a detailed description of a semantic scalable image coding method, system, device and storage medium provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professionals in the field. If no specific conditions are specified in the embodiments of the present invention, they shall be carried out according to the conventional conditions in the field or the conditions recommended by the manufacturer.
[0039] Embodiment 1
[0040] The embodiment of the present invention provides a semantic scalable image coding method, such as Figure 2 As shown, it mainly includes:
[0041] Step 1: Obtain an original input image and its corresponding semantic features of two different granularities. The two semantic features of different granularities are respectively referred to as a first semantic feature and a second semantic feature.
[0042] In the embodiment of the present invention, the granularity of the first semantic feature is greater than that of the second semantic feature. For example, the first semantic feature and the second semantic feature are the coarse-grained semantic feature and the fine-grained semantic feature introduced in the above background technology section, respectively.
[0043] In the embodiment of the present invention, the semantic features of two different granularities can be extracted by an existing feature extraction network, and the present invention does not limit the type of the feature extraction network.
[0044] Step 2: Construct a semantically scalable image coding framework based on a tree structure, encode the original input image and its corresponding semantic features of two different granularities, and obtain three levels of bit streams.
[0045] Figure 3 The coding schematic diagram of the semantic scalable image coding framework based on the tree structure is shown. The semantic scalable image coding framework based on the tree structure includes two branches, the front ends of the two branches share the same basic layer, the first branch includes the first-level enhancement layer and the second-level enhancement layer, and the second branch includes the first-level enhancement layer; the basic layer and the two branches constitute a three-level bit stream.
[0046] In the embodiment of the present invention, the base layer + first branch is similar to the existing chain structure bit stream; the second branch has only one level of enhancement layer, which contains image signal information and can directly reconstruct the input image based on the base layer.
[0047] In the embodiment of the present invention, in order to reduce the redundancy of information in the joint compression of semantic scalable coding, the present invention proposes corresponding compression models (i.e., the second feature compression model and the image compression model) for the second semantic feature and the original input image, respectively. The overall semantic scalable image coding process can be described as follows:
[0048] Step 21: compress the first semantic feature using a first feature compression model, and use the first bit stream obtained by compression as a base layer.
[0049] Step 22, in the first branch, the first bit stream is reconstructed, and the second semantic feature is compressed by using the reconstructed first semantic feature in combination with the second feature compression model, and the compressed second bit stream is used as the first-level enhancement layer of the first branch; the second bit stream is reconstructed, and the original input image is compressed by using the reconstructed second semantic feature in combination with the image compression model, and the compressed third bit stream is used as the second-level enhancement layer of the first branch.
[0050] Step 23, in the second branch, the first bit stream is reconstructed, and the original input image is compressed using the first semantic feature obtained by reconstruction in combination with an image compression model, and the fourth bit stream obtained by compression is used as the first-level enhancement layer of the second branch.
[0051] In an embodiment of the present invention, in order to improve the efficiency of feature joint compression, a feature transformation module for feature compression is proposed, which is applied to two parts of semantic scalable image coding. Specifically: in the aforementioned step 21, the first semantic feature is firstly subjected to feature transformation, and then compressed using the first feature compression model. In the aforementioned step 22, the second semantic feature is subjected to feature transformation, and then compressed using the reconstructed first semantic feature combined with the second feature compression model.
[0052] like Figure 4 As shown, the feature transformation module for feature compression includes: a convolutional layer with a learnable output channel number of C dimensions, and a shuffling operation layer. Figure 4 An example of a convolutional layer with a convolution kernel size of 1×1 is given. The first semantic feature or the second semantic feature with an input size of H×W×C is fused in the channel dimension through a convolutional layer with a learnable output channel number of C. The channel dimension of the feature is then converted into a spatial dimension through a reordering operation layer, and a reordered feature map with a higher spatial resolution and fewer channels is output, with a size of Among them, H, W, and C are the height, width, and number of channels of the first semantic feature or the second semantic feature of the input respectively; r H With r W are all parameters in the rearrangement layer, which are used to control the size and number of channels of the rearranged features. The rearranged features can be compressed using the corresponding compression model.
[0053] As previously introduced, the embodiments of the present invention mainly include three types of compression models, namely: a first feature compression model, a second feature compression model and an image compression model.
[0054] 1) The first feature compression model can be a compression model based on a decomposition probability model, which mainly includes: a first encoding network and a first decoding network. As mentioned above, the first semantic feature needs to be transformed first, and the feature transformation result of the first semantic feature is compressed through the first encoding network to obtain a first bit stream; the first bit stream is processed through the first decoding network to reconstruct the first semantic feature.
[0055] 2) The second feature compression model. Figure 5As shown, the second feature compression model includes: a first encoder, a first quantization module, a first arithmetic encoder, a first arithmetic decoder, a first decoder and a conditional probability model; wherein the input of the conditional probability model is the reconstructed first semantic feature, and the output is the probability parameter required by the first arithmetic encoder and the first arithmetic decoder; firstly, the second semantic feature is subjected to feature transformation, and the feature transformation result of the second semantic feature is encoded by the first encoder, and then compressed by the first arithmetic encoder after passing through the first quantization module to obtain a second bit stream; the second bit stream is processed in turn by the first arithmetic decoder and the first decoder, and then the second semantic feature is reconstructed by feature inverse transformation.
[0056] Figure 5 In the figure, Q, AE, and AD represent the quantization module, the arithmetic encoder, and the arithmetic decoder, respectively; and They represent the first semantic feature and the first semantic feature reconstructed after compression, respectively, and F f is the second semantic feature, Z f represents the encoding result output by the first encoder, represents the decoded output of the first arithmetic decoder, is the conditional probability output by the conditional probability model, μ f ,σ f is the conditional probability The probability parameter of .
[0057] Figure 5 The dashed box on the right is the main structure of the conditional probability model, including: five convolutional layers arranged in sequence, wherein the first convolutional layer and the second convolutional layer are provided with a rearrangement operation layer, and an activation function layer (e.g., LeakyReLU layer) is provided between adjacent convolutional layers from the second convolutional layer to the fifth convolutional layer. Among them, K×K Conv, T indicates that the convolution kernel size of the convolutional layer is K×K (e.g., 1×1, 5×5 in the figure), the number of convolution kernels is T, T=Cf,N,M, Cf,N,M are all positive integers, and the specific values can be set according to actual conditions or experience, ↓2 indicates a downsampling operation.
[0058] 3) Image compression model. Figure 6As shown, the image compression model includes: a second encoder, a second quantization module, a second arithmetic encoder, a prediction module, and a super a priori probability network; wherein the input of the prediction module is the first semantic feature obtained by reconstruction, and the output is the mean and scale parameters required for compressing the original input image in the first branch; or, the input of the prediction module is the second semantic feature obtained by reconstruction, and the output is the parameters (mean m and scale s) required for compressing the original input image in the second branch; the second encoder encodes the original input image in combination with the mean and scale parameters output by the prediction module, and the encoding result is processed by the second quantization module and the second arithmetic encoder to obtain a first part of the bit stream, and the encoding result is processed by the super a priori probability network to obtain a second part of the bit stream, and the first part of the bit stream and the second part of the bit stream are combined as the third bit stream or the fourth bit stream.
[0059] Figure 6 The lower part also includes a series of network parts for image decoding, such as a decoder and an arithmetic decoder AD, etc. The relevant principles can refer to conventional technologies, so they are not described in detail. Figure 6 In, X, They represent the original input image and the reconstructed original input image respectively. The dotted line on the right is the structure of the prediction model. Figure 6 The image coding model shown is applied to two branches, so the prediction model can be or ResBlock represents the residual module, and ↑2 represents the upsampling operation.
[0060] It should be noted that the first and second involved in the present invention are mainly used to distinguish the same type of features. For example, the first encoder and the second encoder are mainly used to distinguish and represent two encoders; but due to Figure 5 and Figure 6 Each corresponds to a different compression model, so Figure 5 and Figure 6 In addition, the specific structures of the feature transformation module, feature compression model, and image compression model provided by the present invention are only examples. In practical applications, users can select existing feature transformation modules, feature compression models, and image compression models with other structures to apply to the present invention.
[0061] For ease of understanding, four examples of semantically scalable coding are provided below.
[0062] Example 1
[0063] In the example, the proposed semantic scalable coding method is combined with a feature extraction network based on an invertible neural network (i-RevNet). The compressed features in this example are the features of the last layer of i-RevNet, where the first semantic feature is 384-dimensional and the second semantic feature is 1536-dimensional. Features of different dimensions are obtained by modifying the number of output channels of the last convolutional layer.
[0064] The main steps of this example are as follows:
[0065] Step 1: Compress the first semantic feature. The parameter r of the feature transformation module is H and r W The values are 8 and 16 respectively. The transformed features are compressed using a compression model based on a decomposition probability model. The compressed bitstream is used as the base layer of the scalable bitstream.
[0066] Step 2: compress the second semantic feature based on the compressed and reconstructed first semantic feature. H and r W are taken as 16 and 32 respectively, and the transformed features are used as follows Figure 5 The compressed bit stream is used as the first level enhancement layer in the first branch of the tree structure bit stream.
[0067] Step 3: compress the image signal based on the first semantic feature reconstructed after compression in the first step. The image compression model is as follows: Figure 6 The compressed bitstream is used as the first level enhancement layer in the second branch of the tree structure bitstream.
[0068] Step 4: Compress the image signal based on the second semantic feature reconstructed after compression in the second step. The image compression model is as follows: Figure 6 The compressed bit stream is used as the second level enhancement layer in the first branch of the tree structure bit stream.
[0069] Example 2
[0070] In this example, the proposed semantic scalable coding method is combined with a feature extraction network based on a residual network (ResNet). The compressed features in this example are the features output by the fifth stage of ResNet, where the first semantic feature is 512-dimensional and the second semantic feature is 2048-dimensional.
[0071] The main steps of this example are as follows:
[0072] Step 1: Compress the first semantic feature. The parameter r of the feature transformation module is H and r WThe transformed features are compressed using a compression model based on a decomposition probability model. The compressed bitstream is used as the base layer of the scalable bitstream.
[0073] Step 2: compress the second semantic feature based on the compressed and reconstructed first semantic feature. H and r W are taken as 32 and 32 respectively, and the transformed features are used as follows Figure 5 The compressed bit stream is used as the first level enhancement layer in the first branch of the tree structure bit stream.
[0074] Step 3: compress the image signal based on the first semantic feature reconstructed after compression in the first step. The image compression model is as follows: Figure 6 The compressed bitstream is used as the first level enhancement layer in the second branch of the tree structure bitstream.
[0075] Step 4: Compress the image signal based on the second semantic feature reconstructed after compression in the second step. The image compression model is as follows: Figure 6 The compressed bit stream is used as the second level enhancement layer in the first branch of the tree structure bit stream.
[0076] Example 3
[0077] In this example, the proposed semantic scalable coding method is combined with a feature extraction network based on a residual network (ResNet). The compressed first semantic feature in this example is the 1024-dimensional feature output by the fourth stage of ResNet, and the second semantic feature is the 256-dimensional feature output by the second stage of ResNet. In this example, the second semantic feature has fewer dimensions because the feature space scale of the second stage is much larger than that of the fourth stage.
[0078] The main steps of this example are as follows:
[0079] Step 1: Compress the first semantic feature. The parameter r of the feature transformation module is H and r W The values are 16 and 32 respectively. The transformed features are compressed using a compression model based on a decomposition probability model. The compressed bitstream is used as the base layer of the scalable bitstream.
[0080] Step 2: compress the second semantic feature based on the compressed and reconstructed first semantic feature. H and r W are taken as 8 and 16 respectively, and the transformed features are used as follows Figure 5 The compressed bit stream is used as the first level enhancement layer in the first branch of the tree structure bit stream.
[0081] Step 3: compress the image signal based on the first semantic feature reconstructed after compression in the first step. The image compression model is as follows: Figure 6 The compressed bitstream is used as the first level enhancement layer in the second branch of the tree structure bitstream.
[0082] Step 4: Compress the image signal based on the second semantic feature reconstructed after compression in the second step. The image compression model is as follows: Figure 6 The compressed bit stream is used as the second level enhancement layer in the first branch of the tree structure bit stream.
[0083] In order to illustrate the effect of the above scheme of the present invention, the scheme of Example 1 is taken as an example to compare the performance with the image coding method that does not use semantic scalable coding. The scheme of Example 1 is abbreviated as Ours, and the image coding methods compared include: 1) the traditional image encoder BPG (Better Portable Graphics); 2) the image encoder based on deep learning proposed by Minnen et al. (denoted as Minnenetal.).
[0084] The first semantic feature is used for the image description generation task, and the evaluation index selected is CIDEr. Figure 7 Compression results of the first semantic feature on the image description generation task.
[0085] The second semantic feature is used for the target detection task, and the evaluation index selected is AP (Average Precision). Figure 8 It is the compression result of the second semantic feature in the object detection task, where Base Layer refers to the basic layer and Enh.Layer refers to the enhanced layer.
[0086] The evaluation index of image compression is PSNR (Peak Signal-to-Noise Ratio). The bit rate of image and feature compression is measured using bpp (bits per pixel). Fig. 9 is the image compression result, where B 1 Represents the first branch, B 2 Represents the second branch.
[0087] from Figure 7 and Figure 8 It can be seen that the solution proposed in the present invention has obvious advantages over traditional image compression methods in semantic analysis tasks. Fig. 9 It is shown in the figure that the solution proposed in the present invention can also achieve similar results as traditional image compression methods in image reconstruction tasks.
[0088] Embodiment 2
[0089] The present invention also provides a semantically scalable image coding system, which is mainly implemented based on the method provided in the above-mentioned embodiment 1. Fig.10 As shown, the system mainly includes:
[0090] A data information acquisition unit, used for acquiring an original input image and its corresponding semantic features of two different granularities, wherein the semantic features of the two different granularities are respectively referred to as a first semantic feature and a second semantic feature;
[0091] A semantic scalable image coding framework construction and coding unit, used to construct a semantic scalable image coding framework based on a tree structure, the semantic scalable image coding framework based on the tree structure includes two branches, the front ends of the two branches share the same basic layer, the first branch includes a first-level enhancement layer and a second-level enhancement layer, and the second branch includes the first-level enhancement layer; the basic layer and the two branches constitute three-level bit streams; for the first semantic feature, a first feature compression model is used for compression, and the first bit stream obtained by compression is used as the basic layer; in the first branch, the first bit stream is reconstructed, and the second semantic feature is compressed by combining the reconstructed first semantic feature with the second feature compression model, and the second bit stream obtained by compression is used as the first-level enhancement layer of the first branch; the second bit stream is reconstructed, and the original input image is compressed by combining the reconstructed second semantic feature with the image compression model, and the third bit stream obtained by compression is used as the second-level enhancement layer of the first branch; in the second branch, the first bit stream is reconstructed, and the original input image is compressed by combining the reconstructed first semantic feature with the image compression model, and the fourth bit stream obtained by compression is used as the first-level enhancement layer of the second branch.
[0092] Technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0093] Embodiment 3
[0094] The present invention also provides a processing device, such as Fig.11 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the aforementioned embodiments.
[0095] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0096] In the embodiment of the present invention, the specific types of the memory, input device and output device are not limited; for example:
[0097] The input device may be a touch screen, an image acquisition device, a physical button or a mouse, etc.;
[0098] The output device may be a display terminal;
[0099] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.
[0100] Embodiment 4
[0101] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.
[0102] In the embodiment of the present invention, the readable storage medium is a computer-readable storage medium and can be set in the aforementioned processing device, for example, as a memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk, etc., which can store program codes.
[0103] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A semantically scalable image coding method, characterized in that: include: Obtaining an original input image and its corresponding semantic features of two different granularities, the two semantic features of different granularities are respectively referred to as a first semantic feature and a second semantic feature; Constructing a semantic scalable image coding framework based on a tree structure, the semantic scalable image coding framework based on the tree structure includes two branches, the front ends of the two branches share the same basic layer, the first branch includes a first-level enhancement layer and a second-level enhancement layer, and the second branch includes the first-level enhancement layer; the basic layer and the two branches constitute a three-level bit stream; For the first semantic feature, compress it using a first feature compression model, and use a first bit stream obtained by compression as a base layer; In the first branch, the first bit stream is reconstructed, and the second semantic feature is compressed by using the first semantic feature obtained by reconstruction in combination with the second feature compression model, and the second bit stream obtained by compression is used as the first level enhancement layer of the first branch; Reconstructing the second bit stream, compressing the original input image by using the second semantic feature obtained by reconstruction in combination with an image compression model, and using a third bit stream obtained by compression as a second-level enhancement layer of the first branch; In the second branch, the first bit stream is reconstructed, and the original input image is compressed using the first semantic feature obtained by reconstruction in combination with an image compression model, and the fourth bit stream obtained by compression is used as the first-level enhancement layer of the second branch.
2. A semantically scalable image coding method according to claim 1, characterized in that: The granularity of the first semantic feature is greater than the granularity of the second semantic feature.
3. A semantically scalable image coding method according to claim 1, characterized in that: The first feature compression model includes: a first encoding network and a first decoding network; wherein, firstly, feature transformation is performed on the first semantic feature, and then the feature transformation result of the first semantic feature is compressed through the first encoding network to obtain a first bit stream; and the first bit stream is processed through the first decoding network to reconstruct the first semantic feature.
4. A semantically scalable image coding method according to claim 1, characterized in that: The second feature compression model includes: a first encoder, a first quantization module, a first arithmetic encoder, a first arithmetic decoder, a first decoder and a conditional probability model; wherein the input of the conditional probability model is the reconstructed first semantic feature, and the output is the probability parameter required by the first arithmetic encoder and the first arithmetic decoder; First, the second semantic feature is transformed, the feature transformation result of the second semantic feature is encoded by the first encoder, and then compressed by the first arithmetic encoder after passing through the first quantization module to obtain a second bit stream; The second bit stream is processed in sequence by the first arithmetic decoder and the first decoder, and the second semantic feature is obtained by reconstructing the feature inverse transformation.
5. A semantically scalable image coding method according to claim 4, characterized in that: The conditional probability model comprises: five convolutional layers arranged in sequence, wherein a rearrangement operation layer is provided in the first convolutional layer and the second convolutional layer, and an activation function layer is provided between adjacent convolutional layers in the second convolutional layer to the fifth convolutional layer.
6. A semantically scalable image coding method according to claim 1, 3 or 4, characterized in that: Two feature transformations are performed in the process of semantic scalable image coding; in the first feature transformation, the first semantic feature is transformed, and then compressed using the first feature compression model; in the second feature transformation, the second semantic feature is transformed, and then compressed using the reconstructed first semantic feature combined with the second feature compression model; The feature transformation is implemented by a feature transformation module for feature compression, and the feature transformation module for feature compression includes: a convolution layer with a learnable output channel number of C dimensions, and a rearrangement operation layer; The first semantic feature or the second semantic feature with an input size of H×W×C is fused in the channel dimension through a convolutional layer with a learnable output channel number of C. The channel dimension of the feature is then converted into a spatial dimension through a rearrangement operation layer. The output after rearrangement is of size feature map; where H, W, and C are the height, width, and number of channels of the first semantic feature or the second semantic feature of the input respectively; r H With r W These are parameters in the rearrangement operation layer, used to control the size and number of channels of the rearranged features.
7. A semantically scalable image coding method according to claim 1, characterized in that: The image compression model includes: a second encoder, a second quantization module, a second arithmetic encoder, a prediction module, and a hyper-prior probability network; wherein the input of the prediction module is the first semantic feature obtained by reconstruction, and the output is the mean and scale parameter required for compressing the original input image in the first branch; or the input of the prediction module is the second semantic feature obtained by reconstruction, and the output is the mean and scale parameter required for compressing the original input image in the second branch; The second encoder encodes the original input image in combination with the mean and scale parameters output by the prediction module, and the encoding result is processed by the second quantization module and the second arithmetic encoder to obtain a first part of the bit stream, and the encoding result is passed through the super prior probability network to obtain a second part of the bit stream, and the first part of the bit stream and the second part of the bit stream are combined as a third bit stream or a fourth bit stream.
8. A semantically scalable image coding system, characterized in that The method according to any one of claims 1 to 7 is implemented, and the system comprises: A data information acquisition unit, used for acquiring an original input image and its corresponding semantic features of two different granularities, wherein the semantic features of the two different granularities are respectively referred to as a first semantic feature and a second semantic feature; A semantic scalable image coding framework construction and coding unit, used to construct a semantic scalable image coding framework based on a tree structure, the semantic scalable image coding framework based on the tree structure includes two branches, the front ends of the two branches share the same basic layer, the first branch includes a first-level enhancement layer and a second-level enhancement layer, and the second branch includes the first-level enhancement layer; the basic layer and the two branches constitute three-level bit streams; for the first semantic feature, a first feature compression model is used for compression, and the first bit stream obtained by compression is used as the basic layer; in the first branch, the first bit stream is reconstructed, and the second semantic feature is compressed by combining the reconstructed first semantic feature with the second feature compression model, and the second bit stream obtained by compression is used as the first-level enhancement layer of the first branch; the second bit stream is reconstructed, and the original input image is compressed by combining the reconstructed second semantic feature with the image compression model, and the third bit stream obtained by compression is used as the second-level enhancement layer of the first branch; in the second branch, the first bit stream is reconstructed, and the original input image is compressed by combining the reconstructed first semantic feature with the image compression model, and the fourth bit stream obtained by compression is used as the first-level enhancement layer of the second branch.
9. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method for utilizing subject content analysis for producing a compressed bit stream from a digital image
US20030044078A1
Efficient scalable coding concept
US20150304667A1