Power image compression method, system and code stream transmission method for region of interest protection

By segmenting the region of interest and background regions of power images and performing bilinear downsampling, combined with a content-weighted attention module and intelligent bitstream management, the problems of inflexible bitrate allocation and block artifacts in power image compression are solved, achieving more efficient power image compression and bandwidth utilization.

CN120151534BActive Publication Date: 2025-11-14ZHEJIANG GUANGYAO DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510388191.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-11-14
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

Existing power image compression technologies struggle to flexibly allocate bit rates according to the actual needs of power equipment and backgrounds, resulting in poor image compression performance in regions of interest. Furthermore, traditional methods suffer from block artifacts and color distortion at high compression rates, leading to low bandwidth utilization during the transmission of multiple power image streams.

Method used

The method employs a region of interest and background region segmentation approach, smooths the background region through bilinear downsampling, combines a content-weighted attention module and prior information, dynamically allocates the bit rate, and achieves staggered and orderly transmission through intelligent bitstream management.

Benefits of technology

While maintaining the fidelity of the region of interest, this method reduces the consumption of compressed bitstream, improves the overall image compression effect, and optimizes the utilization of transmission bandwidth, thus solving the problems of inflexible bit rate allocation and block artifacts in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151534B_ABST
    Figure CN120151534B_ABST
Patent Text Reader

Abstract

This invention provides a power image compression method, system, and bitstream transmission method for region-of-interest (ROI) protection, relating to the field of power image compression technology. This invention distinguishes between the ROI and the background region, smoothing the background region through bilinear downsampling to remove spatial redundancy information in the background region of the power image, effectively reducing bit allocation in the background region. This achieves lower bitstream consumption while maintaining ROI fidelity and the same background quality. Simultaneously, through a content-weighted attention module, adaptive bit allocation at different spatial locations is implemented, rationally allocating bits according to the importance of image content to improve ROI encoding quality and overall compression performance. The bitstream transmission method proposed in this invention forwards the bitstream required for each image in a staggered and orderly manner to better utilize transmission bandwidth resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power image compression technology, and in particular to a power image compression method, system, and code stream transmission method for region of interest protection. Background Technology

[0002] In the operation and management of power systems, power image monitoring plays a crucial role. With the continuous expansion of power facilities and the increasing demand for intelligent monitoring, the volume of power image data is growing exponentially. For example, numerous cameras are deployed in widely distributed substations and transmission lines, continuously collecting image data. This data is essential for real-time monitoring of power equipment operation, timely detection of potential faults, and ensuring the safe and stable operation of the power system. To achieve effective storage, transmission, and subsequent analysis and processing of power image data, image compression technology has become a key element. Traditional image compression technologies mainly include methods based on standard encoding formats (such as H.264 and H.265). Although these have been applied to some extent in power image processing, they are insufficient to meet the demand for high-quality recovery of regions of interest (ROIs) in certain specific scenarios.

[0003] Several ROI-based image compression techniques have been proposed in the existing technology, such as:

[0004] The JPEG2000 ROI-based scheme has several limitations. The ROI-based maximum displacement method for JPEG2000 cannot control the relative importance of ROI and non-ROI regions. While the ROI-based general scaling method for JPEG2000 can control region importance, it can only handle regular ROI shapes, such as rectangles and ellipses. This restricts its application in power image compression, as it cannot flexibly allocate bitrates based on the actual needs of power equipment and background.

[0005] The High Efficiency Image Coding (HEVC) scheme addresses some limitations of the JPEG2000 image compression method based on the Region of Interest (ROI), but it also suffers from the following problems. First, its ROI method is developed at the coding unit (CTU) level, disallowing element-wise ROI spatial domain compression, which is detrimental to bit rate savings. Second, this method exhibits significant block artifacts and color distortion at high compression rates.

[0006] Learning-based ROI image compression methods rely solely on loss functions to constrain the bitstream allocation of the background region by masking the ROI region and the background region, making it difficult to effectively remove spatial redundancy information in the background region.

[0007] In short, JPEG2000 based on ROI struggles to achieve flexible bitrate allocation according to the actual needs of power equipment and background, significantly limiting its application in power image compression. While HEVC solves the ROI bit allocation problem in JPEG2000-based compression, it requires development at the CTU level, making pixel-level bit allocation difficult to achieve effectively. Learning-based ROI image compression algorithms also struggle to adequately address the optimization problem of object-fidelity image compression. Furthermore, existing methods rarely comprehensively consider the bitstreams of multiple power images, leading to data transmission congestion and low bandwidth utilization under certain bandwidth conditions.

[0008] Based on this, the present invention is proposed. Summary of the Invention

[0009] The purpose of this invention is to provide a power image compression method, system, and bitstream transmission method for region of interest (ROI) protection, which preserves ROI fidelity, consumes less compressed bitstream at the same quality, achieves better ROI and overall image compression effects, and has high transmission bandwidth utilization.

[0010] To achieve the above objectives, the present invention provides the following technical solution:

[0011] In a first aspect, the present invention provides a power image compression method for region of interest protection, comprising: S1. dividing an image x input from a compression terminal into regions of interest x. or and background area x br Partition such that x = x or +x br S2. Perform bilinear downsampling on the segmented background region to construct a background prior and obtain the intermediate features of the background region; S3. Encode the segmented region of interest to obtain the intermediate features of the region of interest, and merge and fuse the intermediate features of the region of interest with the intermediate features of the background region. Capture global features and weights through the internal content-weighted attention module to obtain the latent representation y of the overall image; S4. Quantize the latent representation y of the overall image to obtain the quantized latent representation of the overall image. S5. Implicit representation of the quantized overall image S6. Perform arithmetic encoding to form a bitstream; S7. Perform arithmetic decoding on the bitstream to restore the implicit representation of the entire image. S7. Implicit representation of the overall image after restoring the bitstream Decode to obtain the decoded image.

[0012] Preferably, the compression method further includes the following: S8. Further encoding the latent representation y of the overall image to introduce prior information to obtain the spatial dependency of y, thereby obtaining a further encoded latent representation z of the overall image; S9. Quantizing the prior latent representation z of the overall image to obtain the quantized prior latent representation of the overall image. S10. Hyperprior latent representation of the quantized overall image S11. Perform arithmetic encoding to form a bitstream; S12. Perform arithmetic decoding on the bitstream to restore the quantized overall image to its hyperprior latent representation. S12. To Perform advanced prior decoding and feed it into the spatial-channel context module to predict the latent representation of the entire image. The probability distribution.

[0013] In a second aspect, the present invention provides a power image compression system for region-of-interest (ROI) protection, employing the aforementioned compression method, comprising: an image segmentation module for segmenting the image x input from the compression end into a region of interest x. or and background area x br Partition such that x = x or +x br The prior module performs bilinear downsampling on the segmented background region to construct a background prior and obtain intermediate features of the background region. The encoder encodes the segmented region of interest (ROI) to obtain intermediate features of the RIO, and merges and fuses these intermediate features with the intermediate features of the background region. An internal content-weighted attention module captures global features and weights to obtain the latent representation y of the overall image. The first quantization module quantizes the latent representation y of the overall image to obtain the quantized latent representation of the overall image. The first arithmetic coding module is used for the implicit representation of the quantized overall image. Arithmetic encoding is performed to form a bitstream; the first arithmetic decoding module is used to perform arithmetic decoding on the bitstream to restore the implicit representation of the entire image. A decoder is used to reconstruct the implicit representation of the overall image from the bitstream. Decode to obtain the decoded image.

[0014] Preferably, the compression system further includes: a super encoder for further encoding the latent representation y of the overall image to introduce prior information to capture the spatial dependencies of y, thereby obtaining a further encoded prior latent representation z of the overall image; and a second quantization module for quantizing the prior latent representation z of the overall image to obtain a quantized prior latent representation of the overall image. The second arithmetic coding module is used for the implicit representation of the quantized overall image. The first module performs arithmetic encoding to form a bitstream; the second module performs arithmetic decoding to restore the bitstream to the implicit representation of the quantized overall image. Super decoder, used for Perform advanced prior decoding and feed it into the spatial-channel context module to predict the latent representation of the entire image. The probability distribution.

[0015] The present invention provides a bitstream transmission method in a third aspect, based on the above-mentioned power image compression method for region of interest protection, comprising the following steps: S21. Decoding the bitstream transmitted from each camera to obtain a decoded image; S22. Inputting the decoded image into the power image compression system for region of interest protection for data compression to obtain a compressed bitstream; S23. Calculating the bitstream size of each channel, and allocating transmission according to the current bandwidth through an intelligent bitstream management module to achieve staggered and orderly intelligent bitstream transmission.

[0016] Compared with existing technologies, the above technical solution has the following advantages:

[0017] This invention provides a power image compression method for region of interest (ROI) protection. It distinguishes between the ROI (target region) and the background region, smoothing the background region through bilinear downsampling to remove spatial redundancy and effectively reduce bit allocation in the background region. This achieves lower compressed bitstream consumption while maintaining ROI fidelity and the same background quality. Simultaneously, a content-weighted attention module enables adaptive bit allocation at different spatial locations, rationally allocating bits based on the importance of image content to improve ROI encoding quality and overall compression performance.

[0018] The present invention proposes a bitstream transmission method to forward the bitstream required for each image in a staggered and orderly manner, so as to better utilize transmission bandwidth resources. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 A block diagram of a power image compression system for region-of-interest protection provided in a specific embodiment of the present invention;

[0021] Figure 2This is a specific model block diagram of each module in a power image compression system for region of interest protection provided in a specific embodiment of the present invention.

[0022] Figure 3 A block diagram of a bitrate allocation model for a content-weighted attention module in a power image compression system with region-of-interest protection, provided as a specific embodiment of the present invention;

[0023] Figure 4 A block diagram of a block-based local attention module in a power image compression system for region-of-interest protection, provided as a specific embodiment of the present invention;

[0024] Figure 5 A flowchart illustrating a specific embodiment of the code stream transmission method provided by the present invention;

[0025] Figure 6 A more detailed flowchart of a code stream transmission method provided in a specific embodiment of the present invention is shown below.

[0026] Figure 1 In the accompanying figures, the following are the reference numerals: Image segmentation module 100, Prior module 200, Encoder 300, First quantization module 410, First arithmetic coding module 420, First arithmetic decoding module 430, Decoder 500, Super encoder 600, Second quantization module 710, Second arithmetic coding module 720, Second arithmetic decoding module 730, Super decoder 700, and Spatial-channel context module 900. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] In power substation monitoring, different cameras have different monitoring requirements for their respective areas of interest (such as switchgear, transmission lines, and background). This invention prioritizes bandwidth allocation to the switchgear area through ROI protection, while moderately compressing the background area to ensure the clarity of the critical target area (also referred to as the region of interest or ROI in this embodiment). To achieve this, in a preferred embodiment of this invention, a power image compression system oriented towards ROI protection is provided, such as... Figure 1As shown, it mainly consists of an image segmentation module, a priori module, an encoder, a first quantization module, a first arithmetic coding module, a first arithmetic decoding module, a decoder, a super encoder, a second quantization module, a second arithmetic coding module, a super decoder, and a spatial-channel context module. Based on this compression system, in a preferred embodiment of the present invention, a power image compression method (ROI-IC) for region of interest protection is provided. Please refer to [reference needed]. Figure 2 This is mainly achieved through the following steps:

[0029] S1. Determine the region of interest (ROI) for the image x input from the compression end. or and background area x br Partition such that x = x or +x br .

[0030] S2. Perform bilinear downsampling on the segmented background region to construct a background prior and obtain the intermediate features of the background region.

[0031] S3. Encode the segmented region of interest to obtain the intermediate features of the region of interest, and merge and fuse the intermediate features of the region of interest with the intermediate features of the background region. Capture global features and weights through the internal content-weighted attention module to obtain the latent representation y of the overall image;

[0032] S4. Quantize the latent representation y of the entire image to obtain the quantized latent representation of the entire image.

[0033] S5. Implicit representation of the quantized overall image Arithmetic encoding is performed to form a bitstream;

[0034] S6. Perform arithmetic decoding on the bitstream to restore the implicit representation of the entire image.

[0035] S7. Implicit representation of the overall image after restoring the bitstream Decode to obtain the decoded image.

[0036] S8. Further encode the latent representation y of the whole image to introduce super-prior information to obtain the spatial dependency of y, and obtain the further encoded super-prior latent representation z of the whole image;

[0037] S9. Quantize the hyperprior latent representation z of the overall image to obtain the quantized hyperprior latent representation of the overall image.

[0038] S10. Hyperprior latent representation of the quantized overall image Arithmetic encoding is performed to form a bitstream;

[0039] S11. Perform arithmetic decoding on the bitstream to restore the quantized overall image to its hyperprior latent representation.

[0040] S12. To Perform prior decoding and feed the data into the spatial-channel context module for prediction. The probability distribution;

[0041] S13. Utilize the information decoded from the super-prior in S12 and the partially decoded information in S6 to predict. The probability distribution is used for S5 pairs. Perform bitstream encoding.

[0042] The steps S1 to S7 above are described in more detail below:

[0043] At the compression end, given the input image An object detection algorithm, such as R-CNN (OROB), is used to segment the ROI region and the background region. Here, we use... and Let x represent the ROI region and the background region, such that x = x or +x br In a more preferred embodiment, to address the difficulty of achieving object-fidelity compression at low bit rates due to the transmission of background information, a scaling factor α is introduced to expand the proportion of the detected ROI region by object-oriented detection. Therefore, if the task requires greater focus on the neighborhood of a particular object, a higher α value can be chosen. Assuming the width and height of the ROI region are w and h, respectively, the width and height of the expanded ROI region will become (1+α)w and (1+α)h, respectively. In this way, the proposed method can adaptively focus on the expanded ROI region.

[0044] Subsequently, to reduce bit allocation in the background region, in a more preferred implementation, a 1 / 8 bilinear downsampling smoothing operation is applied to the background region. Following this, a background prior g is constructed. d To construct a background projection subspace, and then input it into encoder g. a In the middle, this determines the background information essential for compression. At the decompression end, the image is decoded. This was obtained under the condition of optimizing the regional difference loss. Specifically, it was achieved using g d The included background information, encoder g a The aim is to obtain the implicit representation y in a compact manner. Subsequently, y is quantized using the quantization function Q(·), resulting in... Finally, using and decoder gs You can then obtain the decoded image. Figure 2 In the code, AE and AD represent arithmetic code and arithmetic decoder, respectively.

[0045] The above processing procedure can be described as follows:

[0046] y = g a (x or ,g d (x br );φ),

[0047]

[0048] Among them, g a and g s These represent the encoder network and decoder network, respectively, with φ and θ referring to the parameters of the encoder network and decoder network, respectively.

[0049] like Figure 2 As shown, "2RB" refers to the sequential concatenation of two residual blocks; "Conv|k5|s2" represents a convolutional layer with a kernel size of 5×5 and a stride of 2 for downsampling; "TConv" represents a transposed convolutional layer for upsampling; and Attn represents a spatial attention module consisting of two sets of residual blocks combined with a sigmoid function.

[0050] The intermediate features of the background region and the intermediate features of the region of interest are first initially merged by adding them element by element, and then fed into... Figure 3 This approach aims to fuse information from background and target region features. Specifically, it first obtains a fused feature F through convolutional fusion; then, it normalizes the fused feature F, performs 3×3 convolution, height pooling, and width pooling to obtain query, key, and value matrices; given the query, key, and value matrices, it performs multi-head attention computation on the query and key matrices, and performs a dot product operation on the result with the value matrix; finally, it introduces a model enhancement network, inputting the result of the dot product operation into this network to obtain the long-range dependencies of the fused feature F. The global features and attention weights are reweighted to obtain the feature vector. Figure 3 The output is and Figure 3 Right now Figure 2 CWAM in the context, that is, CWAM is Figure 2 The encoder uses a content-weighted attention module to encode the input image to obtain the latent representation y. A more detailed explanation follows:

[0051] like Figure 2 As shown, in step S3, the merging of intermediate features is specifically achieved by adding elements from g. dand g a The intermediate features are initially merged, and through an addition operation, the background region and the ROI region are merged into a small-scale overall image feature. In a more preferred embodiment, a Content Weighted Attention (CWAM) module is introduced into the encoder to rebalance the importance of the ROI region and the background region in terms of bit allocation. Generally, the information content in an image varies spatially, and bit rate allocation based on the importance of different spatial contents of the image plays an important role in improving the rate-distortion performance of image compression. CWAM captures global features and dynamically allocates bits at different spatial locations. Here, CWAM uses a global attention network structure to realize global spatial domain feature association. Its bit rate allocation is implicitly constrained by a loss function. Figure 2 The CWAM in the encoder can be viewed as a global fusion module for the background and target regions. Guided by a loss function that differentiates between regions, it adaptively emphasizes the target region while discarding background redundancy to a certain extent, thereby ultimately achieving dynamic bit allocation in spatial location. The loss function is described in detail below. In a more preferred embodiment, two CWAMs are introduced into the encoder, which can achieve bit rate allocation at different scales.

[0052] The CWAM model or process for code rate (bit rate) allocation is as follows: Figure 3 As shown, the key process is as follows: A convolutional layer with a 1×1 kernel is used to fuse features from the ROI region and the background region to obtain the fused feature F. Considering the powerful modeling ability of the extended Transformer module in capturing global dependencies, it is used as the backbone structure of the content-weighted attention module design. Furthermore, to reduce computational complexity, average pooling operations are performed on the height and width dimensions of the intermediate features generated in the Transformer module, namely "height pooling" and "width pooling." Pooling the height and width reduces the complexity of matrix operations in the Transformer module. Query, key, and value matrix (W) Q ), (W K ) and (W V (can be accessed via W) Q =f Q (F), W K =f K (F) and W V =f V (F) is used to calculate, where f Q (·), f K (·) and f V (·) denotes a function based on a convolutional layer. Given W Q W K and W VThe matrix is ​​used to perform multi-head attention computation to achieve global attention. Then, a gated depthwise separable convolutional feedforward network (GDFN) with dilated convolutions is further employed to enhance the model's capabilities. This allows for the fusion of long-range dependencies of features. It can be derived as follows:

[0053]

[0054] Among them, W Q W K and W V These are the learned parameters, d k is a scaling factor used to avoid order-of-magnitude problems in the dot product results to stabilize the training process; it is learned through the network. RC(·) denotes the residual connection function, and T denotes the matrix transpose operation.

[0055] After obtaining the global features, attention operations are used to implement global bit allocation, as shown in the following expression:

[0056]

[0057] Among them, W A This refers to the learned attention weights, where ⊙ represents element-wise multiplication. This indicates that the feature vector output by the CWAM module is used as the input to subsequent networks.

[0058] The steps S8 to S13 above are described in more detail below:

[0059] via super encoder h a Advanced prior information is introduced to capture the spatial dependency of element y. After obtaining z through quantization, a super decoder h is used. s based on To estimate the conditional probability distribution. The above processing procedure can be described as follows:

[0060] z = h a (y;φ h ),

[0061]

[0062] Among them, h a and h s They refer to the super-encoder network and the decoder network, respectively. h and θ h Indicates model parameters. This represents the estimated probability distribution.

[0063] Using the Spatial-Channel Context Module (SCCTX) to Predict Latent Variables The probability distribution is estimated using the parameter π = (u, σ). Using this estimate, the first arithmetic encoder and decoder can... Compress it into a bitstream.

[0064] In a more preferred embodiment, considering the key role of the entropy model in improving rate distortion performance, a block-based local attention module is proposed in the design of the super-encoder and decoder networks to improve the estimation accuracy of the entropy model.

[0065] The super encoder, second quantization module, second arithmetic encoding module, second arithmetic decoding module, super decoder, and spatial-channel context module together constitute an entropy model based on a convolutional neural network. Both the super encoder and super decoder incorporate block-based Local Attention (PLAM) modules to dynamically reweight local information within the convolutional neural network-based entropy model, thereby improving the overall prediction accuracy of this entropy model. The accuracy of data distribution. A key aspect of lossy learning-based image compression lies in the entropy model, which is responsible for predicting the quantized latent representation. The probability distribution is calculated and fed into the second arithmetic coding module. This preferred embodiment provides an integrated block-based attention mechanism to dynamically reweight local information in the entropy model based on convolutional neural networks, thereby improving its estimation accuracy. To this end, a block-based local attention module is proposed, the model or process of which is as follows: Figure 4 As shown, this module mainly includes:

[0066] The segmentation operation is used to divide the intermediate feature image M generated in the super encoder or super decoder into P blocks in a B×B manner in a non-overlapping manner.

[0067] The Block Attention Submodule (PAM) is used to perform attention processing on each block;

[0068] The splicing operation is used to splice together the individual blocks that have undergone attention processing;

[0069] The Residual Block Submodule (2RB) is used to process the stitched features by applying two residual block layers;

[0070] The softmax activation function (S) is used to generate attention weights from features processed through two residual block layers using the softmax function.

[0071] The more detailed process is as follows:

[0072] First, the intermediate feature map M is divided into P blocks in a B×B format without overlap. p,i and r p,jLet be the feature values ​​of the i-th and j-th elements in the p-th block, respectively. The output of the block attention submodule can be expressed by the following formula:

[0073]

[0074] The definitions of g(·) and f(·) are as follows:

[0075]

[0076] in And τ(r) p,j ) = w τ r p .

[0077] Therefore, the final output can be obtained in the following way:

[0078] t p =w l s p +r p

[0079] Next, the attention-processed blocks are concatenated and two residual block layers are applied, followed by a softmax function to generate attention weights. The output of the block-based local attention module is used... It can be represented as follows:

[0080]

[0081] Where T = {t1, t2, ..., t} p ,...,t P} refers to the splicing result.

[0082] According to the formula The output of a block-based local attention module can be obtained, that is... Figure 2 The output characteristics of PLAM in the module. This module is used as a design... Figure 2 The main components of the super encoder and decoder are used to improve the prediction accuracy of the entropy model. Figure 4 In the middle, μ: 1×1, τ:1×1 and l:1×1 represent four convolutional layers with a kernel size of 1×1. p Let s represent the set of all feature elements of the p-th block. p This represents the output feature of the p-th block after being computed through the attention mechanism.

[0083] In a more preferred embodiment, the power image compression method and system for region of interest (ROI) protection provided by the above embodiment includes a convolutional neural network-based entropy model that requires training through two training phases: Phase 1: Pre-training the model at a low compression ratio. Phase 2: Training using parameters initialized from the pre-trained model, and adjusting the weights of the ROI and background regions to obtain a training model suitable for different compression ratios; the convolutional neural network entropy model implicitly constrains the bit allocation of global spatial domain features through a loss function. More detailed explanation follows:

[0084] Model training driven by a power image dataset: The entire compression system / method (ROI-IC) was trained using the power image dataset (RailFOD23, a publicly available dataset for foreign object detection on transmission lines, containing 14,615 images and 40,541 labeled objects, including four common types of foreign objects: bird nests, balloons, plastic bags, and floating debris). To train the proposed method, the mean squared error (MSE) metric and the Adam optimizer (with its two hyperparameters β1 = 0.9 and β2 = 0.999) were used for parameter optimization. Since the model needs to be trained at different λ values ​​to obtain models suitable for different compression ratios, it was first trained at a low compression ratio (i.e., λ = 4.5 × 10⁻⁶). -2 We pre-train one model. Afterward, other models with the target compression ratio are trained using parameters initialized from the pre-trained model. We use the default settings for the number of channels for all models and set λ = {8, 16, 32, 150, 450} × 10 for each model in the first training phase. -4 Under each λ, the weights η of the target region and the background region can be further adjusted to control the compression bit rate.

[0085] Furthermore, the hybrid quantization technique of SCCTX is employed to optimize the proposed model, namely the compression system / method (ROI-IC) proposed in this invention. Specifically, each Q(yu) is encoded as a bit instead of Q(y), and during decoding, Q(yu) + u is input to the decoder g. s Image decoding is performed. However, due to the decoder g... s The input is Q(yu) + u instead of Q(u), and this inconsistency makes training challenging. To address this, training is divided into two steps. In the first 80% of the total training rounds, noise n is sampled from a uniform distribution U = {0.5, 0.5} to generate... The analog quantized value is input to SCCTX, and S(y) is fed to the decoder g. s , where S is the pass-through estimator. In the remaining training rounds, S(yu)+u is input into gs and SCCTX for model training.

[0086] The probability distribution estimation parameters are π = (u, σ), where u represents the mean. Q(yu) represents the quantization of the residual of y relative to the estimated mean u. (By quantizing Q(yu), u is used as prior information to capture the local statistical characteristics of y, thus making the quantized representation more compact. In this way, the quantization error is mainly distributed within a small range of yu, reducing the overall error magnitude.)

[0087] Q(yu)+u represents adding the mean u (subtracted from the encoder) at the decoding end to recover the signal. S(y) represents the Straight-Through Estimator (STE). S(y) is an approximate quantization method that bypasses the non-differentiability problem of quantization during backpropagation. Its role is to simulate the quantization process in the early stages of training, avoiding the high nonlinear error caused by directly using hard quantization Q(y). Furthermore, while quantization Q(y) is non-differentiable, the straight-through estimator S(y) can propagate gradients during backpropagation, preventing training stagnation. S(yu)+u adds residual compensation during quantization. In the later stages of training, the model gradually transitions to directly using the quantized value S(yu)+u, thus maintaining consistency with the real input during the inference stage.

[0088] In this process, there is no explicit bit allocation. The model will automatically allocate bits through this global network during the learning and training process. This is because the loss function will give the target region a greater fidelity constraint and the background a smaller fidelity constraint. Under the constraints of the loss function, the global network will automatically allocate weights according to the loss function.

[0089] Loss Function: The proposed method requires two training phases. The first phase focuses on training a pre-trained model that achieves a low compression rate across the entire image. The second phase aims to introduce importance weights for the ROI and background regions to fine-tune the model obtained in the first phase.

[0090] In the first stage, the loss function can be defined as follows:

[0091]

[0092] Where R(·) represents the implicit representation after quantization. and The bitrate is obtained by calculating the entropy; D(·,·) represents distortion, which is used to measure the similarity between the input image and the decoded image, and is measured by mean square error; λ is a hyperparameter that balances bitrate R and distortion D.

[0093] In the second stage, rate-distortion optimization is performed by employing a region-discriminating loss function, which can be expressed as follows:

[0094]

[0095] Where N refers to the number of images input to the proposed model in each training iteration, i is the index, and x is the number of images input to the model. i The latter represents the i-th input image. This is the corresponding decoded image. This represents a weighted graph used for control. Bit allocation for the ROI region and background region. It is at element position. The value at this location can be represented as follows:

[0096]

[0097] Where, x bri It refers to the i-th input image x i The background region is defined as η ∈ [0.0, 0.5], which is a hyperparameter used to control the degree to which the background region is compressed.

[0098] Please refer to Figure 5 In one embodiment, a code stream transmission method is provided, based on the above-mentioned power image compression method for region of interest protection (ROI-IC), which mainly includes the following steps:

[0099] S21. First, decode the bitstream transmitted from each camera to obtain the decoded image;

[0100] S22. Input the decoded image into the power image compression system for protection of the region of interest for data compression to obtain a compressed bitstream;

[0101] S23. Calculate the bitstream size of each channel, and allocate the transmission through the intelligent bitstream management module according to the current bandwidth to achieve staggered and orderly transmission of intelligent bitstreams.

[0102] For a more detailed process, please refer to Figure 6 .

[0103] Based on the above preferred embodiments, the power image compression method, system, and code stream transmission method for region-of-interest protection provided by the present invention have the following beneficial technical effects.

[0104] 1. This invention proposes a power image compression method for Region of Interest (ROI-IC), integrating multiple techniques to achieve object-fidelity remote sensing image compression at low bit rates. The background smoothing operation reduces background bit allocation through bilinear downsampling, which is beneficial for achieving high overall rate-distortion performance. Furthermore, an intelligent bitstream management mechanism is incorporated into the ROI-IC method to better utilize transmission bandwidth.

[0105] 2. This invention employs a unique method for dividing and processing regions of interest (ROIs) and background regions. By using bilinear downsampling to smooth the background region, it effectively reduces bit allocation and solves the problem of excessive background information affecting object-fidelity compression at low bit rates. Simultaneously, the proposed CWAM (Coordinated Written Image) mechanism enables adaptive bit allocation at different spatial locations, rationally allocating bits based on the importance of image content to improve ROI encoding quality and overall compression performance.

[0106] 3. The PLAM module provided by this invention overcomes the limitation of traditional CNNs in treating local information equally when processing entropy models by incorporating a block-based local attention mechanism into the traditional CNN model and reweighting the local importance of each element. This improvement significantly enhances the estimation accuracy of the entropy model, thereby further improving rate-distortion performance and enabling the compressed image to better preserve object details and overall quality while maintaining a low bit rate.

[0107] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0109] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A power image compression method for region-of-interest protection, characterized in that, include: S1. Segment the region of interest and background region of the image input from the compressed end; S2. Perform bilinear downsampling on the segmented background region to construct a background prior and obtain the intermediate features of the background region; S3. Encode the segmented region of interest to obtain the intermediate features of the region of interest, and fuse the intermediate features of the region of interest with the intermediate features of the background region. Capture global features and weights through the internal content-weighted attention module to obtain the latent representation of the overall image. S4. Quantize the implicit representation of the whole image to obtain the quantized implicit representation of the whole image; S5. Arithmetic coding is performed on the implicit representation of the quantized overall image to form a bitstream; S6. Perform arithmetic decoding on the bitstream to restore the implicit representation of the entire image; S7. Decode the implicit representation of the overall image after the bitstream is restored to obtain the decoded image; The compression method is based on a power image compression system oriented towards region of interest protection, the system comprising: The image segmentation module is used to segment the region of interest and background region of the image input from the compression end; The prior module is used to perform bilinear downsampling on the segmented background region to construct the background prior and obtain the intermediate features of the background region. The encoder is used to encode the segmented region of interest, obtain the intermediate features of the region of interest, and merge and fuse the intermediate features of the region of interest with the intermediate features of the background region. The internal content-weighted attention module captures global features and weights to obtain the latent representation of the overall image. The first quantization module is used to quantize the implicit representation of the whole image to obtain the quantized implicit representation of the whole image. The first arithmetic coding module is used to perform arithmetic coding on the implicit representation of the quantized overall image to form a bit stream; The first arithmetic decoding module is used to perform arithmetic decoding on the bitstream to restore the implicit representation of the whole image; A decoder is used to decode the implicit representation of the overall image after the bitstream is restored, in order to obtain a decoded image; In the encoder, the intermediate features of the region of interest (ROI) and the intermediate features of the background region are merged and fused. The internal content-weighted attention module captures global features and weights to obtain the latent representation of the overall image. Specifically, the intermediate features of the background region and the intermediate features of the ROI are first initially merged by adding elements one by one, and then fused by convolution to obtain fused features. The fused features are then normalized, convolved, height pooled, and width pooled to obtain query, key, and value matrices. Given the query, key, and value matrices, multi-head attention is calculated on the query and key matrices, and the result is multiplied by the value matrix. A model enhancement network is introduced, and the result of the multiplication is input into the model enhancement network to obtain the long-distance dependencies, global features, and attention weights of the fused features. After reweighting, the feature vector is obtained.

2. The power image compression method for region-of-interest protection according to claim 1, characterized in that, It also includes the following: S8. Further encode the latent representation of the whole image to introduce super-prior information to obtain the spatial dependency of the latent representation of the whole image, and obtain the further encoded super-prior latent representation of the whole image. S9. Quantize the hyperprior latent representation of the whole image to obtain the quantized hyperprior latent representation of the whole image; S10. Arithmetic coding is performed on the super-prior hidden representation of the quantized overall image to form a bitstream; S11. Perform arithmetic decoding on the bitstream to restore the quantized overall image to its super-prior hidden representation; S12. Perform super-prior decoding on the super-prior latent representation of the quantized overall image and feed it into the spatial-channel context module to predict the probability distribution of the latent representation of the overall image. S13. Use the information decoded by the super prior in S12 and the partially decoded information in S6 to predict the probability distribution of the hidden representation of the whole image, so as to use S5 to perform bitstream encoding of the hidden representation of the whole image.

3. The power image compression method for region-of-interest protection according to claim 2, characterized in that, The system also includes: The super encoder is used to further encode the latent representation of the whole image to introduce super-prior information to obtain the spatial dependency of the latent representation of the whole image, and obtain the further encoded super-prior latent representation z of the whole image. The second quantization module is used to quantize the hyper-prior latent representation of the whole image to obtain the quantized hyper-prior latent representation of the whole image. The second arithmetic coding module is used to perform arithmetic coding on the hyperprior hidden representation of the quantized overall image to form a bit stream; The second arithmetic decoding module is used to perform arithmetic decoding on the bitstream to restore the quantized overall image to the super-prior hidden representation. The super decoder is used to perform super-prior decoding on the super-prior latent representation of the quantized overall image and feed it into the spatial-channel context module to predict the probability distribution of the latent representation of the overall image. The spatial-channel context module is used to predict the probability distribution of the latent representation of the whole image by using the information decoded by the superdecoder prior and the partial information already decoded by the first arithmetic decoding module, so as to enable the first arithmetic coding module to perform bitstream coding of the latent representation of the whole image.

4. The power image compression method for region-of-interest protection according to claim 3, characterized in that, The super encoder, second quantization module, second arithmetic encoding module, second arithmetic decoding module, super decoder, and spatial-channel context module together constitute an entropy model based on a convolutional neural network. Both the super encoder and super decoder incorporate block-based local attention modules to dynamically reweight local information within the convolutional neural network-based entropy model. Specifically, these modules include: Segmentation operations are used to divide intermediate feature images generated in a super encoder or super decoder into multiple blocks in a non-overlapping manner; The block attention submodule is used to perform attention processing on each block; The splicing operation is used to splice together the individual blocks that have undergone attention processing; The residual block submodule is used to process the stitched features by applying two residual blocks; The softmax activation function is used to generate attention weights from features that have been processed through two residual blocks.

5. The power image compression method for region-of-interest protection according to claim 4, characterized in that, The entropy model based on the convolutional neural network needs to be trained through two training phases; Phase 1: Pre-training the model at a low compression rate; The second stage involves training the model using parameters initialized from the pre-trained model and adjusting the weights of the region of interest and the background region to obtain a training model suitable for different compression ratios. The encoder network g implicitly constrains the bit allocation of global spatial domain features through a loss function.

6. A code stream transmission method, based on the power image compression method for region-of-interest protection as described in any one of claims 1 to 5, characterized in that, Including the following: S21. First, decode the bitstream transmitted from each camera to obtain the decoded image; S22. Input the decoded image into the power image compression system for protection of the region of interest for data compression to obtain a compressed bitstream; S23. Calculate the bitstream size of each channel, and allocate the transmission through the intelligent bitstream management module according to the current bandwidth to achieve staggered and orderly transmission of intelligent bitstreams.

Citation Information

Patent Citations

  • Screen content image compression method of two-stage octave convolution based on multi-scale residual error and window attention

    CN117544783A

  • Infrared weak and small target detection method fusing background prior information

    CN118154954A