A roi-aware deep joint source-channel coding image transmission method

CN122845802APending Publication Date: 2026-09-29UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610988517.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供一种ROI感知的深度联合源信道编码图像传输方法,用以解决现有深度联合源信道编码图像传输方法中未充分考虑图像区域语义重要性差异、关键区域传输可靠性低、接收端依赖额外ROI先验信息、重建图像感知质量不足以及难以满足语义通信场景需求等问题

Benefits of technology

[0036]1)关键区域优先保护:通过 ROI 引导的特征增强机制实现隐式功率重分配,在总发射功率约束下提升ROI区域的有效信噪比,0dB信噪比下ROI区域PSNR较基础Deep JSCC提升超过15dB,显著增强了低信噪比条件下关键语义信息的传输可靠性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845802A_ABST
    Figure CN122845802A_ABST
Patent Text Reader

Abstract

This invention belongs to the interdisciplinary field of wireless communication and digital image processing, specifically providing a ROI-aware deep joint source-channel coding (LSC) image transmission method to address numerous problems existing in current deep LSC image transmission methods. This invention achieves implicit power redistribution through an ROI-guided feature enhancement mechanism, improving the effective signal-to-noise ratio (SNR) of the ROI region under total transmit power constraints, significantly enhancing the transmission reliability of key semantic information under low SNR conditions. Simultaneously, it fuses shallow texture, mid-level structure, and deep semantic information through a multi-scale attention aggregation module, effectively improving the feature representation capability of the ROI region through the synergistic effect of channel and spatial attention. Furthermore, the receiver directly predicts the ROI mask from noisy features using a lightweight network, eliminating the need for the transmitter to transmit additional control information, reducing transmission overhead, and improving the system's practicality and deployability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of wireless communication and digital image processing, and specifically relates to an image semantic transmission method for distinguishing regions of interest based on deep joint source channel coding, specifically a ROI-aware deep joint source channel coding image transmission method. Background Technology

[0002] With the rapid development of 5G and future 6G wireless communication technologies, image data is increasingly being used in typical scenarios such as autonomous driving, smart cities, and remote monitoring. Traditional image transmission methods are based on Shannon separation theory, designing source coding and channel coding separately. They have good performance under long code and stable channel conditions, but they suffer from insufficient robustness under finite block length, low signal-to-noise ratio, and dynamic channel environments, and are prone to severe decoding distortion caused by error propagation.

[0003] In recent years, Deep Joint Source-Channel Coding (Deep JSCC) has directly established the mapping relationship between the image and the channel input through end-to-end learning, avoiding the error propagation problem of bit-level transmission and showing superior performance to traditional methods under weak channel conditions. However, existing Deep JSCC methods usually perform uniform optimization on the entire image, assuming that all regions have the same semantic importance, which results in key semantic regions (such as pedestrians, vehicles, and traffic signs) not being adequately protected under limited resources. At the same time, the optimization method based on pixel-level loss function tends to lead to overly smooth reconstructed images, resulting in significant deficiencies in visual perception quality. Furthermore, existing research has proposed some ROI-aware image transmission methods to address the issue of regional importance differences, but these methods still have the following shortcomings: First, most methods only weight ROIs at the loss function level, without truly changing the resource allocation method in the channel transmission stage, resulting in limited improvement for key regions; second, they lack a complete ROI processing closed loop, with the receiver often relying on the additional ROI mask information transmitted by the transmitter, increasing transmission overhead and system complexity; third, the optimization of perception quality and the priority protection of ROI are disconnected, making it impossible to improve visual realism while ensuring the structural accuracy of key regions. Summary of the Invention

[0004] The purpose of this invention is to provide a ROI-aware deep joint source channel coding image transmission method to solve the problems of existing deep joint source channel coding image transmission methods, such as insufficient consideration of the semantic importance differences of image regions, low transmission reliability of key regions, reliance on additional ROI prior information at the receiver, insufficient perceptual quality of reconstructed images, and difficulty in meeting the requirements of semantic communication scenarios.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for transmitting ROI-aware deep joint source channel coding images, characterized by including encoding transmission at the transmitting end and decoding reconstruction at the receiving end;

[0007] The transmitting end encoded transmission includes:

[0008] S1. Obtain the input image and generate a region of interest mask, mark the semantic key regions as regions of interest (ROI), and mark the remaining regions as non-regions of interest (NROI).

[0009] S2. The encoder is used to perform multi-layer convolutional feature extraction on the input image to obtain multi-scale semantic feature maps;

[0010] S3. Input the multi-scale semantic feature map into the multi-scale attention aggregation module, and enhance the feature expression of the ROI region through the synergistic effect of channel attention and ROI-guided spatial attention to obtain multi-scale attention-enhanced features;

[0011] S4. Perform ROI-guided feature enhancement operations on the multi-scale attention-enhanced features to obtain the features before transmission;

[0012] S5. Normalize the power of the pre-transmission features and map them into a complex signal, which is then transmitted to the receiving end through an additive white Gaussian noise channel.

[0013] The receiver decoding and reconstruction includes:

[0014] S6. Receive the noisy complex signal after transmission through the channel and restore it to the tensor form of the received feature Y;

[0015] S7. Input the received feature Y into the ROI recovery network to automatically predict the ROI estimated mask;

[0016] S8. Perform inverse enhancement operation on the received feature Y based on the ROI estimated mask to restore the feature scale to the decoding range;

[0017] S9. Input the inversely enhanced features into the decoder to obtain the reconstructed image.

[0018] Furthermore, the specific process of step S1 is as follows:

[0019] A pre-trained DeepLabv3 semantic segmentation model is used to perform pixel-level semantic parsing on the input image. The semantic key regions include: roads, pedestrians, vehicles, road signs and traffic signs. The corresponding pixels in the mask are marked as 1, and the rest are marked as 0, thus generating a region of interest mask.

[0020] Furthermore, in step S2, the multi-layer convolutional feature extraction is completed by the encoder, which consists of 5 convolutional layers. After each convolutional layer, a generalized division normalization (GDN) layer and a ReLU activation layer are set. The output features of the 1st, 3rd and 5th convolutional layers of the encoder are selected to form three-scale semantic features, which correspond to shallow texture features X1, mid-level structural features X2 and deep semantic features X3, respectively.

[0021] Furthermore, in the multi-scale attention aggregation module of step S3, the shallow texture feature X1 and the mid-level structural feature X2 are downsampled to the same resolution as the deep semantic feature X3, and then concatenated in the channel dimension to obtain the fused feature F;

[0022] Channel attention is applied to the fused feature F, and channel statistics vectors are obtained through global average pooling and global max pooling, respectively. and Channel features are obtained after passing through fully connected networks. and The two are added together and then activated by the Sigmoid function to generate channel weights. The channel enhancement feature is obtained by multiplying the fusion feature F channel by channel. ;

[0023] ROI-guided spatial attention is applied to the fused features, and multi-scale contextual information is extracted through 3×3, 5×5, and 7×7 convolutions, respectively. , and The sum of the three yields a multi-scale spatial representation. The average and maximum projections are calculated along the channel dimension, and the corresponding channel projection features are obtained. and The two are concatenated with the downsampled ROI mask, and the concatenated image is then activated by a Sigmoid function to generate a spatial attention map. The spatial enhancement feature is then obtained by element-wise multiplication with the fusion feature F. ;

[0024] Enhance channel features Spatial Enhancement Features After addition, the features are mapped through a 1×1 convolution to obtain the final multi-scale attention-enhanced features. .

[0025] Furthermore, the specific process of step S4 is as follows:

[0026] The multi-scale attention-enhanced features are divided into two parts according to the channel dimension: the unenhanced feature subset and the unenhanced feature subset. With enhanced feature subset For enhanced feature subsets Perform ROI-guided feature enhancement to obtain features :

[0027] ,

[0028] in, Indicates the ROI enhancement coefficient. Indicates the stabilizing factor. To provide a ROI mask downsampled to feature resolution, and Concatenation yields features before transmission .

[0029] Furthermore, in step S7, the ROI recovery network consists of three 3×3 convolutional layers and a sigmoid activation function, outputting an estimated ROI mask. The resolution is consistent with the ROI mask.

[0030] Furthermore, the specific process of step S8 is as follows:

[0031] Using the same partitioning method as the feature enhancement operation, the received feature Y is divided into an unenhanced feature subset. and inverse enhanced feature subset For inverse enhanced feature subsets Perform inverse enhancement operation to obtain features :

[0032] ,

[0033] Will and Inverse enhancement features are obtained by splicing. .

[0034] Furthermore, in step S9, the decoder and encoder adopt a symmetrical design, inverting the enhancement features. The input is to the decoder, which gradually restores the spatial resolution through multiple upsampling modules to obtain the reconstructed image. .

[0035] Based on the above technical solution, the beneficial effect of the present invention is that it provides a ROI-aware deep joint source channel coding image transmission method, which has the following advantages:

[0036] 1) Priority protection for critical areas: Implicit power redistribution is achieved through a feature enhancement mechanism guided by ROI, which improves the effective signal-to-noise ratio of the ROI area under the constraint of total transmit power. The PSNR of the ROI area is improved by more than 15dB compared with the basic Deep JSCC at a signal-to-noise ratio of 0dB, which significantly enhances the transmission reliability of critical semantic information under low signal-to-noise ratio conditions.

[0037] 2) ROI self-recovery without additional overhead: The receiver predicts the ROI mask directly from noisy features through a lightweight network, without the need for the sender to transmit additional control information, which reduces transmission overhead and improves the system's practicality and deployability.

[0038] 3) Multi-scale semantic feature enhancement: The designed multi-scale attention aggregation module integrates shallow texture, mid-level structure and deep semantic information. Through the synergistic effect of channel and spatial attention, it effectively enhances the feature expression ability of the ROI region.

[0039] 4) Dynamic resource scheduling: It can adaptively adjust the resource allocation strategy according to channel conditions, prioritize the quality of ROI region under low signal-to-noise ratio, and gradually restore the performance of background region under high signal-to-noise ratio, realizing dynamic optimization driven by semantic importance. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the structure of the Multiscale Attention Aggregation Module (MSAA) in this invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0042] This embodiment provides a method and system for ROI-aware deep joint source channel coding image transmission, including encoding transmission at the transmitting end and decoding reconstruction at the receiving end;

[0043] The encoded transmission at the sending end includes:

[0044] S1. Obtain the input image and generate a Region of Interest (ROI) mask. Mark semantically key regions such as roads, pedestrians, vehicles, road signs, and traffic signs as ROIs, and mark the remaining regions as non-ROIs.

[0045] A pre-trained DeepLabv3 semantic segmentation model is used to perform pixel-level semantic parsing on the input urban scene image (resolution 512×1024). Pixels corresponding to categories such as roads, pedestrians, vehicles, road signs, and traffic signs are labeled as 1 (ROI region), and all other pixels are labeled as 0 (NROI region), generating a binary ROI mask M∈{0,1}. 512×1024 ;

[0046] S2. Perform multi-layer convolution feature extraction on the input image to obtain semantic feature maps at different scales;

[0047] The input image is processed by an encoder consisting of 5 convolutional layers for feature extraction. Each convolutional layer is followed by Generalized Divisive Normalization (GDN) and ReLU activation to progressively reduce the spatial resolution and extract high-level semantic features. The output features of the encoder's 1st, 3rd, and 5th layers are selected, corresponding to shallow texture features X1, mid-level structural features X2, and deep semantic features X3, respectively.

[0048] S3. Input semantic feature maps of different scales into the multi-scale attention aggregation module, and enhance the feature representation of the ROI region through the synergistic effect of channel attention and ROI-guided spatial attention;

[0049] like Figure 1 As shown, in the multi-scale attention aggregation module, the shallow texture feature X1 and the mid-level structural feature X2 are downsampled to the same resolution (64×128) as the deep semantic feature X3, and then concatenated in the channel dimension to obtain the fused feature F;

[0050] Channel attention is applied to the fused feature F, and channel statistics vectors are obtained through global average pooling and global max pooling, respectively. and Channel features are obtained after passing through fully connected networks. and The two are added together and then activated by the Sigmoid function to generate channel weights. The channel enhancement feature is obtained by multiplying the fusion feature F channel by channel. ;

[0051] ROI-guided spatial attention is applied to the fused features, and multi-scale contextual information is extracted through 3×3, 5×5, and 7×7 convolutions, respectively. , and The sum of the three yields a multi-scale spatial representation. The average and maximum projections are calculated along the channel dimension, and the corresponding channel projection features are obtained. and The two are concatenated with the downsampled ROI mask, and the concatenated image is then activated by a Sigmoid function to generate a spatial attention map. The spatial enhancement feature is then obtained by element-wise multiplication with the fusion feature F. ;

[0052] Enhance channel features Spatial Enhancement Features After addition, the features are mapped through a 1×1 convolution to obtain the final multi-scale attention-enhanced features. ;

[0053] S4. Perform ROI-guided feature enhancement operations on multi-scale attention-enhanced features to increase the energy proportion of ROI region features before power normalization, thereby achieving implicit channel resource tilting.

[0054] The multi-scale attention-enhanced features are divided into two parts according to the channel dimension: the unenhanced feature subset. (First 8 channels) and enhanced feature subset (Later 8 channels) for enhancing feature subsets Execute ROI-guided feature enhancements:

[0055] ,

[0056] in, Indicates the ROI enhancement factor ( ), Indicates the stability factor ( ), To provide a ROI mask downsampled to feature resolution, and Concatenation yields features before transmission ;

[0057] S5. Normalize the power of the pre-transmission features and map them into a complex signal, which is then transmitted to the receiving end through an additive white Gaussian noise (AWGN) channel.

[0058] Sending features Flattening the vector to a real vector, we divide it into real and imaginary parts to construct a complex signal vector z, and then perform power normalization:

[0059]

[0060] Where n is the number of complex symbols, and the average transmit power of each complex symbol after normalization is 1;

[0061] The normalized signal is transmitted through an AWGN channel, and the noise variance is... , Signal-to-noise ratio (dB);

[0062] The receiving end decoding and reconstruction includes:

[0063] S6. Receive the noisy complex signal after transmission through the channel and restore it to the tensor form of the received feature Y;

[0064] S7. Input the received feature Y into the ROI recovery network to automatically predict the ROI estimation mask without the need for the sending end to transmit ROI information additionally.

[0065] The ROI recovery network consists of three 3×3 convolutional layers and a sigmoid activation function, outputting an estimated ROI mask. The resolution is consistent with the ROI mask;

[0066] S8. Based on the ROI estimated mask, perform inverse enhancement operation on the received features to restore the feature scale to a range suitable for decoding;

[0067] The received feature Y is divided into channels according to the channel dimension. (First 8 channels) and (Later 8 channels), for Perform the inverse enhancement operation:

[0068] ,

[0069] Will and Inverse enhancement features are obtained by splicing. ;

[0070] S9. Input the inversely enhanced features into the decoder to obtain the initial reconstructed image;

[0071] Inverse enhancement features The input is a decoder symmetrical to the encoder. This decoder gradually recovers the spatial resolution through multiple upsampling modules (including deconvolution layers or interpolation layers) to obtain the reconstructed image. .

[0072] The above description is merely a specific embodiment of the present invention. Any feature disclosed in this specification may be replaced by other equivalent or similar features unless otherwise specified. All disclosed features, or steps in all methods or processes, may be combined in any way except for mutually exclusive features and / or steps.

Claims

1. A method for transmitting ROI-aware, deep joint source-channel coded images, characterized in that, This includes encoding and transmission at the sending end and decoding and reconstruction at the receiving end; The transmitting end encoded transmission includes: S1. Obtain the input image and generate a region of interest mask, mark the semantic key regions as regions of interest (ROI), and mark the remaining regions as non-regions of interest (NROI). S2. The encoder is used to perform multi-layer convolutional feature extraction on the input image to obtain multi-scale semantic feature maps; S3. Input the multi-scale semantic feature map into the multi-scale attention aggregation module, and enhance the feature expression of the ROI region through the synergy of channel attention and ROI-guided spatial attention to obtain multi-scale attention-enhanced features; S4. Perform ROI-guided feature enhancement operations on the multi-scale attention-enhanced features to obtain the features before transmission; S5. Normalize the power of the pre-transmission features and map them into a complex signal, which is then transmitted to the receiving end through an additive white Gaussian noise channel. The receiver decoding and reconstruction includes: S6. Receive the noisy complex signal after transmission through the channel and restore it to the tensor form of the received feature Y; S7. Input the received feature Y into the ROI recovery network to automatically predict the ROI estimated mask; S8. Perform inverse enhancement operation on the received feature Y based on the ROI estimated mask to restore the feature scale to the decoding range; S9. Input the inversely enhanced features into the decoder to obtain the reconstructed image.

2. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, The specific process of step S1 is as follows: A pre-trained DeepLabv3 semantic segmentation model is used to perform pixel-level semantic parsing on the input image. The semantic key regions include: roads, pedestrians, vehicles, road signs and traffic signs. The corresponding pixels in the mask are marked as 1, and the rest are marked as 0, thus generating a region of interest mask.

3. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, In step S2, multi-layer convolutional feature extraction is performed by the encoder, which consists of 5 convolutional layers. After each convolutional layer, a generalized division normalization (GDN) layer and a ReLU activation layer are set. The output features of the 1st, 3rd and 5th convolutional layers of the encoder are selected to form three-scale semantic features, which correspond to shallow texture features X1, mid-level structural features X2 and deep semantic features X3, respectively.

4. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, In the multi-scale attention aggregation module of step S3, the shallow texture feature X1 and the mid-level structure feature X2 are downsampled to the same resolution as the deep semantic feature X3, and then concatenated in the channel dimension to obtain the fused feature F; Channel attention is applied to the fused feature F, and channel statistics vectors are obtained through global average pooling and global max pooling, respectively. and Channel features are obtained after passing through fully connected networks. and The two are added together and then activated by Sigmoid to generate channel weights. The channel enhancement feature is obtained by multiplying the fusion feature F channel by channel. ; ROI-guided spatial attention is applied to the fused features, and multi-scale contextual information is extracted through 3×3, 5×5, and 7×7 convolutions, respectively. , and The sum of the three yields a multi-scale spatial representation. The average and maximum projections are calculated along the channel dimension, and the corresponding channel projection features are obtained. and The two are concatenated with the downsampled ROI mask, and the concatenated image is then activated by a Sigmoid function to generate a spatial attention map. The spatial enhancement feature is then obtained by element-wise multiplication with the fusion feature F. ; Enhance channel features Spatial Enhancement Features After addition, the features are mapped through a 1×1 convolution to obtain the final multi-scale attention-enhanced features. .

5. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, The specific process of step S4 is as follows: The multi-scale attention-enhanced features are divided into two parts according to the channel dimension: the unenhanced feature subset and the unenhanced feature subset. With enhanced feature subset For enhanced feature subsets Perform ROI-guided feature enhancement to obtain features : , in, Indicates the ROI enhancement coefficient. Indicates the stabilizing factor. To provide a ROI mask downsampled to feature resolution, and Concatenation yields features before transmission .

6. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, In step S7, the ROI recovery network consists of three 3×3 convolutional layers and a sigmoid activation layer, outputting an estimated ROI mask. The resolution is consistent with the ROI mask.

7. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, The specific process of step S8 is as follows: Using the same partitioning method as the feature enhancement operation, the received feature Y is divided into an unenhanced feature subset. and inverse enhanced feature subset For inverse enhanced feature subsets Perform inverse enhancement operation to obtain features : , Will and Inverse enhancement features are obtained by splicing. .

8. The ROI-aware deep joint source-channel coded image transmission method according to claim 1, characterized in that, In step S9, the decoder and encoder adopt a symmetrical design to inversely enhance the features. The input is to the decoder, which gradually restores the spatial resolution through multiple upsampling modules to obtain the reconstructed image. .