Semantic masking image transmission method based on uncertainty perception in bandwidth-limited wireless channel
By employing an uncertainty-aware semantic masking image transmission method, the problem of image transmission under bandwidth constraints and channel fluctuations is solved, achieving high-quality image reconstruction in low-bandwidth environments and improving image transmission efficiency and reconstruction quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-04
AI Technical Summary
In wireless environments with limited bandwidth and large channel fluctuations, traditional image transmission methods struggle to simultaneously achieve both high-efficiency transmission and high-quality reconstruction. Especially under low bandwidth and high noise conditions, key and structural information in images is easily lost, and existing technologies lack channel adaptability and structural awareness.
The image transmission method using uncertainty-aware semantic masking utilizes a semantic segmentation network to calculate the semantic uncertainty, model uncertainty, and structural uncertainty of the image, generates a structure-aligned stripe mask map, dynamically adjusts the transmission ratio, and combines a generative reconstruction model to recover missing regions, thereby optimizing bandwidth utilization and channel adaptation.
Under conditions of limited bandwidth and unstable channels, this study aims to improve image reconstruction quality, maintain the coherence of key information and structure, enhance PSNR and MS-SSIM metrics, and achieve efficient and robust image transmission.
Smart Images

Figure CN122510366A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication and image transmission technology, and in particular to a semantic masking image transmission method based on uncertainty awareness in bandwidth-limited wireless channels. Background Technology
[0002] With the rapid development of wireless communication technology and the widespread adoption of smart terminals and Internet of Things (IoT) networks, how to achieve real-time visual data transmission in wireless environments with limited bandwidth and large channel fluctuations has become an increasingly important research problem in mobile multimedia systems. Traditional image communication systems typically rely on a "source coding + channel coding" structure based on Shannon information theory, compressing image redundancy and using channel coding for bit-level error correction. These traditional methods can approach Shannon capacity under high signal-to-noise ratio (SNR) conditions, but they face challenges such as limited bandwidth and dynamic fluctuations in channel noise in mobile applications. Especially when the upload bandwidth of IoT devices is extremely limited, transmitting high-resolution visual data becomes very difficult.
[0003] Traditional bit-level coding-based transmission methods cannot fully utilize the semantic information of images, nor can they effectively address the impact of bandwidth limitations and channel noise fluctuations on image transmission quality. Especially under low bandwidth and high noise conditions, once structurally complex or semantically important regions in an image are lost, conventional compression and transmission methods often cannot simultaneously meet the demands of bandwidth efficiency and high-quality image reconstruction. These problems prevent traditional technologies from achieving ideal visual effects under bandwidth-constrained and unstable wireless channel conditions.
[0004] Existing technical solutions and their disadvantages: The first category is the traditional separate source-channel coding method, which involves first compressing and encoding the image, and then using channel coding to achieve interference-resistant transmission. Examples include combining compression methods such as JPEG, JPEG2000, HEVC, and BPG with channel coding methods such as low-density parity-check codes. This type of method primarily focuses on pixel redundancy compression and error control, and has a mature application foundation when bandwidth is relatively sufficient and channel conditions are good. However, under conditions of significantly limited bandwidth and significant link fluctuations, this type of method often struggles to balance compression efficiency, error resistance, and receiver reconstruction quality.
[0005] The second category is semantic communication methods. These methods no longer focus solely on pixel-level fidelity but attempt to extract task-related features, semantic features, or high-level representation information from the image. They reduce communication overhead by transmitting semantic features instead of complete pixel data. While this type of method has certain advantages in task-oriented scenarios such as classification, recognition, and detection, for scenarios requiring the output of a complete visual image, transmitting only semantic features often fails to retain sufficient texture and local structural information, resulting in problems such as blurred edges, loss of detail, and structural distortion in the reconstructed image.
[0006] The third category is masking or selective transmission methods. Some methods use fixed masking, random masking, block-level sampling, or empirical region selection to transmit only a portion of pixels or local image regions, which are then completed by the receiver using interpolation or generative models. While this type of method can reduce transmission volume to some extent, existing solutions often employ fixed sampling rules or random masking strategies, making it difficult to accurately distinguish which regions are semantically important, which are difficult to predict, and which are more difficult to recover. Therefore, limited bandwidth resources cannot be prioritized for truly critical content regions.
[0007] In addition, existing technologies generally have the following shortcomings: 1. Insufficient determination of the importance of image regions: Most existing solutions do not simultaneously consider semantic boundary ambiguity, model prediction reliability and local structure recoverability, thus making it difficult to accurately model the importance of image content.
[0008] 2. Masking methods lack structural awareness: Random point masking or random block masking can easily destroy the continuous edges and texture directions of the image, making the receiving end restoration model lack stable structural guidance information, thus affecting the restoration quality.
[0009] 3. Lack of channel adaptation capability in transmission strategies: Many methods are designed with the assumption that the channel state is stable and cannot dynamically adjust the transmission ratio and transmission mode according to conditions such as signal-to-noise ratio and bandwidth utilization. This results in insufficient robustness in low signal-to-noise ratio scenarios and insufficient bandwidth utilization in high signal-to-noise ratio scenarios.
[0010] Therefore, there is an urgent need for an image transmission method that can perform fine perception of key areas in images for bandwidth-constrained wireless channels, and can dynamically adjust the transmitted content and masking ratio according to the real-time channel status, so as to improve both transmission efficiency and receiver reconstruction quality under limited communication resources. Summary of the Invention
[0011] To address the shortcomings of traditional technologies, the purpose of this invention is to solve the problem of efficient transmission of semantic and structural information during image transmission in bandwidth-constrained and channel-fluctuating environments. Specifically, the technical framework proposed in this invention, through an innovative uncertainty-aware semantic masking strategy, can maintain the visual quality of images, particularly the preservation of key information and structural regions, under low bandwidth and dynamic channel conditions, achieving high-quality image reconstruction. This overcomes the limitations of traditional methods in handling the semantic complexity and structural uncertainty of images while ensuring bandwidth efficiency.
[0012] To achieve the above objectives, the present invention provides a semantic masking image transmission method based on uncertainty awareness in bandwidth-constrained wireless channels, comprising the following steps: (1) Acquire the input image and perform preprocessing; (2) Input the preprocessed image into the semantic segmentation network and calculate the semantic uncertainty, model uncertainty, and structural uncertainty of its pixel position; then perform weighted fusion of the semantic uncertainty, model uncertainty, and structural uncertainty to obtain the joint uncertainty; (3) Define the local principal direction of the pixel position; based on the joint uncertainty, generate a structure alignment stripe mask map by controlling the target masking rate and the local principal direction information, and select the image pixels according to the mask map to obtain the pixels that need to be retained; (4) The transmitting end encodes the position, pixel value and mask map representation of the retained pixels to construct the transmission data and transmits it to the receiving end through the wireless channel; after receiving the data, the receiving end restores the position and pixel value of the retained pixels, constructs a mask image according to the mask map, and restores the missing area through the reconstruction model to obtain the final reconstructed image.
[0013] Further, step (2) includes: preprocessing the image Input the semantic segmentation network, and let the semantic segmentation network be located at the pixel position. Output the first The predicted probability of the class is Then semantic uncertainty Calculated using Shannon entropy: ; in, Total number of categories; pixel position Belongs to the The predicted probability of a class; pixel position Semantic uncertainty; assuming the same input image is executed After the second random forward propagation, at the pixel position The first obtained at the place The predicted probabilities of the classes are respectively Its mean is: ; Then the model uncertainty Calculated based on the predicted variance: ; in, For each random forward propagation number; For the first The average predicted probability of the class; Indicates pixel position Model uncertainty at the input image; Convert to grayscale Then calculate the horizontal gradient separately. and vertical gradient : ; Constructing a local structure tensor based on gradient information : ; Let the structure tensor eigenvalues and And satisfy Then define the local anisotropic characteristic quantity. for: ; In the formula, To prevent extremely small constants with a denominator of zero; structural uncertainty is defined. for: ; in, The larger the value, the more pronounced the local directionality and the clearer the structure of the region. The larger the value, the weaker the directionality of the region, the less stable the local structure, or the more dependent it is on context recovery.
[0014] Furthermore, in step (2), semantic uncertainty, model uncertainty, and structural uncertainty are weighted and fused to obtain joint uncertainty. : ; in, , , These are the weighting coefficients; satisfying... .
[0015] Further, step (3) includes: defining the target masking rate. Let be the proportion of pixels in the image that are masked and not directly transmitted, then: ; in, For target masking rate; Low concealment rate; High concealment rate; This refers to the current bandwidth resource utilization rate, bandwidth ratio, or equivalent resource usage indicators. To adjust the index.
[0016] Furthermore, in step (3), all pixel positions are sorted according to joint uncertainty, and stripe pixels in areas with higher uncertainty are preferentially retained under the constraint of satisfying the target retention rate.
[0017] Furthermore, in step (3), when the actual retention rate of the initial masking map is inconsistent with the target retention rate, the pixels are supplemented or further removed by adjusting the stripe period or stripe width, or by sorting based on joint uncertainty, so that the final retention rate meets the target constraint.
[0018] Further, in step (4), the data output by the transmitting end includes one or more of the following: the location index of the reserved pixel; the pixel value or feature value of the reserved pixel; the masking map or its compressed representation; the stripe direction related auxiliary parameters; the channel state related control parameters; and the auxiliary features used for reconstruction by the receiving end.
[0019] Furthermore, in step (4), the transmitting end selects different transmission methods according to the current signal-to-noise ratio: semantic feature transmission is selected under low signal-to-noise ratio conditions; pixel-level transmission is selected under medium-to-high signal-to-noise ratio conditions.
[0020] Furthermore, in step (4), the reconstruction model is a LaMa network.
[0021] Furthermore, in step (4), during the reconstruction process at the receiving end, the masking map... The received pixels are preserved as strong constraints, and the masking map is... The missing pixels are recovered by the reconstruction model based on the surrounding context, structural information and semantic cues.
[0022] The beneficial effects of this invention are: 1. Improved bandwidth efficiency: By using uncertainty-aware masking and generative reconstruction, this invention effectively reduces the amount of data transmitted. Under bandwidth-constrained conditions, important image areas are transmitted first, avoiding the waste of transmitting redundant data in traditional methods.
[0023] 2. Enhanced image reconstruction quality: This technology can maintain high-quality image reconstruction in low-bandwidth and unstable channel environments, especially in the recovery of key information and structural regions, and significantly improves evaluation indicators such as PSNR (Peak Signal-to-Noise Ratio) and MS-SSIM (Multi-scale Structural Similarity).
[0024] 3. Adaptive transmission strategy: The transmission strategy is dynamically adjusted according to the channel quality, which achieves robustness under low SNR and improvement of reconstruction details under high SNR, ensuring the stability and efficiency of the image under different channel conditions.
[0025] 4. Maintaining structural and semantic consistency of images: By introducing a structural alignment masking strategy, this technique can maintain the structural coherence and semantic consistency of images, reduce the image quality loss caused by traditional random masking methods, and significantly improve the recovery of texture and detail.
[0026] In summary, this invention optimizes each stage of image transmission within a framework of uncertainty awareness, enabling efficient and accurate transmission of image content in wireless communication environments with limited bandwidth and unstable channels. This significantly improves image reconstruction quality and promotes the application of semantic communication in mobile multimedia systems. Attached Figure Description
[0027] Figure 1 This is a graph showing the influence of the joint uncertainty weight parameters (α, β, γ) in this invention on the reconstructed PSNR under different signal-to-noise ratio conditions; Figure 2 This is a graph showing the variation of the masking ratio error under different signal-to-noise ratio conditions of the present invention; Figure 3 This is a comparison chart of PSNR of image reconstruction under different signal-to-noise ratios and different sampling times according to the present invention; Figure 4 This is the overall architecture diagram of the present invention. Detailed Implementation
[0028] To overcome the shortcomings of traditional methods, this invention proposes an uncertainty-aware semantic masking framework. This framework guides pixel-level mask generation and adaptive content transmission by jointly modeling semantic uncertainty, model prediction uncertainty, and structural uncertainty. Specifically, the proposed structure-aligned stripe masking strategy can preserve key information and structural regions, while recovering recoverable regions through a generative reconstruction model, thereby effectively improving the quality of image reconstruction.
[0029] Unlike traditional methods, the core technology of this invention lies in generating a transmission masking map through joint uncertainty mapping and dynamically adjusting the masking ratio based on bandwidth and channel quality to achieve efficient transmission of semantic and structural information. This innovative approach not only optimizes bandwidth utilization for image transmission but also improves visual quality in low-bandwidth environments, particularly in mobile communication environments.
[0030] Using only a single uncertainty cannot simultaneously reflect semantic boundary ambiguity, model prediction stability, and structural recoverability. This invention achieves multi-dimensional characterization of the importance of image regions through joint modeling of semantics, model, and structure, thereby significantly improving the accuracy of mask assignment.
[0031] Through these innovations, this invention enables high-quality image transmission in wireless network environments with limited bandwidth and unstable channels, and has significant practical value.
[0032] This invention solves the following technical problems: 1. Low image transmission efficiency under bandwidth constraints: Traditional image transmission technologies typically employ uniform bitstream transmission or fixed masking strategies. These methods fail to flexibly adjust transmission resources according to the importance (semantic or structural) of image regions, leading to wasted bandwidth and decreased image reconstruction quality. Especially in bandwidth-constrained mobile communication environments, traditional methods fail to prioritize the transmission of semantic and structural information.
[0033] This invention's solution employs an uncertainty-aware semantic masking strategy to dynamically select key regions for transmission, while other recoverable regions are reconstructed using a generative model, thereby optimizing bandwidth utilization efficiency. This not only reduces the amount of data transmitted but also ensures the effective transmission of critical information.
[0034] 2. Image quality degradation under channel fluctuations: Existing technologies assume stable channel conditions or optimize only under certain conditions, lacking adaptive mechanisms to cope with channel noise fluctuations. Under conditions of low signal-to-noise ratio (SNR) or large channel fluctuations, traditional methods often lead to a sharp decline in image reconstruction quality, failing to effectively guarantee the structural and semantic consistency of the image.
[0035] The solution of this invention introduces a channel-aware masking ratio control mechanism, which dynamically adjusts the masking ratio based on real-time channel quality (such as SNR), ensuring that important image details can be preserved under different channel conditions, while reducing the dependence on the reconstruction model when the channel is poor, thereby improving the image recovery capability in various environments.
[0036] 3. Loss of structural and semantic information of the image: In traditional masking transmission methods, random masking or fixed shape masking is often used. This method fails to consider the internal structure or semantic complexity of the image, resulting in the loss of structural information, especially the loss of detailed areas.
[0037] The solution proposed in this invention is a structure-aligned stripe masking strategy. This strategy generates a mask map based on the texture direction of the image, thereby better preserving the structural consistency of the image. This structured masking method reduces texture and structural distortion caused by random masking, enabling generative models to perform image restoration more effectively.
[0038] See Figure 4 The overall architecture of this invention includes a transmitter, a wireless channel, and a receiver. The transmitter performs data standardization on the input image, calculates semantic uncertainty, model uncertainty, and structural uncertainty, and fuses them to obtain joint uncertainty. Based on this joint uncertainty, a structure-aligned stripe mask map is generated using target masking rate control and local principal direction information. Based on this, image pixels are selected to obtain the pixels to be retained. Subsequently, the positions, pixel values, and mask map representations of the retained pixels are encoded to construct transmission data, which is then transmitted to the receiver via the wireless channel. After receiving the data, the receiver restores the positions and pixel values of the retained pixels, constructs a masked image based on the mask map, and further recovers the missing regions using a reconstruction model to obtain the final reconstructed image.
[0039] This invention provides a semantic masking image transmission method based on uncertainty awareness in bandwidth-constrained wireless channels, comprising the following steps.
[0040] Step S1: Acquire the input image and perform preprocessing.
[0041] The sending end obtains the image to be transmitted. The image is preferably a color image with a size of [size missing]. .
[0042] For the input image Normalization is performed to obtain the preprocessed image. : in, and These represent the mean and standard deviation obtained from the training data, respectively.
[0043] During the training phase, data augmentation can be applied to the input image, including random cropping, random flipping, and random scaling. Preferably, in one embodiment, the input image is uniformly scaled to [a specific scale value]. And randomly cut into Image patches are used for training; during the deployment phase, the image size can be selected based on the terminal's computing power and bandwidth conditions. , Or other preset sizes.
[0044] Step S2: Calculate semantic uncertainty.
[0045] Preprocessed image Input a semantic segmentation network. The semantic segmentation network is preferably a U-Net network (U-Net is a pixel-level segmentation network with an encoder-decoder structure), but other network structures capable of outputting pixel-level category probability distributions can also be used.
[0046] Let the semantic segmentation network be at the pixel position Output the first The predicted probability of the class is Then semantic uncertainty Calculated using Shannon entropy: in: Total number of categories; pixel position Belongs to the The predicted probability of a class; pixel position Semantic uncertainty.
[0047] when When the value is large, it indicates that the category boundary of the region corresponding to the pixel location is more ambiguous or that there is greater ambiguity in the semantic determination. This region is usually more worthy of priority preservation and transmission.
[0048] To avoid logarithmic calculation errors, you can use [the appropriate method] in the actual implementation. ,in It is a very small constant.
[0049] Step S3: Calculate model uncertainty.
[0050] To evaluate the stability of the model's prediction results, a random deactivation mechanism is enabled for the semantic segmentation network during the inference phase, and multiple random forward propagations are performed. Monte Carlo Dropout (MC Dropout) is preferred.
[0051] Assuming the same input image is executed After the second random forward propagation, at the pixel position The first obtained at the place The predicted probabilities of the classes are respectively Its mean is: Then the model uncertainty Calculated based on the predicted variance: in: For each random forward propagation number; For the first The average predicted probability of the class; Indicates pixel position The model uncertainty at that point.
[0052] In actual implementation, The values can be further summed or averaged along the category dimension to obtain the scalar uncertainty value for that pixel location. A large value indicates that the model's prediction for that region is unstable. If the region does not transmit at all at the transmitting end, it will be difficult for the receiving end to recover, so its transmission priority should be increased.
[0053] Preferably, in one embodiment, Choose 8 or 10.
[0054] Step S4: Calculate structural uncertainty.
[0055] To characterize the local texture orientation and structural recoverability of an image, the input image is first... Convert to grayscale Then calculate the horizontal gradient separately. and vertical gradient : Constructing a local structure tensor based on gradient information : Let the structure tensor eigenvalues and And satisfy Then define the local anisotropic characteristic quantity. for: in, To prevent extremely small constants with a denominator of zero.
[0056] Defining structural uncertainty for: in: The larger the value, the more pronounced the local directionality and the clearer the structure of the region. The larger the value, the weaker the directionality of the region, the less stable the local structure, or the more dependent it is on context recovery.
[0057] At the same time, the local principal direction of the pixel position Defined as: in, Used for generating subsequent structure alignment mask maps.
[0058] Step S5: Generate a joint uncertainty diagram.
[0059] The joint uncertainty is obtained by weighted fusion of semantic uncertainty, model uncertainty, and structural uncertainty. : in: , , These are the weighting coefficients; satisfying... .
[0060] Preferably, in one embodiment: When more attention is paid to semantic boundaries, it is advisable to ; When more emphasis is placed on structural restoration, it is advisable to... ; In general image scenarios, the following can be adopted: .
[0061] To unify subsequent scheduling processes, joint uncertainties can be addressed. Normalization to The intervals are divided into high uncertainty, medium uncertainty, and low uncertainty regions based on thresholds.
[0062] Preferably, in one embodiment: High uncertainty region satisfies ; Medium uncertainty region satisfies ; Low uncertainty region satisfies .
[0063] Preferably, .
[0064] Step S6: Generate a structure-aligned stripe masking map.
[0065] According to joint uncertainty and local principal direction Construct a structure to align the stripe masking pattern .
[0066] In this invention, it is agreed that: Indicates pixel position The pixels at that location are retained by the sending end and participate in the transmission; Indicates pixel position The pixels at a given location are not directly transmitted at the sending end, but are reconstructed by the receiving end.
[0067] To ensure the masking direction aligns with the local texture direction of the image, the pixel position is... Define along the local principal direction Projected coordinates : Different stripe periods are set for different uncertainty levels. and stripe width Therefore, the stripe masking rule is defined as follows: in: Indicates the first Stripe spacing at each level of uncertainty; Indicates the first The width of the retained stripes at each level of uncertainty.
[0068] Preferably, in one embodiment: In regions of high uncertainty, a denser retention strategy is employed, for example... ; A moderate retention strategy is adopted in areas of moderate uncertainty, for example... ; Regions with low uncertainty employ a sparser preservation strategy, for example... .
[0069] By using the above method, more real pixels can be retained in high uncertainty areas, while fewer pixels can be retained in low uncertainty areas, thereby achieving bandwidth priority allocation.
[0070] To reduce artifacts caused by stripe boundaries, in one embodiment, the initial mask image can be smoothed, morphologically processed, or edge-aligned to obtain the final binary mask image.
[0071] Step S7: Control the target masking rate according to the channel state.
[0072] This invention introduces a masking rate control mechanism based on channel state to dynamically adjust the proportion of pixels in an image that need to be directly transmitted according to the wireless link state.
[0073] Define target masking rate Let be the proportion of pixels in the image that are masked and not directly transmitted, then: For target masking rate; Low concealment rate; High concealment rate; This refers to the current bandwidth resource utilization rate, bandwidth ratio, or equivalent resource usage indicators. To adjust the index.
[0074] In this invention, the target retention rate can be expressed as: in, This indicates the proportion of pixels directly retained and transmitted by the sending end. Therefore, it can be seen that... and The pixel ratios correspond, and the semantics are consistent.
[0075] Preferably, in one embodiment: Furthermore, the current signal-to-noise ratio (SNR) can be incorporated into the adjustment strategy: When SNR is low, reduce ,improve That is, to retain more real pixels; When SNR is high, increase ,reduce This means transmitting fewer real pixels and relying on the receiver to reconstruct the image.
[0076] Preferably, in one embodiment: When the SNR is less than 4 dB, The range of values can be: ; When the SNR is between 4 dB and 10 dB, The range of values can be: ; When the SNR is greater than 10 dB, The range of values can be: .
[0077] In practice, it can be considered as a combination of uncertainties. Sort all pixel positions and then select the one that meets the target retention rate. Under the constraint of [the relevant factor], stripe pixels in areas with higher uncertainty are preferentially retained.
[0078] When the actual retention rate of the initial masking map is inconsistent with the target retention rate, the pixels can be supplemented or further removed by adjusting the stripe period or stripe width, or by sorting based on joint uncertainty, so that the final retention rate meets the target constraint.
[0079] Step S8: The sending end outputs and executes the transmission.
[0080] The transmitting end uses the final masking map. Select the pixel locations that need to be transmitted directly, and construct the data to be sent.
[0081] The data output by the sending end includes at least one or more of the following: Preserve the pixel's position index; Preserve the pixel value or feature value of the pixel; Masking diagram Or its compressed representation; Auxiliary parameters related to stripe direction; Channel state-related control parameters; Auxiliary features used for receiver reconstruction.
[0082] In one implementation, the transmitter can select different transmission methods based on the current SNR: Method A: Semantic feature transmission under low signal-to-noise ratio conditions When the SNR is low, in order to avoid bit-level transmission being sensitive to bit errors, Joint Source-Channel Coding (JSCC) can be used to send reserved area features and auxiliary guidance information.
[0083] Method B: Pixel-level transmission under medium to high signal-to-noise ratio conditions When SNR is high, it can be used for The corresponding reserved pixels undergo pixel-level bit transmission, and the encoded result of the masking map is sent synchronously. This pixel-level transmission can be achieved through a binary channel or other digital transmission methods.
[0084] To reduce additional overhead, masking maps The representation can be achieved using run-length encoding, block index encoding, sparse index encoding, or other compressed representation methods.
[0085] Step S9: The receiving end constructs a masked image and performs reconstruction.
[0086] After receiving the data transmitted by the sender, the receiver first restores the position and pixel value of the retained pixels, and then uses the masking map... Constructing masked images .
[0087] In this invention, Defined as: for Location, Retrieve the actual pixel values received from the receiver; for Location, Set to missing state, zero value, mean fill value or preset placeholder value.
[0088] The receiving end will mask the image. and concealment diagram Input Reconstruction Network To obtain the reconstructed image : in: To mask the image; For masking purposes; Image restoration network; For the final reconstructed image.
[0089] The reconstruction network is preferably a LaMa network (LaMa, a large-mask image inpainting network), but other convolutional neural networks, generative inpainting networks, or diffusion-based restoration networks can also be used. During the reconstruction process at the receiving end, The received pixels are retained as strong constraints. The missing pixels are recovered by the reconstruction network based on the surrounding context, structural information and semantic cues.
[0090] In one implementation, an uncertainty level is assigned to each pixel based on the joint uncertainty map and the local principal direction of the pixel, and the fringe period and fringe width are determined according to this level. For a pixel position (x, y), its projected coordinates along the local principal direction are first calculated, and then a periodic modulo operation is performed on the projected coordinates to determine whether it falls within the reserved fringe interval; if it falls within the reserved fringe interval, the pixel is marked as reserved; otherwise, it is marked as masked. This forms a fringe sampling pattern with consistent direction in the local texture direction of the image. Finally, the initial mask map is smoothed and its connectivity is corrected to obtain the final structure-aligned fringe mask map.
[0091] Input: Joint uncertainty plot U; local principal orientation plot θ; high threshold Th; medium threshold Tm; fringe periods P1, P2, P3 for each layer; fringe widths W1, W2, W3 for each layer; Output: Structure-aligned stripe masking map M.
[0092] Includes the following steps: 1. Normalize U to obtain ; 2. For each pixel (x, y): 2.1 If If Th >= 1, then level l = 1; 2.2 If Tm <= If Th < Th, then level l = 2; 2.3 If If Tm < Tm, then level l = 3; 2.4 Read the stripe period Pl and stripe width Wl corresponding to this level; 2.5 Calculate the projected coordinates: d = round(x*cosθ(x,y) + y*sinθ(x,y)) 2.6 Calculate the margin: r = d mod Pl 2.7 If r < Wl, then M0(x,y)=1; otherwise, M0(x,y)=0 3. Perform morphological smoothing or median filtering on the initial masking map M0; 4. Output the final masking map M.
[0093] The stripe period, stripe width, and uncertainty threshold can be preset based on the reconstruction quality index and target transmission ratio on the validation set, or obtained automatically through the training process.
[0094] Example 1 1. Experimental Environment and Equipment Hardware environment: Experiments were conducted on a computer with GPU acceleration, including at least one NVIDIA GPU, for deep learning training and inference.
[0095] Software environment: Python 3.8, PyTorch framework, used to implement Joint Source Channel Coding (JSCC), uncertainty calculation, and structure-guided LaMa image inpainting network.
[0096] Channel simulation: Simulate bandwidth-constrained wireless channels, including: White Noise Gaussian Channel (AWGN) with SNR ranging from 1 dB to 15 dB. Semantic Feature Transmission (JSCC) is used at low SNR (1–4 dB), and Pixel-Level Bit Transmission (BSC) is used at medium to high SNR.
[0097] 2. Dataset Training dataset: DIV2K, used to train the JSCC network and uncertainty estimation module.
[0098] Test datasets: Kodak (24 images, resolution 768×512); CLIC2021 (resolution uniformly adjusted to 512×512).
[0099] Data augmentation: Randomly crop 256×256 patches, randomly flip and scale to enhance the robustness of the model.
[0100] 3. Experimental Methods (1) Input image standardization.
[0101] Given an input RGB image Normalized to: in and These are the image mean and standard deviation, respectively. The input is an RGB color image with dimensions H×W×3 (height × width × number of color channels).
[0102] (2) Uncertainty calculation.
[0103] Semantic uncertainty: Calculating the class probability per pixel using a pre-trained semantic segmentation network. And calculate Shannon entropy: C represents the pixel coordinates in the image; C represents the total number of semantic categories. The segmentation network predicts the probability that the pixel belongs to category i; The semantic uncertainty of a pixel (x,y) is measured by Shannon entropy. A larger value indicates greater uncertainty in semantic prediction.
[0104] Model uncertainty: Enable Monte Carlo Dropout during the inference phase. Second forward propagation, calculate prediction variance: The model's prediction variance for this pixel is obtained through multiple forward inferences using Monte Carlo Dropout. Structural uncertainty: Calculating local anisotropy based on structural tensors : The local anisotropy index obtained from the structure tensor calculation measures the structural strength and directionality around a pixel. Structural uncertainty. A larger value indicates that the structure near the pixel is not obvious or is difficult to recover using a generative model.
[0105] Joint uncertainty: Weighting coefficients control the contribution of each uncertainty to the joint uncertainty. + + =1 weight normalization.
[0106] Pixels are categorized into low, medium, and high uncertainty levels using quantization methods for spatial transmission scheduling.
[0107] (3) Structure-guided stripe mask generation.
[0108] According to the local principal direction Generate stripe mask Different fringe spacings are used for different uncertainty layers: in This indicates that the pixel is retained and sent to the receiving end.
[0109] (4) Transmission strategy.
[0110] Low SNR: JSCC semantic feature transmission; Medium-high SNR: Reserved area pixel bit transmission (BSC); Adaptive masking rate control: Target masking rate (the proportion of pixels to be masked). The maximum and minimum masking rates are controlled by channel conditions. Current bandwidth percentage or channel resource utilization rate. Adjust the index to control the sensitivity of the masking rate to changes in bandwidth.
[0111] Dynamically adjust based on channel bandwidth and signal-to-noise ratio (5) Reconstruction of the receiving end.
[0112] The LaMa structure is used to guide the image inpainting network to restore the masked areas. The input is a mask and the received pixels, ensuring structural continuity and semantic consistency. in To cover up the image, , Mask indicator matrix (masking map), 1 indicates a pixel that needs to be repaired. Structure-guided image inpainting networks (such as LaMa) use masks and received pixels to generate reconstructed images.
[0113] 4. Experimental Evaluation Indicators Pixel-level reconstruction quality: Peak signal-to-noise ratio (PSNR); Structure and perceived quality: Multiscale structural similarity (MS-SSIM); Mask estimation accuracy: mask ratio error; Computational overhead: runtime of each module.
[0114] 5. Comparison Methods BSC-only: Pixel-bit transmission; JSCC-only: Semantic feature transfer; SCSC: Ultracompressed Semantic Communication Framework; BPG + LDPC: Traditional split-source channel coding; The method of this invention: hybrid semantic-pixel transmission, adaptive mask and structural stripe mask.
[0115] 6. Analysis of Experimental Results like Figure 1 The joint uncertainty weighting parameters (α, β, γ) shown have an impact on the reconstructed PSNR under different signal-to-noise ratio conditions; from Figure 1 As can be seen, the image reconstruction PSNR under each set of joint uncertainty weight parameters gradually increases with the increase of signal-to-noise ratio, indicating that the joint uncertainty-aware scheduling mechanism proposed in this invention can adapt to different channel conditions. Furthermore, the parameter combination with higher structural uncertainty weights achieves better reconstruction results, indicating that strengthening the characterization of image structural recoverability during mask generation and transmission scheduling helps improve the overall quality of the reconstructed image at the receiving end. Figure 2 The trend of masking ratio error with signal-to-noise ratio is shown, and the impact of different Monte Carlo sampling times on masking generation accuracy is compared to illustrate the effectiveness of the method of the present invention in terms of uncertainty estimation stability and spatial scheduling accuracy. Figure 3 The trend of reconstructed PSNR with signal-to-noise ratio is shown, and the impact of different sampling times on image reconstruction quality is compared to illustrate the effectiveness of the method of the present invention in improving image restoration accuracy.
[0116] Experimental results show that both PSNR and MS-SSIM steadily improve with increasing SNR. The joint uncertainty + structured fringe mask exhibits the best reconstruction performance under all SNR conditions, especially at medium to high SNR. The mask ratio error decreases significantly with increasing SNR, and MC Dropout ( It provides more stable estimates of uncertainty. The overall computational cost increases by only about 3%, which is acceptable.
[0117] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A method for uncertainty-aware semantic cover image transmission in bandwidth-constrained wireless channels, characterized in that, Includes the following steps: (1) Acquire the input image and perform preprocessing; (2) Input the preprocessed image into the semantic segmentation network and calculate the semantic uncertainty, model uncertainty, and structural uncertainty of its pixel position; then perform weighted fusion of the semantic uncertainty, model uncertainty, and structural uncertainty to obtain the joint uncertainty; (3) Define the local principal direction of the pixel position; based on the joint uncertainty, generate a structure alignment stripe mask map by controlling the target masking rate and the local principal direction information, and select the image pixels according to the mask map to obtain the pixels that need to be retained; (4) The transmitting end encodes the position, pixel value and mask map representation of the retained pixels to construct the transmission data and transmits it to the receiving end through the wireless channel; after receiving the data, the receiving end restores the position and pixel value of the retained pixels, constructs a mask image according to the mask map, and restores the missing area through the reconstruction model to obtain the final reconstructed image.
2. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, Step (2) includes: preprocessing the image Input the semantic segmentation network, and let the semantic segmentation network be located at the pixel position. Output the first The predicted probability of the class is Then semantic uncertainty Calculated using Shannon entropy: ; in, Total number of categories; pixel position Belongs to the The predicted probability of a class; pixel position Semantic uncertainty; assuming the same input image is executed After the second random forward propagation, at the pixel position The first obtained at the place The predicted probabilities of the classes are respectively Its mean is: ; Then the model uncertainty Calculated based on the predicted variance: ; in, For each random forward propagation number; For the first The average predicted probability of the class; Indicates pixel position Model uncertainty at the input image; Convert to grayscale Then calculate the horizontal gradient separately. and vertical gradient : ; Constructing a local structure tensor based on gradient information : ; Let the structure tensor eigenvalues and And satisfy Then define the local anisotropic characteristic quantity. for: ; In the formula, To prevent extremely small constants with a denominator of zero; structural uncertainty is defined. for: ; in, The larger the value, the more pronounced the local directionality and the clearer the structure of the region. The larger the value, the weaker the directionality of the region, the less stable the local structure, or the more dependent it is on context recovery.
3. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 2, characterized in that, In step (2), semantic uncertainty, model uncertainty, and structural uncertainty are weighted and fused to obtain joint uncertainty. : ; in, , , These are the weighting coefficients; satisfying... .
4. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, Step (3) includes: defining the target masking rate. Let be the proportion of pixels in the image that are masked and not directly transmitted, then: ; in, For target masking rate; Low concealment rate; High concealment rate; This refers to the current bandwidth resource utilization rate, bandwidth ratio, or equivalent resource usage indicators. To adjust the index.
5. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, In step (3), all pixel positions are sorted according to joint uncertainty, and stripe pixels in areas with higher uncertainty are preferentially retained under the constraint of satisfying the target retention rate.
6. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, In step (3), when the actual retention rate of the initial masking image is inconsistent with the target retention rate, the pixels are supplemented or further removed by adjusting the stripe period or stripe width, or by sorting based on joint uncertainty, so that the final retention rate meets the target constraint.
7. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, In step (4), the data output by the transmitting end includes one or more of the following: the location index of the reserved pixel; the pixel value or feature value of the reserved pixel; the masking map or its compressed representation; the stripe direction related auxiliary parameters; the channel state related control parameters; and the auxiliary features used for reconstruction by the receiving end.
8. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, In step (4), the transmitting end selects different transmission methods according to the current signal-to-noise ratio: semantic feature transmission is selected under low signal-to-noise ratio conditions; pixel-level transmission is selected under medium-to-high signal-to-noise ratio conditions.
9. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, In step (4), the reconstructed model is a LaMa network.
10. The semantic masking image transmission method based on uncertainty awareness in a bandwidth-constrained wireless channel according to claim 1, characterized in that, In step (4), during the reconstruction process at the receiving end, the masking map The received pixels are preserved as strong constraints, and the masking map is... The missing pixels are recovered by the reconstruction model based on the surrounding context, structural information and semantic cues.