A semantic communication method and system for high-resolution images
By transforming the image from the pixel domain to the feature domain and using the hyper-prior distribution estimation and quadtree spatial entropy model for step-by-step prior estimation, the cliff effect problem of traditional high-resolution image communication systems under harsh channel conditions is solved, and efficient image reconstruction and semantic accuracy are achieved.
Patent Information
- Application Number
- CN202411494669.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-10-24
AI Technical Summary
When traditional high-resolution image communication systems are subjected to harsh channel conditions or limited bandwidth, a mismatch between the communication transmission code rate and the channel capacity will lead to a cliff effect, making it impossible to effectively guarantee the accuracy of the source semantic level.
The image is transformed from the pixel domain to the feature domain, and super-prior distribution estimation and quadtree spatial entropy model step-by-step prior estimation are performed. The image is reconstructed using a symbol mapping and denoising enhancement module through a four-way semantic feature encoder and decoder combined with channel state for encoding and decoding.
It improves the encoding and decoding efficiency under harsh channel conditions and the model's ability to understand video content, enhances the ability to reconstruct high-quality images, and controls the scale of model parameters.
Smart Images

Figure CN119383363B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of semantic communication technology, and in particular to a semantic communication method and system for high-resolution images. Background Art
[0002] The amount of data required to transmit for high-resolution image-related services is extremely large, which will bring huge communication overhead. It is necessary to implement a compact coding and transmission scheme to reduce the use of channel bandwidth.
[0003] Traditional image communication systems with a separate design have the following problems when transmitting high-resolution images: in the case of poor channel conditions or limited bandwidth, once there is a mismatch between the communication transmission bit rate and the channel capacity, it will lead to a cliff effect. That is, the communication performance will drop sharply when the channel capacity is lower than the communication transmission bit rate, and the receiving end is often unable to recover the image data; the system is only an independent optimization at the symbol level. In the case of poor channel environment or limited communication bandwidth, the accuracy at the symbol level cannot effectively guarantee the accuracy at the source semantic level.
[0004] Based on this, a new semantic communication method for high-resolution image transmission is urgently needed. Summary of the Invention
[0005] The present application provides a semantic communication method and system for high-resolution images to solve the above problems.
[0006] In a first aspect of the present application, a semantic communication method for high-resolution images is provided, the method comprising:
[0007] Performing a nonlinear transformation on an image to be encoded from a pixel domain to a feature domain to obtain semantic features, wherein the image to be encoded is a video I frame;
[0008] Estimating a super-prior distribution of the semantic features to obtain a super-prior;
[0009] Evenly dividing the semantic features into four groups along the channel dimension to obtain four groups of feature elements;
[0010] Based on the four groups of characteristic elements and the super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior;
[0011] estimating four-step semantic importance based on the channel state, the four groups of feature elements, and the four-step spatial domain prior;
[0012] For the i-th path in the four paths, the combination of feature elements corresponding to the estimated index in the i-th step, the semantic importance in the i-th step, the semantic feature encoding results of the paths before the i-th path, and the channel state are input into the i-th semantic feature encoder, and the i-th semantic feature encoding result is output. The semantic feature encoder is a four-path semantic feature encoder, corresponding to the four-step prior estimation of the spatial entropy model based on the quadtree, and the value range of i is 0-3;
[0013] The semantic feature encoding results of each channel are mapped into symbols of different lengths for wireless transmission through the corresponding symbol mapping modules of each channel;
[0014] The symbol mapping module at the receiving end restores the symbols according to their dimensions;
[0015] For the i-th channel among the four channels, the channel state, the received i-th channel semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th channel are input into the i-th channel semantic feature decoder to obtain the i-th channel semantic feature decoding result, wherein the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature;
[0016] Inputting the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling;
[0017] The noisy features and the channel state are input into a denoising and enhancement module to reconstruct an image.
[0018] In an optional embodiment of the present application, performing a super-prior distribution estimation on the semantic feature to obtain a super-prior includes:
[0019] The semantic features are input into a super-prior estimation module to obtain the super-prior. The super-prior estimation module is only deployed at the transmitting end, and the transmitting end does not transmit the super-prior.
[0020] In an optional embodiment of the present application, when a quadtree-based spatial entropy model is used for step-by-step prior estimation, each group of characteristic elements in each step of the step-by-step prior estimation is different;
[0021] The prior used by the feature element in step 0 is the initial fusion prior, which is generated based on the super prior;
[0022] The feature elements in step 1 will use the priors estimated based on the feature elements in step 0 and the initial fusion priors;
[0023] The feature elements in step 2 will use the priors estimated based on the feature elements in step 0, the feature elements in step 1, and the initial fusion priors;
[0024] The feature elements in step 3 will use the priors estimated based on the feature elements in step 0, step 1, step 2, and the initial fusion priors;
[0025] The prior used for the characteristic element at each step is used to estimate the symbol length corresponding to the encoding of the characteristic element at that step.
[0026] In an optional embodiment of the present application, based on the four groups of characteristic elements and the super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior, including:
[0027] Generate an initial fusion prior based on the super prior and use it as the spatial domain estimation result of the spatial domain estimation in step 0;
[0028] In the process of spatial domain estimation from steps 1 to 3, the estimated distribution features obtained in the previous steps and the initial fusion prior are used as references, and the estimation is completed using a network with shared weights to obtain the spatial domain prior for each step.
[0029] In an optional embodiment of the present application, each semantic feature encoder includes: a channel perception module and a semantic feature encoding module, and the semantic feature encoding module includes: a first cross attention layer and a second cross attention layer;
[0030] The combination of feature elements corresponding to the estimated index of the i-th step, the semantic importance of the i-th step, the semantic feature encoding results of each path before the i-th path, and the channel state are input into the i-th semantic feature encoder, and the i-th semantic feature encoding result is output, including:
[0031] Inputting the channel state and the characteristic element corresponding to the i-th estimation index into the channel sensing module to obtain a first channel sensing result;
[0032] Inputting the first channel perception result and the i-th step semantic importance into the first cross attention layer to obtain a first cross attention result;
[0033] Based on the first cross-attention result and the semantic feature encoding results of each path before the i-th path, the second cross-attention layer is used to obtain the i-th semantic feature encoding result.
[0034] In an optional embodiment of the present application, the channel perception module includes a spatial attention module and a channel attention module;
[0035] Inputting the channel state and the characteristic element corresponding to the i-th estimation index into the channel sensing module to obtain a first channel sensing result, including:
[0036] Input the channel state and the characteristic element corresponding to the i-th step estimation index into the channel attention module to obtain a channel attention result;
[0037] Based on the feature element combination corresponding to the estimated index in the i-th step and the channel attention result, the spatial attention result is obtained using the spatial attention module;
[0038] Based on the feature element combination corresponding to the i-th step estimation index and the spatial attention result, the first channel perception result is obtained.
[0039] In an optional embodiment of the present application, each semantic feature decoder includes: a channel perception module and a semantic feature decoding module, and the semantic feature decoding module includes: a third cross attention layer;
[0040] Inputting the channel state, the received i-th semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th semantic feature decoder to obtain the i-th semantic feature decoding result, including:
[0041] Based on the received i-th semantic feature encoding result and the semantic feature decoding results before the i-th path, using the third cross attention layer to obtain a second cross attention result;
[0042] Based on the channel state and the second cross-attention result, the channel perception module is used to obtain the i-th semantic feature decoding result.
[0043] In an optional embodiment of the present application, the semantic information restorer includes: a first semantic information restoration module, a channel perception module, and a second semantic information restoration module;
[0044] Inputting the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling, comprising:
[0045] Inputting the received semantic features into the first semantic information recovery module to obtain a semantic information recovery result;
[0046] Inputting the semantic information recovery result and the channel state into the channel sensing module to obtain a second channel sensing result;
[0047] The second channel perception result is input into the second semantic information recovery module to obtain a noisy feature.
[0048] In an optional embodiment of the present application, the denoising enhancement module includes: multiple Unet2 layers, Conv2d layers and a channel perception module;
[0049] Inputting the noisy features and the channel state into a denoising and enhancement module to reconstruct an image includes:
[0050] Input the noisy features into the first Unet2 layer in the denoising enhancement module to obtain the noise features;
[0051] Based on the noise characteristics and channel state, multiple Unet2 layers, Conv2d layers and channel perception modules in the denoising enhancement module are used to obtain the noise characteristics based on the channel state;
[0052] An image is reconstructed based on the noisy feature and the noise feature based on the channel state.
[0053] In a second aspect of the present application, a semantic communication system for high-resolution images is provided, the system comprising:
[0054] A semantic feature extraction module is used to perform a nonlinear transformation from a pixel domain to a feature domain on an image to be encoded to obtain semantic features, wherein the image to be encoded is a video I frame;
[0055] A super prior distribution estimation module, configured to perform super prior distribution estimation on the semantic features to obtain a super prior;
[0056] A feature segmentation module is used to evenly segment the semantic features into four groups along the channel dimension to obtain four groups of feature elements;
[0057] a distribution prior estimation module, configured to perform step-by-step prior estimation based on the four groups of characteristic elements and the super prior using a quadtree-based spatial entropy model to obtain a four-step spatial prior;
[0058] a semantic importance estimation module, configured to estimate the four-step semantic importance based on the channel state, the four groups of feature elements, and the four-step spatial domain prior;
[0059] A semantic feature encoding module is configured to input, for the i-th path among the four paths, a combination of feature elements corresponding to the i-th step estimated index, the i-th step semantic importance, the semantic feature encoding results of the paths before the i-th path, and the channel state, into the i-th path semantic feature encoder, and output the i-th path semantic feature encoding result. The semantic feature encoder is a four-path semantic feature encoder corresponding to the four-step prior estimation of the quadtree-based spatial entropy model, and the value range of i is 0-3;
[0060] The transmission module is used to map the semantic feature encoding results of each channel into symbols of different lengths for wireless transmission through the corresponding symbol mapping modules of each channel;
[0061] The restoration module is used for the symbol mapping module of each channel at the receiving end to restore the symbols according to their dimensions;
[0062] A semantic feature decoding module is configured to input the channel state, the received semantic feature encoding result of the i-th channel, and the semantic feature decoding results of each channel before the i-th channel into the i-th semantic feature decoder for the i-th channel, to obtain the i-th semantic feature decoding result, wherein the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature;
[0063] A semantic recovery module, configured to input the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling;
[0064] The reconstruction module is used to input the noisy features and the channel state into the denoising and enhancement module to reconstruct the image.
[0065] The present application has the following advantages: a semantic communication method and system for high-resolution images are provided in the present application, the image to be encoded is transformed from the pixel domain to the feature domain through nonlinear transformation to obtain semantic features; the semantic features are estimated by super-prior distribution to obtain super-prior; the semantic features are evenly divided into four groups along the channel dimension to obtain four groups of feature elements; based on the four groups of feature elements and the super-prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior; based on the channel state, the four groups of feature elements and the four-step spatial prior, the four-step semantic importance is estimated; for the i-th path in the four paths, the feature element combination corresponding to the i-th step estimation index, the i-th step semantic importance, the semantic feature encoding results of each path before the i-th path and the channel state are input into the i-th path semantic feature encoder, and the i-th path semantic feature encoding result is output, and the semantic feature encoder is a four-path semantic feature encoder. , corresponding to the four-step prior estimation of the spatial entropy model based on the quadtree, the value range of i is 0-3; the semantic feature encoding result of each channel is mapped into symbols of different lengths for wireless transmission through the corresponding symbol mapping module of each channel; the symbol mapping module of each channel at the receiving end restores the symbol according to the dimension of the symbol; for the i-th channel among the four channels, the channel state, the received i-th semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th channel are input into the i-th semantic feature decoder to obtain the i-th semantic feature decoding result, the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature; the received semantic feature and the channel state are input into the semantic information restorer, and the semantic information restorer is used to restore the received semantic feature to a noisy feature after upsampling; the noisy feature and the channel state are input into the denoising enhancement module to reconstruct the image. The module built using CNN extracts and restores the semantic features of video frames, and the codec module processes the semantic features of video frames. This not only improves the encoding and decoding efficiency, but also enhances the model's ability to understand video content, helps to better reconstruct high-quality video content, and controls the size of model parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0067] Figure 1 This is a flowchart of the steps of a semantic communication method for high-resolution images provided in an embodiment of the present application;
[0068] Figure 2This is a flowchart of a semantic communication method for high-resolution images provided in an embodiment of the present application;
[0069] Figure 3 This is a network structure diagram of a channel perception module in a semantic communication method for high-resolution images provided in an embodiment of the present application;
[0070] Figure 4 This is a network structure diagram of a symbol mapping module in a semantic communication method for high-resolution images provided in an embodiment of the present application;
[0071] Figure 5 This is an architectural diagram of a semantic communication system for high-resolution images provided in an embodiment of the present application. DETAILED DESCRIPTION
[0072] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0073] In order to solve the following problems caused by the traditional image communication system with a separate design when transmitting high-resolution images: under poor channel conditions or limited bandwidth, once a mismatch occurs between the communication transmission code rate and the channel capacity, it will lead to a cliff effect, that is, the communication performance will drop sharply when the channel capacity is lower than the communication transmission code rate, and the receiving end is often unable to recover the image data; the system is only an independent optimization at the symbol level. Under poor channel conditions or limited communication bandwidth, the accuracy at the symbol level is difficult to effectively guarantee the accuracy at the semantic level of the source. This application proposes a semantic communication method for high-resolution images to solve the above problems. The image frames in the high-resolution images transmitted in this application are I frames of video image frames, and any I frame can be randomly decoded to play the video stream.
[0074] See Figure 1 , Figure 1 This is a flowchart of a method for semantic communication of high-resolution images provided in an embodiment of the present application, the method comprising the following steps:
[0075] Step 101: performing a nonlinear transformation on an image to be encoded from a pixel domain to a feature domain to obtain semantic features, wherein the image to be encoded is a video I frame;
[0076] Step 102: Estimating the hyper-prior distribution of the semantic features to obtain a hyper-prior;
[0077] Step 103: Evenly divide the semantic features into four groups along the channel dimension to obtain four groups of feature elements;
[0078] Step 104: Based on the four groups of characteristic elements and the super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior.
[0079] Step 105: estimating the four-step semantic importance based on the channel state, the four groups of feature elements and the four-step spatial domain prior;
[0080] Step 106: For the i-th path among the four paths, the feature element combination corresponding to the estimated index in the i-th step, the semantic importance in the i-th step, the semantic feature encoding results of the paths before the i-th path, and the channel state are input into the i-th path semantic feature encoder, and the i-th path semantic feature encoding result is output. The semantic feature encoder is a four-path semantic feature encoder corresponding to the four-step prior estimation of the spatial entropy model based on the quadtree, and the value range of i is 0-3;
[0081] Step 107: Mapping the semantic feature encoding results of each channel into symbols of different lengths for wireless transmission through the corresponding symbol mapping modules of each channel;
[0082] Step 108: The symbol mapping module of each channel at the receiving end restores the symbol according to the symbol dimension;
[0083] Step 109: For the i-th channel among the four channels, input the channel state, the received i-th channel semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th channel into the i-th channel semantic feature decoder to obtain the i-th channel semantic feature decoding result, wherein the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature;
[0084] Step 110: inputting the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features to noisy features through upsampling;
[0085] Step 111: Input the noisy features and the channel state into a denoising and enhancement module to reconstruct an image.
[0086] In order to clearly illustrate the semantic communication method of high-resolution images proposed in this application, Figure 2 To explain, Figure 2 This is a flow chart of a semantic communication method for high-resolution images provided in an embodiment of the present application.
[0087] When specifically implementing step 101, first obtain any I frame in the ultra-high-definition video to be transmitted as a high-resolution image to be encoded, wherein the high-resolution image is an image with a large number of pixels (Pixel) or dots (Dot) in the digital image, and these pixels or dots together constitute the details and clarity of the image. High-resolution images can usually provide finer details, clearer edges and less blur, making the objects, scenes or texts in the image more realistic and easy to identify. The image to be encoded is transformed nonlinearly from the pixel domain to the feature domain to obtain semantic features. The image semantic feature extractor includes a first downsampling module, a channel perception module, a second downsampling module and a Conv2d layer, which can transform the image into an implicit feature space through nonlinear transformation. Each downsampling module contains a downsampling layer and a DepthConvBlock2 layer. Specifically, the acquired image to be encoded Input the first downsampling module to obtain the first downsampling result, input the first downsampling result and the channel state (such as SNR) into the channel perception module to obtain the channel perception result, and based on the channel perception result, use the second downsampling module and the Conv2d layer to obtain the semantic feature .
[0088] When implementing step 102, in order to compact the semantic features, this application introduces a super-prior to perform a preliminary estimate of the potential distribution of the semantic features. The super-prior distribution estimation is performed using a single neural network model. The super-prior distribution estimation is performed on the above-mentioned semantic features to obtain the super-prior. Specifically, the above-mentioned semantic features are input into the super-prior estimation module to obtain the above-mentioned super-prior. The super-prior estimation module is deployed only at the transmitting end, and the transmitting end does not transmit the super-prior.
[0089] In an optional embodiment of the present application, the super prior estimation module includes two Conv2d layers and one Leaky Relu layer. The super prior estimation module is used to estimate the super prior distribution of the semantic features to obtain the super prior. Specifically, the semantic features Input the hyper prior estimation module to obtain the hyper prior. The above hyper prior estimation module is only deployed at the transmitting end, and the transmitting end will not transmit the hyper prior distribution. and letters The meanings represented are all semantic features, but the letters representing the semantic features are different in different processing scenarios.
[0090] When implementing step 103, a quadtree entropy model is introduced in this application to perform subsequent spatial domain prior distribution estimation and semantic importance estimation. First, the extracted semantic features are evenly divided into four groups along the channel dimension to obtain four groups of feature elements. The four groups of feature elements are used for subsequent spatial domain prior distribution estimation and semantic importance estimation. The channel dimension of the semantic features is channel, which can be abbreviated as C.
[0091] Specifically, when performing step 104, a quadtree-based spatial entropy model is used to estimate the spatial prior distribution and semantic importance. The quadtree-based spatial entropy model can fully identify the redundancy of spatial information in semantic features. First, based on four groups of feature elements and the above-mentioned super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior. According to the step-by-step estimation strategy of the entropy model, a four-way semantic feature coding transmission link with mutual reference is realized; the semantic importance is estimated based on the output result of the entropy model to guide the feature encoder to fully reduce the channel bandwidth occupancy and realize unequal error protection transmission.
[0092] In an optional embodiment of the present application, when a quadtree-based spatial entropy model is used for step-by-step prior estimation, each group of feature elements in each step-by-step prior estimation is different; specifically, the prior used for the feature element of step 0 is the initial fusion prior, which is generated based on the super prior; the prior used for the feature element of step 1 is estimated based on the feature element of step 0 and the initial fusion prior; the prior used for the feature element of step 2 is estimated based on the feature element of step 0, the feature element of step 1, and the initial fusion prior; the prior used for the feature element of step 3 is estimated based on the feature element of step 0, the feature element of step 1, the feature element of step 2, and the initial fusion prior; since the prior used for the feature element of each step is different, the feature element of each step is also different. The prior used for the feature element of each step is used to estimate the symbol length corresponding to the feature element of that step after encoding.
[0093] In an optional embodiment of the present application, based on the four groups of feature elements and the super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation, and the process of obtaining a four-step spatial prior specifically includes: first, generating an initial fusion prior based on the above super prior, and using the initial fusion prior as the spatial estimation result of the 0th step spatial estimation; then, in the process of spatial estimation in steps 1 to 3, the estimated distribution characteristics obtained in the previous steps and the initial fusion prior are used as references, and a shared weight network is used to complete the estimation to obtain the spatial prior for each step. The network structure of the quadtree-based spatial entropy model used in this application is roughly the same as that of the prior art, but the output quantization step length is removed.
[0094] In an optional embodiment of the present application, the spatial domain prior distribution estimation is divided into four steps, which can be found in Figure 2 , along the channel dimension C, the dimension is The semantic features are evenly divided into four groups of dimensions: Each step will estimate the feature element of the corresponding index. C represents the channel dimension, H represents the height dimension, and W represents the width dimension. For example, in step 0 ( Figure 2 In Step 0) each group will Figure 2 The position with index 0 in the [H, W] matrix is estimated. Each group has a different index, but the combined indexes of the groups cover the dimension [H, W]. The indexes at each estimation step are also different, and the combination of feature elements used in each estimation step is the combination of all feature elements corresponding to the estimated index at each step. Therefore, each step of spatial prior distribution estimation covers every spatial position in the dimension [H, W], and the estimated feature elements can form a feature with dimensions [C / 4, H, W]. In subsequent steps 1, 2, and 3, the spatial prior distribution estimation uses the feature elements estimated in the previous steps as a reference. This allows the estimated position within the same group at each step to refer to its neighbors in the same group and the same position in different groups, thereby guiding the model to fully identify spatial information redundancy in the semantic features. The entropy model output parameters are also used to guide the semantic feature encoder to implement unequal error protection. The semantic importance estimation module uses the prior distribution parameters estimated by the entropy model and the semantic features to calculate the semantic importance of each feature element. The output of this module interacts with the semantic feature encoding process through the cross-attention layer in the semantic feature encoder.
[0095] Specifically, when performing semantic importance estimation in step 105 , the channel state, the four groups of feature elements, and the four-step spatial domain priors are input into a semantic importance estimation module to estimate the four-step semantic importance.
[0096] In an optional embodiment of the present application, the network structure of the semantic importance module used in each step of semantic importance estimation is the same. The following is an exemplary description of the semantic importance estimation of the i-th step. The semantic importance estimation module includes multiple DepthConvBlock layers, multiple ResBlock layers and a channel perception module (AF Module), which converts the above i-th group of feature elements into The four groups of feature elements are obtained by segmenting the semantic features, and the alphabetical interpretations of the semantic features are still used here to represent the feature elements) and the i-th step spatial domain prior (Etropy prior) are input into the first DepthConvBlock layer to obtain the first processing result; based on the first processing result and the channel state (such as SNR), the channel sensing module is used to obtain the channel sensing processing result; based on the channel sensing processing result, the second ResBlock layer is used to obtain the i-th step semantic importance, where the value of i ranges from 0 to 3.
[0097] Specifically, when step 106 is implemented, for the i-th path among the four paths, the feature element combination corresponding to the i-th step estimation index, the i-th step semantic importance, the semantic feature encoding results of the paths before the i-th path, and the channel state are input into the i-th path semantic feature encoder, and the i-th path semantic feature encoding result is output. The above-mentioned semantic feature encoder is a four-path semantic feature encoder, corresponding to the four-step prior estimation of the spatial entropy model based on the quadtree, and the value range of i is 0-3;
[0098] The feature spatial domain composition dimension estimated according to the spatial domain prior distribution of each step is: The feature elements of the proposed method are used to construct a four-way semantic feature encoder, which encodes the features composed of the spatial prior distribution estimation in each step respectively to obtain the four-step semantic feature encoding results. The self-attention layer based on Swin-Transformer and the cross-attention layer based on Swin-Transformer-V2 are used to construct a neural network model to remove the information redundancy of the semantic features identified in the spatial distribution estimation.
[0099] The network structure of each semantic feature encoder is the same. Each semantic feature encoder includes: a channel perception module and a semantic feature encoding module. The above semantic feature encoding module includes: a first cross-attention layer and a second cross-attention layer, and also includes a first self-attention layer, a second self-attention layer and a DepthConvBlock layer.
[0100] The characteristic element combination corresponding to the estimated index of step i, the semantic importance of step i, the semantic feature encoding results of each path before the i-th path and the channel state are input into the i-th semantic feature encoder, and the i-th semantic feature encoding result is output, specifically including: the channel state (such as SNR) and the characteristic element combination corresponding to the estimated index of step i Input the above-mentioned channel perception module AF module to obtain the first channel perception result; input the above-mentioned first channel perception result and the semantic importance of the i-th step into the above-mentioned first cross attention layer to obtain the first cross attention result; based on the above-mentioned first cross attention result and the semantic feature encoding results of each path before the i-th path , using the second cross attention layer, the i-th semantic feature encoding result is obtained .
[0101] In an optional embodiment of the present application, the network structure of the channel perception module AF Module in the semantic feature encoder can be referred to Figure 3 , Figure 3 This is a network structure diagram of a channel perception module in a semantic communication method for high-resolution images provided in an embodiment of the present application. The channel perception module includes a spatial attention module and a channel attention module. The channel state and the characteristic element combination corresponding to the i-th step estimation index are input into the channel perception module to obtain a first channel perception result, specifically including: the channel state (e.g., SNR) and the characteristic element combination corresponding to the i-th step estimation index are input into the channel perception module. Input the channel attention module to obtain the channel attention result; based on the feature element combination corresponding to the estimated index in step i above And the above channel attention results, use the above spatial attention module to get the spatial attention results; based on the feature element combination corresponding to the estimated index in the above step i Combined with the spatial attention results, we obtain the first channel perception result. By integrating the channel perception module, we can adaptively adjust the encoding and decoding strategy based on real-time channel conditions (channel state), and use the attention mechanism to enhance important features of the video data, thereby improving overall transmission efficiency and quality.
[0102] When specifically implementing steps 107 and 108, after obtaining the i-th semantic feature encoding result, each semantic feature encoding result is mapped into symbols of different lengths by the corresponding symbol mapping module for wireless transmission; then, each symbol mapping module at the receiving end restores the symbol according to the symbol dimension and transmits the restored encoding result. Figure 4 , Figure 4 This is a network structure diagram of the symbol mapping module in the semantic communication method of a high-resolution image provided by the embodiment of the present application. The dimension of the output result after the semantic feature passes through the feature encoder is , for the dimension Each element of has a length of , according to the hyper-prior distribution estimated by the entropy model, the symbol length required for the element is calculated, and a linear layer matching the symbol length is selected for mapping. At the receiving end, the corresponding linear layer will be used to map it back to length , thereby recovering the dimension from the received symbols The coded transmission result is then transmitted to the semantic feature decoder. Based on the symbol mapping module, a semantic communication method for high-resolution images is implemented that dynamically adjusts the video codec output bit rate according to the channel status and semantic information.
[0103] When step 109 is specifically implemented, for the i-th path among the four paths, the channel state, the received i-th semantic feature encoding result, and the semantic feature decoding results of each path before the i-th path are input into the i-th semantic feature decoder to obtain the i-th semantic feature decoding result. The semantic feature decoder is a four-path semantic feature decoder.
[0104] In an optional embodiment of the present application, each semantic feature decoder includes: a channel perception module and a semantic feature decoding module, and the semantic feature decoding module includes: a third cross-attention layer, and also includes: a third self-attention layer, a fourth self-attention layer and a DepthConvBlock layer.
[0105] The channel state, the received i-th semantic feature encoding result, and the i-th semantic feature decoding results are input into the i-th semantic feature decoder to obtain the i-th semantic feature decoding result, specifically including: based on the i-th semantic feature encoding result received , the decoding results of the semantic features before the above i-th path , using the third cross attention layer, obtain the second cross attention result; based on the channel state (such as SNR) and the second cross attention result, use the channel perception module AF Module to obtain the i-th semantic feature decoding result , four-way semantic feature decoding results Composition of received semantic features .
[0106] Specifically, when step 110 is implemented, the received semantic features and the channel state are input into a semantic information restorer, and the semantic information restorer is used to restore the received semantic features into noisy features through upsampling.
[0107] In an optional embodiment of the present application, the above-mentioned semantic information recoverer includes: a first semantic information recovery module, a channel perception module and a second semantic information recovery module; the first semantic information recovery module and the second semantic information recovery module both include a DepthConvBlock2 layer and an upsampling layer.
[0108] The received semantic features and the channel state are input into a semantic information restorer, which is used to restore the received semantic features to noisy features through upsampling, specifically including: Input the above-mentioned first semantic information recovery module to obtain the semantic information recovery result; input the above-mentioned semantic information recovery result and the above-mentioned channel state (such as SNR) into the channel perception module to obtain the second channel perception result; input the above-mentioned second channel perception result into the above-mentioned second semantic information recovery module to obtain the noisy feature .
[0109] Specifically, in order to further improve the image reconstruction capability at the receiving end, a feature denoising and enhancement module is constructed to further remove the interference of noise on image reconstruction. The noisy features and the channel state are fed into the denoising and enhancement module to reconstruct a high-quality image.
[0110] In an optional embodiment of the present application, the denoising enhancement module includes: multiple Unet2 layers, Conv2d layers and a channel perception module. Specifically, the noise-containing features are Input the first Unet2 layer in the I-frame denoising enhancement module to obtain the noise characteristics ; Based on noise characteristics and channel status (such as SNR), using multiple Unet2 layers, Conv2d layers and channel perception module AF Module in the I frame denoising enhancement module to obtain noise characteristics based on the channel status; based on the noise characteristics and the noise characteristics based on the channel state to reconstruct the image .
[0111] In an embodiment of the present application, a semantic communication method for high-resolution images is provided, wherein an image to be encoded is nonlinearly transformed from a pixel domain to a feature domain to obtain semantic features, wherein the image to be encoded is a video I frame; a super-prior distribution estimation is performed on the semantic features to obtain a super-prior; the semantic features are evenly divided into four groups along the channel dimension to obtain four groups of feature elements; based on the four groups of feature elements and the super-prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior; based on the channel state, the four groups of feature elements and the four-step spatial prior, the four-step semantic importance is estimated; for the i-th path in the four paths, the feature element combination corresponding to the i-th step estimation index, the i-th step semantic importance, the semantic feature encoding results of each path before the i-th path and the channel state are input into the i-th path semantic feature encoder, and the i-th path semantic feature encoding result is output, and the semantic feature encoder is a four-path semantic feature encoder. The device corresponds to the four-step prior estimation of the spatial entropy model based on the quadtree, and the value range of i is 0-3; the semantic feature encoding result of each channel is mapped into symbols of different lengths for wireless transmission through the corresponding symbol mapping module of each channel; the symbol mapping module of each channel at the receiving end restores the symbol according to the dimension of the symbol; for the i-th channel among the four channels, the channel state, the received i-th semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th channel are input into the i-th semantic feature decoder to obtain the i-th semantic feature decoding result, the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature; the received semantic feature and the channel state are input into the semantic information restorer, and the semantic information restorer is used to restore the received semantic feature to a noisy feature after upsampling; the noisy feature and the channel state are input into the denoising enhancement module to reconstruct the image. The module built using CNN extracts and restores the semantic features of high-resolution image frames, and the codec module processes the semantic features of the image. This not only improves the coding and decoding efficiency, but also enhances the model's ability to understand the image content, helps to better reconstruct high-quality image frequency content, and control the size of model parameters.
[0112] In a second aspect of the present application, a semantic communication system for high-resolution images is provided. Figure 5 , Figure 5 : This is an architecture diagram of a semantic communication system for high-resolution images provided in an embodiment of the present application. The system includes:
[0113] Semantic feature extraction module 501, used for performing nonlinear transformation from pixel domain to feature domain on the image to be encoded to obtain semantic features, wherein the image to be encoded is a video I frame;
[0114] A super prior distribution estimation module 502 is used to estimate the super prior distribution of the semantic features to obtain a super prior;
[0115] A feature segmentation module 503 is configured to evenly segment the semantic features into four groups along the channel dimension to obtain four groups of feature elements;
[0116] a distribution prior estimation module 504 for performing step-by-step prior estimation based on the four groups of characteristic elements and the super prior using a quadtree-based spatial entropy model to obtain a four-step spatial prior;
[0117] A semantic importance estimation module 505 is configured to estimate four-step semantic importance based on the channel state, the four groups of feature elements, and the four-step spatial domain prior;
[0118] A semantic feature encoding module 506 is configured to input, for the i-th path among the four paths, a combination of feature elements corresponding to the estimated index in the i-th step, the semantic importance in the i-th step, the semantic feature encoding results of the paths before the i-th path, and the channel state, into an i-th path semantic feature encoder, and output the i-th path semantic feature encoding result. The semantic feature encoder is a four-path semantic feature encoder corresponding to the four-step prior estimation of the quadtree-based spatial entropy model, and the value of i ranges from 0 to 3.
[0119] The transmission module 507 is used to map the semantic feature encoding results of each channel into symbols of different lengths for wireless transmission through the corresponding symbol mapping modules of each channel;
[0120] Restoration module 508, used for each symbol mapping module of the receiving end, restores the symbol according to the symbol dimension;
[0121] The semantic feature decoding module 509 is used to input the channel state, the received semantic feature encoding result of the i-th channel, and the semantic feature decoding results of each channel before the i-th channel into the i-th semantic feature decoder to obtain the i-th semantic feature decoding result. The semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature;
[0122] A semantic recovery module 510 is configured to input the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling;
[0123] The reconstruction module 511 is configured to input the noisy features and the channel state into a denoising and enhancement module to reconstruct an image.
[0124] Wherein, the super prior distribution estimation module includes:
[0125] The distribution estimation submodule is used to input the semantic features into the super-prior estimation module to obtain the super-prior. The super-prior estimation module is only deployed at the transmitting end, and the transmitting end will not transmit the super-prior.
[0126] Wherein, each group of characteristic elements in each step of the prior estimation of the distribution is different; the prior estimation module of the distribution includes:
[0127] A priori generation submodule for step 0, wherein the priori used for the feature elements of step 0 is an initial fusion priori, and the initial fusion priori is generated based on the super priori;
[0128] The priori generation submodule of step 1 uses the priori estimated based on the feature elements of step 0 and the initial fusion priori for the feature elements of step 1.
[0129] The priori generation submodule of step 2 uses the priori estimated based on the feature elements of step 0, the feature elements of step 1, and the initial fusion priori for the feature elements of step 2;
[0130] The priori generation submodule in step 3 uses the priori estimated based on the feature elements in step 0, step 1, step 2, and the initial fusion priori for the feature elements in step 3;
[0131] The symbol estimation submodule is used to estimate the symbol length corresponding to the encoded characteristic element of each step using the prior information used in the characteristic element of each step.
[0132] Wherein, the distribution prior estimation module includes:
[0133] An initial estimation module, configured to generate an initial fusion prior based on the super prior and use it as a spatial domain estimation result of the spatial domain estimation in step 0;
[0134] The submodule for estimating the spatial domain prior at each step is used to utilize the estimated distribution features obtained in the previous steps and the initial fusion prior as references during the spatial domain estimation process from the first to the third step, and to complete the estimation using a network with shared weights to obtain the spatial domain prior at each step.
[0135] Each semantic feature encoder in the semantic feature encoding module includes: a channel perception module and a semantic feature encoding module, the semantic feature encoding module includes: a first cross attention layer and a second cross attention layer, and the semantic feature encoding module includes:
[0136] A first channel sensing submodule, configured to input the channel state and the characteristic element combination corresponding to the i-th estimation index into the channel sensing module to obtain a first channel sensing result;
[0137] A first cross attention submodule, configured to input the first channel perception result and the i-th step semantic importance into the first cross attention layer to obtain a first cross attention result;
[0138] The second cross-attention sub-module is used to obtain the semantic feature encoding result of the i-th path based on the first cross-attention result and the semantic feature encoding results of each path before the i-th path using the second cross-attention layer.
[0139] The channel perception module in the first channel perception submodule includes a spatial attention module and a channel attention module; the first channel perception submodule includes:
[0140] A channel attention unit is used to input the channel state and the feature element corresponding to the i-th step estimation index into the channel attention module to obtain a channel attention result;
[0141] A spatial attention unit, configured to obtain a spatial attention result using the spatial attention module based on the feature element combination corresponding to the estimated index in the i-th step and the channel attention result;
[0142] A combining unit is used to obtain the first channel perception result based on the combination of feature elements corresponding to the i-th step estimation index and the spatial attention result.
[0143] Each semantic feature decoder in the semantic feature decoding module includes: a channel perception module and a semantic feature decoding module, and the semantic feature decoding module includes: a third cross attention layer; the semantic feature decoding module includes:
[0144] A first decoding submodule is configured to obtain a second cross attention result by using the third cross attention layer based on the received i-th semantic feature encoding result and the semantic feature decoding results before the i-th path;
[0145] The second decoding submodule is used to obtain the i-th semantic feature decoding result based on the channel state and the second cross-attention result using the channel perception module.
[0146] The semantic information restorer in the semantic restoration module includes: a first semantic information restoration module, a channel perception module, and a second semantic information restoration module; the semantic restoration module includes:
[0147] a first recovery submodule, configured to input the received semantic features into the first semantic information recovery module to obtain a semantic information recovery result;
[0148] a second recovery submodule, configured to input the semantic information recovery result and the channel state into the channel sensing module to obtain a second channel sensing result;
[0149] The third recovery submodule is used to input the second channel perception result into the second semantic information recovery module to obtain a noisy feature.
[0150] The denoising enhancement module in the reconstruction module includes: multiple Unet2 layers, Conv2d layers and a channel perception module; the reconstruction module includes:
[0151] A first denoising submodule is configured to input the noisy features into the first Unet2 layer in the denoising enhancement module to obtain noise features;
[0152] The second denoising submodule is used to obtain noise features based on the channel state based on the noise features and channel state by using multiple Unet2 layers, Conv2d layers and channel perception modules in the denoising enhancement module;
[0153] The reconstruction submodule is configured to reconstruct an image based on the noisy feature and the noise feature based on the channel state.
[0154] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0155] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0156] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0158] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0159] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0160] The above is a detailed introduction to the semantic communication method and system for high-resolution images provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A semantic communication method for high-resolution images, characterized in that: The method comprises: Performing a nonlinear transformation on an image to be encoded from a pixel domain to a feature domain to obtain semantic features, wherein the image to be encoded is a video I frame; Estimating a super-prior distribution of the semantic features to obtain a super-prior; Evenly dividing the semantic features into four groups along the channel dimension to obtain four groups of feature elements; Based on the four groups of characteristic elements and the super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior; estimating four-step semantic importance based on the channel state, the four groups of feature elements, and the four-step spatial domain prior; For the i-th path in the four paths, the combination of feature elements corresponding to the estimated index in the i-th step, the semantic importance in the i-th step, the semantic feature encoding results of the paths before the i-th path, and the channel state are input into the i-th semantic feature encoder, and the i-th semantic feature encoding result is output. The semantic feature encoder is a four-path semantic feature encoder, corresponding to the four-step prior estimation of the spatial entropy model based on the quadtree, and the value range of i is 0-3; The semantic feature encoding results of each channel are mapped into symbols of different lengths for wireless transmission through the corresponding symbol mapping modules of each channel; The symbol mapping module at the receiving end restores the symbols according to their dimensions; For the i-th channel among the four channels, the channel state, the received i-th channel semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th channel are input into the i-th channel semantic feature decoder to obtain the i-th channel semantic feature decoding result, wherein the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature; Inputting the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling; The noisy features and the channel state are input into a denoising and enhancement module to reconstruct an image.
2. The semantic communication method of high-resolution images according to claim 1, characterized in that: Estimating the super-prior distribution of the semantic features to obtain the super-prior includes: The semantic features are input into a super-prior estimation module to obtain the super-prior. The super-prior estimation module is only deployed at the transmitting end, and the transmitting end does not transmit the super-prior.
3. The semantic communication method of high-resolution images according to claim 1, characterized in that: When the quadtree-based spatial entropy model is used for step-by-step prior estimation, each set of characteristic elements in each step-by-step prior estimation is different; The prior used by the feature element in step 0 is the initial fusion prior, which is generated based on the super prior; The feature elements in step 1 will use the priors estimated based on the feature elements in step 0 and the initial fusion priors; The feature elements in step 2 will use the priors estimated based on the feature elements in step 0, the feature elements in step 1, and the initial fusion priors; The feature elements in step 3 will use the priors estimated based on the feature elements in step 0, step 1, step 2, and the initial fusion priors; The prior used for the characteristic element at each step is used to estimate the symbol length corresponding to the encoding of the characteristic element at that step.
4. The semantic communication method of high-resolution images according to claim 1, characterized in that: Based on the four sets of characteristic elements and the super prior, a quadtree-based spatial entropy model is used to perform step-by-step prior estimation to obtain a four-step spatial prior, including: Generate an initial fusion prior based on the super prior and use it as the spatial domain estimation result of the spatial domain estimation in step 0; In the process of spatial domain estimation from steps 1 to 3, the estimated distribution features obtained in the previous steps and the initial fusion prior are used as references, and the estimation is completed using a network with shared weights to obtain the spatial domain prior for each step.
5. The semantic communication method of high-resolution images according to claim 1, characterized in that: Each semantic feature encoder includes: a channel perception module and a semantic feature encoding module, and the semantic feature encoding module includes: a first cross attention layer and a second cross attention layer; The combination of feature elements corresponding to the estimated index of the i-th step, the semantic importance of the i-th step, the semantic feature encoding results of each path before the i-th path, and the channel state are input into the i-th semantic feature encoder, and the i-th semantic feature encoding result is output, including: Inputting the channel state and the characteristic element corresponding to the i-th estimation index into the channel sensing module to obtain a first channel sensing result; Inputting the first channel perception result and the i-th step semantic importance into the first cross attention layer to obtain a first cross attention result; Based on the first cross-attention result and the semantic feature encoding results of each path before the i-th path, the second cross-attention layer is used to obtain the i-th semantic feature encoding result.
6. The semantic communication method of high-resolution images according to claim 5, characterized in that: The channel perception module includes a spatial attention module and a channel attention module; Inputting the channel state and the characteristic element corresponding to the i-th estimation index into the channel sensing module to obtain a first channel sensing result, including: Input the channel state and the characteristic element corresponding to the i-th step estimation index into the channel attention module to obtain a channel attention result; Based on the feature element combination corresponding to the estimated index in the i-th step and the channel attention result, the spatial attention result is obtained using the spatial attention module; Based on the feature element combination corresponding to the i-th step estimation index and the spatial attention result, the first channel perception result is obtained.
7. The semantic communication method of high-resolution images according to claim 1, characterized in that: Each semantic feature decoder includes: a channel perception module and a semantic feature decoding module, and the semantic feature decoding module includes: a third cross attention layer; Inputting the channel state, the received i-th semantic feature encoding result, and the semantic feature decoding results of each channel before the i-th semantic feature decoder to obtain the i-th semantic feature decoding result, including: Based on the received i-th semantic feature encoding result and the semantic feature decoding results before the i-th path, using the third cross attention layer to obtain a second cross attention result; Based on the channel state and the second cross-attention result, the channel perception module is used to obtain the i-th semantic feature decoding result.
8. The semantic communication method of high-resolution images according to claim 1, characterized in that: The semantic information restorer includes: a first semantic information restoration module, a channel perception module and a second semantic information restoration module; Inputting the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling, including: Inputting the received semantic features into the first semantic information recovery module to obtain a semantic information recovery result; Inputting the semantic information recovery result and the channel state into the channel sensing module to obtain a second channel sensing result; The second channel perception result is input into the second semantic information recovery module to obtain a noisy feature.
9. The semantic communication method of high-resolution images according to claim 1, characterized in that: The denoising enhancement module includes: multiple Unet2 layers, Conv2d layers and a channel perception module; Inputting the noisy features and the channel state into a denoising and enhancement module to reconstruct an image includes: Input the noisy features into the first Unet2 layer in the denoising enhancement module to obtain the noise features; Based on the noise characteristics and channel state, multiple Unet2 layers, Conv2d layers and channel perception modules in the denoising enhancement module are used to obtain the noise characteristics based on the channel state; An image is reconstructed based on the noisy feature and the noise feature based on the channel state.
10. A semantic communication system for high-resolution images, characterized in that: The system comprises: A semantic feature extraction module is used to perform a nonlinear transformation from a pixel domain to a feature domain on an image to be encoded to obtain semantic features, wherein the image to be encoded is a video I frame; A super prior distribution estimation module, configured to perform super prior distribution estimation on the semantic features to obtain a super prior; A feature segmentation module is used to evenly segment the semantic features into four groups along the channel dimension to obtain four groups of feature elements; a distribution prior estimation module, configured to perform step-by-step prior estimation based on the four groups of characteristic elements and the super prior using a quadtree-based spatial entropy model to obtain a four-step spatial prior; a semantic importance estimation module, configured to estimate the four-step semantic importance based on the channel state, the four groups of feature elements, and the four-step spatial domain prior; A semantic feature encoding module is configured to input, for the i-th path among the four paths, a combination of feature elements corresponding to the i-th step estimated index, the i-th step semantic importance, the semantic feature encoding results of the paths before the i-th path, and the channel state, into the i-th path semantic feature encoder, and output the i-th path semantic feature encoding result. The semantic feature encoder is a four-path semantic feature encoder corresponding to the four-step prior estimation of the quadtree-based spatial entropy model, and the value range of i is 0-3; The transmission module is used to map the semantic feature encoding results of each channel into symbols of different lengths for wireless transmission through the corresponding symbol mapping modules of each channel; The restoration module is used for the symbol mapping module of each channel at the receiving end to restore the symbols according to their dimensions; A semantic feature decoding module is configured to input the channel state, the received semantic feature encoding result of the i-th channel, and the semantic feature decoding results of each channel before the i-th channel into the i-th semantic feature decoder for the i-th channel, to obtain the i-th semantic feature decoding result, wherein the semantic feature decoder is a four-channel semantic feature decoder, and the four-channel semantic feature decoding results constitute the received semantic feature; A semantic recovery module, configured to input the received semantic features and the channel state into a semantic information restorer, wherein the semantic information restorer is configured to restore the received semantic features into noisy features through upsampling; The reconstruction module is used to input the noisy features and the channel state into the denoising and enhancement module to reconstruct the image.
Citation Information
Patent Citations
Joint information source channel coding method for image semantic communication
CN117879765A
Channel adaptive semantic communication method and system for multi-vision task
CN118691937A