A lightweight semantic communication method and system for image transmission

By adopting shifted convolution and multiple downsampling techniques in the semantic communication system, the problem of high resource consumption is solved, the efficient application of lightweight semantic communication in IoT devices is realized, and the efficiency of semantic feature extraction and image reconstruction quality are improved.

CN120238649BActive Publication Date: 2025-10-10NANCHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510713473.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-10
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Existing semantic communication systems consume a lot of resources and are difficult to be efficiently applied to IoT devices, which limits the development of IoT technology.

Method used

Shifted convolution is used to replace traditional convolution, combined with external shifted convolution and multiple downsampling, and feature extraction is optimized through feature splicing and residual connection, which reduces computational complexity and improves the efficiency of semantic feature extraction.

Benefits of technology

It effectively reduces the demand for computing resources, improves the efficiency of semantic feature extraction and the quality of reconstructed images, and is suitable for lightweight semantic communication of IoT devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238649B_ABST
    Figure CN120238649B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of communication, and provides a lightweight semantic communication method and system for image transmission, which replaces traditional convolution in down-sampling by shift convolution, and performs outer shift convolution processing on an input target image, splices a first feature map obtained through the outer shift convolution into input data of down-sampling, and forms twice shift convolution operation, so that the extraction efficiency of semantic features can be effectively improved; the down-sampling is continuously executed for multiple times, the outer shift convolution can be reused, the demand for computing resources is further reduced, and the lightweight degree is improved; the first feature map is original information obtained according to the target image, the original information is spliced into each down-sampling, the influence of original information loss caused by continuous multiple shift convolutions can be reduced, and the quality of a reconstructed image is improved. The lightweight semantic communication method and system for image transmission provided by the application can meet the demand for lightweight through the combination of inner and outer shift convolutions, and the quality of image reconstruction can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a lightweight semantic communication method and system for image transmission. Background Art

[0002] With the rapid development of bandwidth-intensive applications such as the Internet of Things and virtual reality, the demand for data transmission in wireless communication networks has increased exponentially, making the design of efficient wireless communication systems a research hotspot. Semantic communication, with its potential to improve transmission efficiency, reduce redundancy, and adapt to complex channel environments, has become a key research direction in 6G communication technology.

[0003] Research on improving the performance of semantic communication systems has made significant progress, but efficient semantic extraction remains a key challenge. While traditional semantic extraction methods based on convolutional neural networks (CNNs) can achieve high accuracy, they consume significant computing resources. Deep learning (DL), with its significant advantage in automatically extracting semantic features, has been applied in semantic communication systems. However, its demands on power consumption and computing resources remain high, and the network models employed are parameter-heavy. The limited power consumption and computing resources of IoT devices make it difficult to efficiently apply existing semantic communication systems to IoT devices, hindering the development of IoT technology. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a lightweight semantic communication method and system for image transmission to solve the problem that the semantic communication system in the existing technology has high resource consumption, is difficult to be efficiently applied to IoT devices, and limits the development of IoT technology.

[0005] The present invention provides a lightweight semantic communication method for image transmission, comprising:

[0006] The target image is sequentially subjected to a first convolution process, downsampling, and a first dilated spatial pyramid pooling to obtain encoded data;

[0007] performing upsampling, a second convolution process, and a second dilated spatial pyramid pooling in sequence according to the received encoded data to obtain a reconstructed image of the target image;

[0008] and performing an outer shift convolution on the target image to obtain a first feature map, wherein a feature size of the first feature map is consistent with a feature size of the downsampled input data;

[0009] The downsampling includes sequentially performing feature splicing, inner shift convolution, and nonlinear activation, and the feature splicing is used to splice the first feature map into the downsampled input data;

[0010] The first feature map is also superimposed on the output data of the inner shift convolution through a residual connection;

[0011] The outer shift convolution and the inner shift convolution both include shifting, batch normalization and pointwise convolution performed in sequence;

[0012] The downsampling is performed multiple times in succession.

[0013] Optionally, the downsampling further includes:

[0014] Adjusting the channel weights of the downsampled input data according to a channel attention mechanism to obtain first intermediate data;

[0015] The feature splicing is used to splice the first feature map to the first intermediate data to obtain second intermediate data, so as to perform inner shift convolution based on the second intermediate data.

[0016] Optionally, the step of adjusting the channel weight of the input data according to the channel attention mechanism and obtaining the first intermediate data further includes:

[0017] The third convolution process, the first global average pooling, the fully connected layer dimensionality reduction, activation, the fully connected layer dimensionality increase, the second global average pooling and channel weighting are performed in sequence;

[0018] The output of the third convolution processing is also superimposed on the output data of the second global average pooling through a residual connection.

[0019] Optionally, the step of obtaining the encoded data is a sequential processing mode, and also includes: obtaining a scaling factor based on the difference between the size of the currently downsampled input data and the size of the target image, and obtaining the shift operation step size of the outer shift convolution based on the scaling factor, so that the size of the obtained first feature map is consistent with the size of the currently downsampled input data.

[0020] Optionally, the method further includes: splicing the first feature map into the input data of the first dilated spatial pyramid pooling.

[0021] Another aspect of the present invention provides a lightweight semantic communication system for image transmission, comprising:

[0022] An encoder, comprising a first convolution module, a downsampling module, and a first dilated spatial pyramid pooling module connected in sequence, configured to sequentially perform a first convolution process, downsampling, and first dilated spatial pyramid pooling on a target image to obtain encoded data;

[0023] a decoder comprising a sequentially connected up-sampling module, a second convolution module and a second empty space pyramid pooling module, configured to sequentially perform up-sampling, second convolution processing and second empty space pyramid pooling according to the received encoding data to obtain a reconstructed image of the target image;

[0024] The encoder further comprises an outer shift convolution unit, configured to perform outer shift convolution on the target image to obtain a first feature map, the feature size of the first feature map being consistent with the feature size of the down-sampling.

[0025] The down-sampling module comprises a sequentially connected feature splicing unit, an inner shift convolution unit and a nonlinear activation unit, wherein the feature splicing unit is configured to splice the first feature map into the input data of the down-sampling module.

[0026] The outer shift convolution unit is further connected to the output end of the inner shift convolution unit to superimpose the first feature map into the output data of the inner shift convolution unit through a residual connection.

[0027] The outer shift convolution unit and the inner shift convolution unit each comprise a sequentially connected shift layer, a batch normalization layer and a point state convolution layer.

[0028] The down-sampling module is continuously provided with a plurality of.

[0029] Optionally, the down-sampling module further comprises:

[0030] a compression excitation unit configured to adjust the channel weight of the input data of the down-sampling according to a channel attention mechanism to obtain first intermediate data.

[0031] The feature splicing unit is configured to splice the first feature map into the first intermediate data to obtain second intermediate data, and take the second intermediate data as the input data of the inner shift convolution unit.

[0032] Optionally, the compression excitation unit comprises:

[0033] a third convolution layer, a first global average pooling layer, a fully connected layer dimension reduction layer, an activation layer, a fully connected layer dimension increasing layer, a second global average pooling layer and a channel weighting layer connected in sequence.

[0034] The output of the third convolution layer is further superimposed into the output data of the second global average pooling layer through a residual connection.

[0035] Optionally, the working mode of the encoder is a sequential processing mode, and the outer shift convolution unit is further configured to: obtain a scaling factor according to a size difference between input data of the current downsampling module and the target image, and obtain a shift operation step of the outer shift convolution unit according to the scaling factor, so that a size of the obtained first feature map is consistent with a size of the input data of the current downsampling module.

[0036] Optionally, the first feature map is further spliced into the input data of the first empty space pyramid pooling module.

[0037] The lightweight semantic communication method for image transmission provided by the present application can effectively reduce the computational complexity by replacing the traditional convolution requirement in downsampling with shift convolution, and the input target image is processed by outer shift convolution, the obtained first feature map is spliced into the input data of downsampling, and two times of shift convolution operation is constructed, which can effectively improve the extraction efficiency of semantic features. And the downsampling is continuously executed for multiple times, and the outer shift convolution can be reused, further reducing the demand for computing resources and improving the lightweight degree; the first feature map is the original information obtained according to the target image, and the original information is spliced into each downsampling, which can reduce the influence of the loss of original information caused by continuous multiple shift convolutions, and improve the quality of the reconstructed image. The lightweight semantic communication method for image transmission provided by the present application can effectively reduce the demand for computing resources of the encoding end by combining outer shift convolution and multiple inner shift convolutions of downsampling, and can reduce the loss of original information and improve the quality of the reconstructed image obtained by decoding. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The main flowchart of the lightweight semantic communication method for image transmission in the embodiment of the present application;

[0039] Figure 2 The downsampling related flowchart of the lightweight semantic communication method for image transmission in the embodiment of the present application;

[0040] Figure 3 The test results of the lightweight semantic communication method for image transmission in the embodiment of the present application.

[0041] The following specific embodiments will further illustrate the present application in conjunction with the above drawings. DETAILED DESCRIPTION

[0042] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings. The drawings show several embodiments of the present application. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0043] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0045] In order to solve the problem that the semantic communication system in the prior art has high resource consumption and is difficult to be efficiently applied to IoT devices, which limits the development of IoT technology. The present application provides a lightweight semantic communication method for image transmission, which can effectively reduce the computational complexity by replacing the traditional convolution requirements in downsampling with shifted convolution, and performs external shifted convolution processing on the input target image, and splices the obtained first feature map into the internal shifted convolution of the downsampling, forming two shifted convolution operations, which can effectively improve the efficiency of semantic feature extraction; and the downsampling is performed continuously for multiple times, and the external shifted convolution can be reused, further reducing the demand for computing resources and improving the degree of lightweight; the first feature map is the original information obtained according to the target image, and the original information is spliced ​​into each downsampling, which can reduce the impact of the original information loss caused by multiple consecutive shifted convolutions and improve the quality of the reconstructed image.

[0046] Please refer to Figure 1 and Figure 2 , shown are the flow charts of the lightweight semantic communication method for image transmission in the embodiments of the invention.

[0047] At the encoder side, the encoding step includes: performing a first convolution process, downsampling, and a first dilated spatial pyramid pooling on the target image in sequence to obtain encoded data, and the encoded data is sent to the receiving end through a physical channel.

[0048] At the decoder side, the decoding steps include upsampling, a second convolution process, and a second dilated spatial pyramid pooling to obtain a reconstructed image of the target image.

[0049] Among them, the first convolution processing and the second convolution processing can be implemented using convolutional neural networks (CNN). The first convolution processing is used to increase the number of feature channels of the data, such as increasing the R, G, and B three-color feature channels of the RGB image to 32 channels; the second convolution processing is used for multiple tasks such as feature refinement, fusion, noise reduction, and dimensionality adjustment to obtain the original information of the target image.

[0050] The downsampling selection shift convolution (ShiftConv) technology can effectively reduce the demand for computing resources and achieve lightweight design compared to traditional CNN convolution.

[0051] like Figure 2 As shown in the figure, downsampling mainly includes inner shift convolution operation, and outer shift convolution operation includes shift (Shift), batch normalization (BatchNorm) and point-wise convolution (1×1 convolution) performed in sequence. Shift is used to enhance feature diversity through spatial translation, batch normalization is used to ensure the stability of training, and point-wise convolution is used to optimize channel interaction, which can improve the overall training efficiency and robustness of the inner shift convolution model.

[0052] In order to improve the efficiency of semantic feature extraction, in this embodiment, downsampling is performed continuously for multiple times, and the encoder side also performs outward shift convolution. The outward shift convolution processes the target image to obtain a first feature map. The first feature map is spliced ​​to the input data of each downsampled inner shift convolution through a feature splicing operation, which can reduce the loss of the original information of the target image in multiple shift convolutions.

[0053] The operation and architecture of the outer shift convolution are consistent with those of the inner shift convolution, so that the feature size of the first feature map obtained by the outer shift convolution is consistent with the feature size of the inner shift convolution, so that the first feature map can be effectively spliced ​​into the input data of the inner shift convolution through feature splicing.

[0054] Among them, the feature concatenation operation can be specifically performed by channel dimension concatenation (Channel-wiseConcatenation).

[0055] After the inner shift convolution, nonlinear activation is performed using the PReLU activation function. This introduction of nonlinearity breaks the linear constraint, optimizes gradient propagation, enhances feature expression, and effectively improves the efficiency of semantic feature extraction.

[0056] To improve the downsampling training efficiency, in this embodiment, the first feature map is also superimposed on the output data of the inner shift convolution through a residual connection.

[0057] To improve the model performance, the downsampling further includes a Squeeze-and-Excitation (SE) operation for adjusting the channel weights of the input data of the downsampling according to a channel attention mechanism to obtain first intermediate data. The Squeeze operation mainly performs global average pooling, and the Excitation operation mainly performs a Fully Connected Layer processing. The global average pooling and the Fully Connected Layer are used to enhance the sensitivity of the model to the channels, and improve the network performance while keeping the image feature resolution unchanged.

[0058] The feature concatenation is used to concatenate the first feature map to the first intermediate data output by the SE operation to obtain second intermediate data, and the inner shift convolution is performed according to the second intermediate data.

[0059] As shown in Figure 2 The SE operation specifically includes a third convolution processing (implemented by a CNN), a first global average pooling, a Fully Connected Layer dimension reduction, an activation (activated by a ReLU activation function), a Fully Connected Layer dimension increase, a second global average pooling, and a channel weighting (Scale), which are sequentially performed. The output of the third convolution processing is also superimposed into the output data of the second global average pooling through a residual connection.

[0060] After the downsampling, the size of the output feature map is different from that of the input feature map. To ensure the size matching of the first feature map and the output feature maps of the multiple outer shift convolutions, the step of obtaining the encoded data is in a sequential processing mode (after the encoding of the current image is completed, the encoding of the next image is performed), and further includes obtaining a scaling factor according to the size difference between the input data of the current downsampling and the target image, and obtaining a shift operation step length of the outer shift convolution according to the scaling factor, so that the size of the obtained first feature map is consistent with that of the input data of the current downsampling.

[0061] When the IoT device is applied in a scenario with low communication speed requirement, the sequential processing mode is adopted, and the step length of the outer shift convolution is dynamically adjusted, so that the size matching requirement of the two feature maps participating in the concatenation is met, the effectiveness of the concatenation is ensured, the reuse of the outer shift convolution is realized, and the lightweight degree of the system is ensured.

[0062] Specifically, the image is generally two-dimensional data, and the two-dimensional size of the obtained feature map after the shift convolution is different from that of the original image. The scaling factors (scaling ratios) in the two dimensions can be obtained respectively, and the average value of the two scaling factors is taken as the scaling factor of the outer shift convolution. The size transformation formula of the outer shift convolution is: wherein, is the feature map size after the shift convolution, For the floor operator, i is the input image size, p is the padding operation, k is the size of the convolution kernel, S is the scaling factor, which is consistent with the shift step of the shift convolution.

[0063] To further improve the feature extraction effect, as shown in the embodiment, the first feature map is spliced into the input data of the first Atrous Spatial Pyramid Pooling (ASPP). Figure 1 Splicing the first feature map containing the original information of the target image into the first Atrous Spatial Pyramid Pooling can reduce the loss of details (such as edges and textures) of deep features and improve the multi-scale feature extraction effect.

[0064] The application also provides a lightweight semantic communication system for image transmission, comprising:

[0065] The encoder comprises a first convolution module, a down-sampling module and a first Atrous Spatial Pyramid Pooling module connected in sequence, and is used for sequentially performing first convolution processing, down-sampling and first Atrous Spatial Pyramid Pooling on the target image to obtain encoded data.

[0066] The decoder comprises an up-sampling module, a second convolution module and a second Atrous Spatial Pyramid Pooling module connected in sequence, and is used for sequentially performing up-sampling, second convolution processing and second Atrous Spatial Pyramid Pooling according to the received encoded data to obtain a reconstructed image of the target image.

[0067] The encoder further comprises an outer shift convolution unit, which is used for performing outer shift convolution on the target image to obtain a first feature map, and the feature size of the first feature map is consistent with the feature size of the down-sampling.

[0068] The down-sampling module comprises a feature splicing unit, an inner shift convolution unit and a nonlinear activation unit connected in sequence, wherein the feature splicing is used for splicing the first feature map into the input data of the down-sampling module.

[0069] The outer shift convolution unit is further connected to the output end of the inner shift convolution unit to superimpose the first feature map through a residual connection into the output data of the inner shift convolution unit.

[0070] The outer shift convolution unit and the inner shift convolution unit each comprise a shift layer, a batch normalization layer and a point state convolution layer connected in sequence.

[0071] The down-sampling module is continuously provided with a plurality of.

[0072] The downsampling module also includes: a compression excitation unit, which is used to adjust the channel weights of the downsampled input data according to the channel attention mechanism to obtain first intermediate data; wherein, the feature splicing unit is used to splice the first feature map to the first intermediate data to obtain second intermediate data, and use the second intermediate data as input data of the inner shift convolution unit.

[0073] The compression excitation unit specifically includes: a third convolutional layer, a first global average pooling layer, a fully connected layer dimensionality reduction layer, an activation layer, a fully connected layer dimensionality increase layer, a second global average pooling layer and a channel weighted layer connected in sequence; wherein, the output of the third convolutional layer is also superimposed on the output data of the second global average pooling layer through a residual connection.

[0074] This embodiment is mainly used in scenarios with low requirements for communication speed. Correspondingly, the working mode of the encoder is a sequential processing mode, and the outer shift convolution unit is also used to: obtain a scaling factor based on the difference between the size of the input data of the current downsampling module and the size of the target image, and obtain the shift operation step of the outer shift convolution unit based on the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling module.

[0075] like Figure 3 As shown, the test results of the lightweight semantic communication method for image transmission in this embodiment are shown. It uses the peak signal-to-noise ratio (PSNR) as the quality evaluation indicator of the reconstructed image. Under low signal-to-noise ratio (SNR) conditions, it can maintain excellent signal recovery capabilities. As the SNR increases, its performance shows a stable growth trend without a sharp decline. The original information loss is small, and the image reconstruction quality is guaranteed.

[0076] The lightweight semantic communication method for image transmission provided by the present invention replaces the traditional convolution requirements in downsampling with shifted convolution, which can effectively reduce the computational complexity, and performs external shifted convolution processing on the input target image, and splices the first feature map obtained by the external shifted convolution into the downsampled input data, constituting two shifted convolution operations, which can effectively improve the efficiency of semantic feature extraction; and the downsampling is performed continuously for multiple times, and the external shifted convolution can be reused, further reducing the demand for computing resources and improving the degree of lightweight; the first feature map is the original information obtained according to the target image, and the original information is spliced ​​into each downsampling, which can reduce the impact of the original information loss caused by multiple consecutive shifted convolutions and improve the quality of the reconstructed image.

[0077] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0078] The above-described embodiments only express several specific implementations of the present application, which are described in a more specific and detailed manner, but cannot be understood as a limitation on the protection scope of the present application. It should be noted that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the present application patent should be subject to the appended claims.

Claims

1. A lightweight semantic communication method for image transmission, characterized in that: include: The target image is sequentially subjected to a first convolution process, downsampling, and a first dilated spatial pyramid pooling to obtain encoded data; performing upsampling, a second convolution process, and a second dilated spatial pyramid pooling in sequence according to the received encoded data to obtain a reconstructed image of the target image; and performing an outer shift convolution on the target image to obtain a first feature map, wherein a feature size of the first feature map is consistent with a feature size of the downsampled input data; The downsampling includes sequentially performing feature splicing, inner shift convolution, and nonlinear activation, and the feature splicing is used to splice the first feature map into the downsampled input data; The first feature map is also superimposed on the output data of the inner shift convolution through a residual connection; The outer shift convolution and the inner shift convolution both include shifting, batch normalization and pointwise convolution performed in sequence; The downsampling is performed multiple times continuously; Among them, the step of obtaining the encoded data is a sequential processing mode, and also includes: obtaining a scaling factor based on the difference between the size of the currently downsampled input data and the size of the target image, and obtaining the shift operation step size of the outer shift convolution based on the scaling factor, so that the size of the obtained first feature map is consistent with the size of the currently downsampled input data.

2. The lightweight semantic communication method for image transmission according to claim 1, characterized in that: The downsampling further comprises: Adjusting the channel weights of the downsampled input data according to a channel attention mechanism to obtain first intermediate data; The feature splicing is used to splice the first feature map to the first intermediate data to obtain second intermediate data, so as to perform inner shift convolution based on the second intermediate data.

3. The lightweight semantic communication method for image transmission according to claim 2, characterized in that: The step of adjusting the channel weight of the input data according to the channel attention mechanism, and obtaining the first intermediate data further includes: The third convolution process, the first global average pooling, the fully connected layer dimensionality reduction, activation, the fully connected layer dimensionality increase, the second global average pooling and channel weighting are performed in sequence; The output of the third convolution processing is also superimposed on the output data of the second global average pooling through a residual connection.

4. The lightweight semantic communication method for image transmission according to claim 1, characterized in that: Also includes: The first feature map is spliced ​​into the input data of the first dilated spatial pyramid pooling.

5. A lightweight semantic communication system for image transmission, characterized in that: include: An encoder, comprising a first convolution module, a downsampling module, and a first dilated spatial pyramid pooling module connected in sequence, configured to sequentially perform a first convolution process, downsampling, and first dilated spatial pyramid pooling on a target image to obtain encoded data; A decoder, comprising an upsampling module, a second convolution module, and a second atrous spatial pyramid pooling module connected in sequence, configured to sequentially perform upsampling, a second convolution process, and a second atrous spatial pyramid pooling on the received encoded data to obtain a reconstructed image of the target image; The encoder further includes an outer shift convolution unit, which is used to perform an outer shift convolution on the target image to obtain a first feature map, wherein the feature size of the first feature map is consistent with the feature size of the downsampling; The downsampling module includes a feature splicing unit, an inner shift convolution unit, and a nonlinear activation unit connected in sequence, wherein the feature splicing is used to splice the first feature map into the input data of the downsampling module; The outer shift convolution unit is further connected to the output end of the inner shift convolution unit to superimpose the first feature map onto the output data of the inner shift convolution unit through a residual connection; The outer shift convolution unit and the inner shift convolution unit each include a shift layer, a batch normalization layer and a pointwise convolution layer connected in sequence; The downsampling modules are provided in plurality; In which, the working mode of the encoder is a sequential processing mode, and the external shift convolution unit is also used to: obtain a scaling factor based on the difference between the size of the input data of the current downsampling module and the size of the target image, and obtain the shift operation step of the external shift convolution unit based on the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling module.

6. The lightweight semantic communication system for image transmission according to claim 5, characterized in that: The downsampling module further includes: A compression excitation unit, configured to adjust the channel weights of the downsampled input data according to a channel attention mechanism to obtain first intermediate data; The feature splicing unit is used to splice the first feature map to the first intermediate data to obtain second intermediate data, and use the second intermediate data as input data of the inner shift convolution unit.

7. The lightweight semantic communication system for image transmission according to claim 6, characterized in that: The compression excitation unit comprises: The third convolutional layer, the first global average pooling layer, the fully connected layer dimensionality reduction layer, the activation layer, the fully connected layer dimensionality increase layer, the second global average pooling layer and the channel weighted layer are connected in sequence; The output of the third convolutional layer is also superimposed on the output data of the second global average pooling layer through a residual connection.

8. The lightweight semantic communication system for image transmission according to claim 5, characterized in that: The first feature map is also spliced ​​into the input data of the first atrous spatial pyramid pooling module.

Citation Information

Patent Citations

  • Real-time semantic segmentation method and system for multi-shape pyramid in traffic scene

    CN119399457A