Lightweight semantic communication method and system for image transmission

By using shift convolution and feature stitching technology in semantic communication systems, a lightweight semantic communication system is built, which solves the problem of high resource consumption of IoT devices and realizes efficient semantic feature extraction and image reconstruction.

CN120238649AActive Publication Date: 2025-07-01NANCHANG UNIV
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510713473.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Existing semantic communication systems consume high resources in IoT devices and are difficult to apply efficiently, which limits the development of IoT technology.

Method used

Shift convolution is used to replace the traditional convolution in downsampling, combining outer shift convolution and inner shift convolution, and through feature stitching and residual connection, a lightweight semantic communication system is built to reduce the computational complexity and improve the semantic feature extraction efficiency.

Benefits of technology

It effectively reduces the demand for computing resources, improves semantic feature extraction efficiency and reconstructs image quality, and meets the lightweight needs of IoT devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238649A_ABST
    Figure CN120238649A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication, and provides an image transmission-oriented lightweight semantic communication method and system, and the method comprises the steps: replacing the conventional convolution in down-sampling through shift convolution, carrying out the external shift convolution processing of an input target image, splicing a first feature map obtained through the external shift convolution into the input data of down-sampling, and carrying out the external shift convolution processing of the input target image, two shift convolution operations are formed, so that the extraction efficiency of semantic features can be effectively improved; downsampling is continuously executed for multiple times, and external shift convolution can be reused, so that the demand on computing resources is further reduced, and the lightweight degree is improved; the first feature map is original information obtained according to the target image, and the original information is spliced into each downsample, so that the original information loss influence caused by continuous multiple shift convolution can be reduced, and the quality of the reconstructed image is improved. According to the image transmission-oriented lightweight semantic communication method and system provided by the invention, through the combination of internal and external shift convolution, the lightweight requirement can be met, and the image reconstruction quality can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a lightweight semantic communication method and system for image transmission. Background Art

[0002] With the rapid development of bandwidth-intensive applications such as the Internet of Things and virtual reality, the demand for data transmission in wireless communication networks has increased exponentially, which makes the design of efficient wireless communication systems a research hotspot. Among them, semantic communication has become an important research direction for 6G communication technologies due to its potential in improving transmission efficiency, reducing redundancy, and adapting to complex channel environments.

[0003] Currently, many progresses have been made in the research on improving the performance of semantic communication systems, but efficient semantic extraction remains a key challenge to be solved. Although traditional semantic extraction methods based on Convolutional Neural Network (CNN) can achieve high accuracy, they consume a large amount of computing resources. Deep Learning (DL) has been applied in semantic communication systems due to its significant advantage of automatically extracting semantic features, but its demand for power consumption and computing resources is still high, and the number of parameters of the adopted network model is large. Moreover, the power consumption and computing resources of Internet of Things devices are limited, making it difficult for existing semantic communication systems to be efficiently applied to Internet of Things devices, which restricts the development of Internet of Things technologies. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a lightweight semantic communication method and system for image transmission to solve the problem that the existing semantic communication systems consume high resources, are difficult to be efficiently applied to Internet of Things devices, and restrict the development of Internet of Things technologies.

[0005] On the one hand, the present invention provides a lightweight semantic communication method for image transmission, including: Performing first convolution processing, downsampling, and first atrous spatial pyramid pooling on a target image in sequence to obtain encoded data; Performing upsampling, second convolution processing, and second atrous spatial pyramid pooling on the received encoded data in sequence to obtain a reconstructed image of the target image; And, performing external shifted convolution on the target image to obtain a first feature map, where the feature size of the first feature map is the same as the feature size of the input data of the downsampling; Wherein, the downsampling includes feature splicing, internal shifted convolution, and non-linear activation performed in sequence, and the feature splicing is used to splice the first feature map into the input data of the downsampling; The first feature map is also superimposed on the output data of the inner shifted convolution through a residual connection; Both the outer shifted convolution and the inner shifted convolution include shifting, batch normalization, and pointwise convolution performed in sequence; The downsampling is continuously performed multiple times.

[0006] Optionally, the downsampling further includes: Adjusting the channel weights of the input data of the downsampling according to the channel attention mechanism to obtain first intermediate data; Wherein, the feature concatenation is used to concatenate the first feature map to the first intermediate data to obtain second intermediate data, and the inner shifted convolution is performed according to the second intermediate data.

[0007] Optionally, the step of adjusting the channel weights of the input data according to the channel attention mechanism to obtain first intermediate data further includes: Third convolution processing, first global average pooling, fully connected layer dimensionality reduction, activation, fully connected layer dimensionality increase, second global average pooling, and channel weighting performed in sequence; Wherein, the output of the third convolution processing is also superimposed on the output data of the second global average pooling through a residual connection.

[0008] Optionally, the step of obtaining the encoded data is in sequential processing mode, and further includes: obtaining a scaling factor according to the size difference between the input data of the current downsampling and the size of the target image, and obtaining the shifting operation step size of the outer shifted convolution according to the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling.

[0009] Optionally, it further includes: concatenating the first feature map to the input data of the first atrous spatial pyramid pooling.

[0010] On the other hand, the present invention provides a lightweight semantic communication system for image transmission, including: An encoder, including a first convolution module, a downsampling module, and a first atrous spatial pyramid pooling module connected in sequence, for performing first convolution processing, downsampling, and first atrous spatial pyramid pooling on a target image in sequence to obtain encoded data; A decoder, including an upsampling module, a second convolution module, and a second atrous spatial pyramid pooling module connected in sequence, for performing upsampling, second convolution processing, and second atrous spatial pyramid pooling on the received encoded data in sequence to obtain a reconstructed image of the target image; Among them, the encoder further includes an outer shift convolution unit, which is used to perform outer shift convolution on the target image to obtain a first feature map, and the feature size of the first feature map is the same as the feature size of the downsampling; The downsampling module includes a feature splicing unit, an inner shift convolution unit, and a non-linear activation unit connected in sequence. Among them, the feature splicing is used to splice the first feature map into the input data of the downsampling module; The outer shift convolution unit is also connected to the output end of the inner shift convolution unit to superimpose the first feature map on the output data of the inner shift convolution unit through residual connection; Both the outer shift convolution unit and the inner shift convolution unit include a shift layer, a batch normalization layer, and a pointwise convolution layer connected in sequence; A plurality of the downsampling modules are continuously arranged.

[0011] Optionally, the downsampling module further includes: A squeeze-and-excitation unit, which is used to adjust the channel weights of the input data of the downsampling according to the channel attention mechanism to obtain first intermediate data; Among them, the feature splicing unit is used to splice the first feature map into the first intermediate data to obtain second intermediate data, and use the second intermediate data as the input data of the inner shift convolution unit.

[0012] Optionally, the squeeze-and-excitation unit includes: A third convolution layer, a first global average pooling layer, a fully connected layer dimensionality reduction layer, an activation layer, a fully connected layer dimensionality increase layer, a second global average pooling layer, and a channel weighting layer connected in sequence; Among them, the output of the third convolution layer is also superimposed on the output data of the second global average pooling layer through residual connection.

[0013] Optionally, the working mode of the encoder is a sequential processing mode, and the outer shift convolution unit is further used to: obtain a scaling factor according to the size difference between the input data of the current downsampling module and the size of the target image, and obtain the shift operation step size of the outer shift convolution unit according to the scaling factor, so that the size of the obtained first feature map is the same as the size of the input data of the current downsampling module.

[0014] Optionally, the first feature map is also spliced into the input data of the first atrous spatial pyramid pooling module.

[0015] The lightweight semantic communication method for image transmission provided by the present invention replaces the traditional convolution requirement in downsampling through shifted convolution, which can effectively reduce the computational complexity. Moreover, the input target image is processed by outer shifted convolution, and the obtained first feature map is spliced into the input data of downsampling to form two shifted convolution operations, which can effectively improve the extraction efficiency of semantic features. And the downsampling is continuously executed multiple times, and the outer shifted convolution can be reused, further reducing the demand for computing resources and improving the degree of lightweight. The first feature map is the original information obtained according to the target image. Splicing the original information into each downsampling can reduce the influence of the loss of original information caused by multiple consecutive shifted convolutions and improve the quality of the reconstructed image. The lightweight semantic communication method for image transmission provided by the present invention can effectively reduce the demand for computing resources at the encoding end through the combination of outer shifted convolution and inner shifted convolution of multiple downsamplings, and can reduce the loss of original information and improve the quality of the reconstructed image obtained by decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is the main flowchart of the lightweight semantic communication method for image transmission in the embodiment of the present invention; Figure 2 It is the flowchart related to downsampling of the lightweight semantic communication method for image transmission in the embodiment of the present invention; Figure 3 It is the test result of the lightweight semantic communication method for image transmission in the embodiment of the present invention.

[0017] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0019] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be a middle element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a middle element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0021] To solve the problems in the prior art that the semantic communication system has high resource consumption, is difficult to be efficiently applied to Internet of Things (IoT) devices, and restricts the development of IoT technology. This application provides a lightweight semantic communication method for image transmission. By replacing the traditional convolution in downsampling with shift convolution, the computational complexity can be effectively reduced. And the input target image is processed by external shift convolution, and the obtained first feature map is spliced into the internal shift convolution of downsampling to form two shift convolution operations, which can effectively improve the extraction efficiency of semantic features. And the downsampling is continuously executed multiple times, and the external shift convolution can be reused, further reducing the demand for computing resources and improving the degree of lightweight. The first feature map is the original information obtained from the target image. Splicing the original information into each downsampling can reduce the influence of the loss of original information caused by consecutive multiple shift convolutions and improve the quality of the reconstructed image.

[0022] Please refer to Figure 1 and Figure 2 , which shows the flowcharts of the lightweight semantic communication method for image transmission in the embodiments of the invention.

[0023] At the encoder end, the encoding steps include: sequentially performing a first convolution process, downsampling, and a first atrous spatial pyramid pooling on the target image to obtain encoded data, and the encoded data is sent through a physical channel to the receiving end.

[0024] At the decoder end, the decoding steps include: upsampling, a second convolution process, and a second atrous spatial pyramid pooling to obtain a reconstructed image of the target image.

[0025] Among them, the first convolution process and the second convolution process can be implemented by a Convolutional Neural Networks (CNN). The first convolution process is used to increase the number of feature channels of the data. For example, the R, G, and B color feature channels of an RGB image are increased to 32 channels. The second convolution process is used for multiple tasks such as feature refinement, fusion, noise reduction, and dimension adjustment to obtain the original information of the target image.

[0026] The downsampling selects the ShiftConv (Shift Convolution) technology, which can effectively reduce the demand for computing resources compared with the traditional CNN convolution and achieve lightweight design.

[0027] Such as Figure 2As shown, downsampling mainly includes inner shifted convolution operations. The outer shifted convolution operations include Shift, BatchNorm, and pointwise convolution (1×1 convolution) performed in sequence. Shift is used to enhance feature diversity through spatial translation, BatchNorm is used to ensure the stability of training, and pointwise convolution is used to optimize channel interaction, which can overall improve the model training efficiency and robustness of inner shifted convolution.

[0028] To improve the extraction efficiency of semantic features, in this embodiment, downsampling is continuously performed multiple times, and outer shifted convolution is also performed at the encoder end. The outer shifted convolution processes the target image to obtain a first feature map. The first feature map is concatenated into the input data of the inner shifted convolution of each downsampling through a feature concatenation operation, which can reduce the loss of the original information of the target image in multiple shifted convolutions.

[0029] The operations and architectures of the outer shifted convolution and the inner shifted convolution are consistent, so that the feature size of the first feature map obtained by the outer shifted convolution is the same as that of the inner shifted convolution, enabling the first feature map to be effectively concatenated into the input data of the inner shifted convolution through feature concatenation.

[0030] Among them, the feature concatenation operation can specifically be Channel-wise Concatenation.

[0031] After the inner shifted convolution, nonlinear activation is also performed through the PReLU activation function. By introducing nonlinearity, linear constraints can be broken, gradient propagation can be optimized, and feature expression can be enhanced, which can effectively improve the semantic feature extraction efficiency.

[0032] To improve the downsampling training efficiency, in this embodiment, the first feature map is also superimposed on the output data of the inner shifted convolution through a residual connection.

[0033] To improve the model performance, downsampling also includes: Squeeze-and-Excitation (SE operation), which is used to adjust the channel weights of the input data of downsampling according to the channel attention mechanism to obtain a first intermediate data. Among them, the Squeeze operation is mainly global average pooling, and the Excitation operation is mainly processed by a fully connected layer. By global average pooling and the fully connected layer, the model's sensitivity to channels is enhanced, and the network performance is improved while maintaining the image feature resolution unchanged.

[0034] Among them, feature concatenation is used to concatenate the first feature map into the first intermediate data output by the squeeze-and-excitation operation to obtain a second intermediate data, and the inner shifted convolution is performed according to the second intermediate data.

[0035] Such asFigure 2 As shown in Figure 2 , the compression excitation operation specifically includes: third convolution processing (implemented by CNN) performed sequentially, first global average pooling, dimensionality reduction by a fully connected layer, activation (activated by the ReLU activation function), dimensionality increase by a fully connected layer, second global average pooling, and channel weighting (Scale); wherein, the output of the third convolution processing is also superimposed on the output data of the second global average pooling through a residual connection.

[0036] After downsampling, there is a scaling difference in the sizes of the output feature map and the input feature map. To ensure the size matching between the first feature map and the output feature maps of each external shifted convolution performed multiple times, the steps to obtain the encoded data are in sequential processing mode (encoding the next image after the current image encoding is completed), and it also includes: obtaining a scaling factor based on the size difference between the input data of the current downsampling and the size of the target image, and obtaining the shift operation step size of the external shifted convolution according to the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling.

[0037] When the Internet of Things device is applied to scenarios with relatively low requirements for communication speed, adopting the sequential processing mode and dynamically adjusting the step size of the external shifted convolution can meet the size matching requirements of the two feature maps participating in the splicing, ensure the effectiveness of the splicing, realize the reuse of the external shifted convolution, and ensure the lightweight degree of the system.

[0038] Specifically, an image is generally two-dimensional data. After the shifted convolution, the two-dimensional size of the obtained feature map is different from that of the original image. Scaling factors (scaling ratios) in two dimensions can be obtained respectively, and the average value of the two scaling factors is used as the scaling factor of the external shifted convolution. The size transformation formula of the shifted convolution is: where is the size of the feature map after the shifted convolution, is the floor operator, i is the size of the input image, p is the padding operation, k is the size of the convolution kernel, and S is the scaling factor, which is consistent with the shift step size of the shifted convolution.

[0039] To further improve the feature extraction effect, as Figure 1 shown in Figure 1 , in this embodiment, it further includes: splicing the first feature map into the input data of the first Atrous Spatial Pyramid Pooling (ASPP). Splicing the first feature map containing the original information of the target image into the first Atrous Spatial Pyramid Pooling can reduce the details (such as edges and textures) lost in the deep features and improve the multi-scale feature extraction effect.

[0040] The present invention also provides a lightweight semantic communication system for image transmission, including: An encoder, including a first convolution module, a downsampling module, and a first atrous spatial pyramid pooling module connected in sequence, is used to perform first convolution processing, downsampling, and first atrous spatial pyramid pooling on a target image in sequence to obtain encoded data; A decoder, including an upsampling module, a second convolution module, and a second atrous spatial pyramid pooling module connected in sequence, is used to perform upsampling, second convolution processing, and second atrous spatial pyramid pooling on the received encoded data in sequence to obtain a reconstructed image of the target image; Among them, the encoder further includes an outer shift convolution unit, which is used to perform outer shift convolution on the target image to obtain a first feature map, and the feature size of the first feature map is consistent with the feature size of the downsampling; The downsampling module includes a feature splicing unit, an inner shift convolution unit, and a non-linear activation unit connected in sequence. Among them, the feature splicing is used to splice the first feature map into the input data of the downsampling module; The outer shift convolution unit is also connected to the output end of the inner shift convolution unit to superimpose the first feature map on the output data of the inner shift convolution unit through a residual connection; Both the outer shift convolution unit and the inner shift convolution unit include a shift layer, a batch normalization layer, and a pointwise convolution layer connected in sequence; A plurality of downsampling modules are continuously arranged.

[0041] The downsampling module further includes: a squeeze-and-excitation unit, which is used to adjust the channel weights of the input data of the downsampling according to the channel attention mechanism to obtain first intermediate data; among them, the feature splicing unit is used to splice the first feature map into the first intermediate data to obtain second intermediate data, and use the second intermediate data as the input data of the inner shift convolution unit.

[0042] The squeeze-and-excitation unit specifically includes: a third convolution layer, a first global average pooling layer, a fully connected layer dimensionality reduction layer, an activation layer, a fully connected layer dimensionality increase layer, a second global average pooling layer, and a channel weighting layer connected in sequence; among them, the output of the third convolution layer is also superimposed on the output data of the second global average pooling layer through a residual connection.

[0043] This embodiment is mainly used in scenarios with relatively low requirements for communication speed. Correspondingly, the working mode of the encoder is a sequential processing mode. The outer shift convolution unit is further used to: obtain a scaling factor according to the size difference between the input data of the current downsampling module and the size of the target image, and obtain the shift operation step size of the outer shift convolution unit according to the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling module.

[0044] Such as Figure 3As shown, it is the test result of the lightweight semantic communication method for image transmission in this embodiment. It uses the peak signal-to-noise ratio (PSNR) as the quality evaluation index of the reconstructed image. Under the condition of low signal-to-noise ratio (SNR), it can maintain excellent signal recovery ability. As the SNR increases, its performance shows a stable growth trend, without sharp decline, with small loss of original information and guaranteed image reconstruction quality.

[0045] The lightweight semantic communication method for image transmission provided by the present invention replaces the traditional convolution requirement in downsampling through shifted convolution, which can effectively reduce the computational complexity. And it performs shifted convolution processing on the input target image, and splices the first feature map obtained by the shifted convolution to the input data of the downsampling to form two shifted convolution operations, which can effectively improve the extraction efficiency of semantic features. And the downsampling is continuously executed multiple times, and the shifted convolution can be reused, further reducing the demand for computing resources and improving the degree of lightweight. The first feature map is the original information obtained from the target image. Splicing the original information into each downsampling can reduce the influence of the loss of original information caused by multiple consecutive shifted convolutions and improve the quality of the reconstructed image.

[0046] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0047] The above-described embodiments only represent several specific implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A lightweight semantic communication method for image transmission, characterized in that, Including: Performing first convolution processing, downsampling, and first atrous spatial pyramid pooling on the target image in sequence to obtain encoded data; Performing upsampling, second convolution processing, and second atrous spatial pyramid pooling on the received encoded data in sequence to obtain a reconstructed image of the target image; And performing external shifted convolution on the target image to obtain a first feature map, the feature size of the first feature map being consistent with the feature size of the input data of the downsampling; Wherein the downsampling includes feature concatenation, internal shifted convolution, and non-linear activation performed in sequence, and the feature concatenation is used to concatenate the first feature map into the input data of the downsampling; The first feature map is also superimposed on the output data of the internal shifted convolution through a residual connection; Both the external shifted convolution and the internal shifted convolution include shifting, batch normalization, and pointwise convolution performed in sequence; The downsampling is continuously executed multiple times.

2. The lightweight semantic communication method for image transmission according to claim 1, characterized in that The downsampling further includes: Adjusting the channel weights of the input data of the downsampling according to the channel attention mechanism to obtain first intermediate data; Wherein the feature concatenation is used to concatenate the first feature map into the first intermediate data to obtain second intermediate data, and the internal shifted convolution is performed according to the second intermediate data.

3. The lightweight semantic communication method for image transmission according to claim 2, characterized in that The step of adjusting the channel weights of the input data according to the channel attention mechanism to obtain first intermediate data further includes: Performing third convolution processing, first global average pooling, fully connected layer dimensionality reduction, activation, fully connected layer dimensionality increase, second global average pooling, and channel weighting in sequence; Wherein the output of the third convolution processing is also superimposed on the output data of the second global average pooling through a residual connection.

4. The lightweight semantic communication method for image transmission according to claim 1, characterized in that The step of obtaining the encoded data is in a sequential processing mode, and further includes: obtaining a scaling factor according to the size difference between the input data of the current downsampling and the size of the target image, and obtaining the shifting operation step size of the external shifted convolution according to the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling.

5. The lightweight semantic communication method for image transmission according to claim 1, wherein Also including: Concatenating the first feature map into the input data of the first atrous spatial pyramid pooling.

6. A lightweight semantic communication system for image transmission, characterized in that, Including: An encoder, including a first convolution module, a downsampling module, and a first atrous spatial pyramid pooling module connected in sequence, for performing first convolution processing, downsampling, and first atrous spatial pyramid pooling on the target image in sequence to obtain encoded data; A decoder, including an upsampling module, a second convolution module, and a second atrous spatial pyramid pooling module connected in sequence, for performing upsampling, second convolution processing, and second atrous spatial pyramid pooling on the received encoded data in sequence to obtain a reconstructed image of the target image; Wherein the encoder further includes an external shifted convolution unit, and the external shifted convolution unit is used to perform external shifted convolution on the target image to obtain a first feature map, the feature size of the first feature map being consistent with the feature size of the downsampling; The downsampling module includes a feature splicing unit, an inner shift convolution unit, and a non-linear activation unit connected in sequence. Among them, the feature splicing is used to splice the first feature map into the input data of the downsampling module; The outer shift convolution unit is also connected to the output end of the inner shift convolution unit to superimpose the first feature map onto the output data of the inner shift convolution unit through a residual connection; Both the outer shift convolution unit and the inner shift convolution unit include a shift layer, a batch normalization layer, and a pointwise convolution layer connected in sequence; A plurality of the downsampling modules are continuously arranged.

7. The lightweight semantic communication system for image transmission according to claim 6, wherein The downsampling module further includes: A squeeze-and-excitation unit for adjusting the channel weights of the input data of the downsampling according to the channel attention mechanism to obtain first intermediate data; Among them, the feature splicing unit is used to splice the first feature map into the first intermediate data to obtain second intermediate data, and use the second intermediate data as the input data of the inner shift convolution unit.

8. The lightweight semantic communication system for image transmission according to claim 7, characterized in that The squeeze-and-excitation unit includes: A third convolution layer, a first global average pooling layer, a fully connected layer dimensionality reduction layer, an activation layer, a fully connected layer dimensionality increase layer, a second global average pooling layer, and a channel weighting layer connected in sequence; Among them, the output of the third convolution layer is also superimposed onto the output data of the second global average pooling layer through a residual connection.

9. The lightweight semantic communication system for image transmission according to claim 6, wherein The working mode of the encoder is a sequential processing mode. The outer shift convolution unit is further used to: obtain a scaling factor according to the size difference between the input data of the current downsampling module and the size of the target image, and obtain the shift operation step size of the outer shift convolution unit according to the scaling factor, so that the size of the obtained first feature map is consistent with the size of the input data of the current downsampling module.

10. The lightweight semantic communication system for image transmission according to claim 6, characterized in that, The first feature map is also spliced into the input data of the first atrous spatial pyramid pooling module.

Citation Information

Patent Citations

  • Skeleton-based shift graph convolutional network human behavior identification method

    CN114463840A

  • Multi-bandwidth separation feature extraction convolutional layer of convolutional neural network

    CN116324811A

  • Lightweight image semantic segmentation method based on spatial shift and convolution

    CN117011527A

  • Real-time semantic segmentation method based on attention and multi-scale feature extraction

    CN117115435A

  • Global-local cooperation lightweight image super-resolution method based on semantic guidance

    CN117314753A