Semantic communication encoding transmission method and device based on sliding window, and equipment

CN117544276BActive Publication Date: 2026-08-21BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311338100.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-16
Publication Date
2026-08-21
Estimated Expiration
2043-10-16

AI Technical Summary

Technical Problem

当前联合编码方法,当信源维度增加时,基于深度学习的信源信道联合编码的模型容量不足以完全地感知信源分布,进而使得语义信息的传输准确率降低,系统性能会严重下降

Benefits of technology

[0018]从上述可以看出,本公开提出了一种基于滑动窗口的语义通信编码传输方法、装置及设备,通过利用滑动窗口编码器对获取到的第一信源图像进行编码处理,得到其对应的第一语义层次表征向量,所述第一语义层次表征向量表征了所述第一信源图像的语义特征,以供后续对抗信道噪声和衰落的影响。根据获取到的目标信道的噪声比,基于信道调制网络对所述第一语义层次表征向量进行信道编码,得到第一信道调制隐式表征向量,而后利用速率调制网络进行速率调节,得到第一信道输入向量,以适根据目标速率实现适时调整。对上述得到的第一信道输入向量进行功率归一化处理,得到第一功率归一化信道输入向量,以适应目标信道,并输入至所述目标信道中进行传输。通过使第一语义层次表征向量分别经过信道调制网络及速率调节网络,提供了图像传输的信噪比自适应策略和速率自适应策略,提高了信源语义信息的感知度,实现了对信源语义信息的全局感知。通过信源语义编码及信道编码的双重编码,将经过编码后的第一信道传输向量传输至信道中,以供其他接收端接收所述第一信道传输向量,识别对应的语义信息,提高了语义信息的传输准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117544276B_ABST
    Figure CN117544276B_ABST
Patent Text Reader

Abstract

The present disclosure provides a sliding window-based semantic communication encoding transmission method, device and equipment, comprising: acquiring a first source image, using a sliding window encoder to encode and process the first source image to obtain a first semantic level representation vector of the first source image; acquiring a signal-to-noise ratio of a target channel, using the channel modulation network to encode and process the first semantic level representation vector according to the signal-to-noise ratio to obtain a first channel modulation implicit representation vector; acquiring a target rate, adjusting the transmission rate of the first channel modulation implicit representation vector through the rate modulation network according to the target rate to obtain a first channel input vector; performing power normalization processing on the first channel input vector to obtain a first power normalized channel input vector, and inputting the first power normalized channel input vector into the target channel. The present disclosure realizes global perception of source semantic information and improves the transmission accuracy of semantic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of semantic communication technology, and in particular to a semantic communication encoding and transmission method, apparatus and device based on a sliding window. Background Technology

[0002] Semantic communication has become a new direction driving the development of information and communication technologies in recent years, and is also a hot topic of innovation in the field of artificial intelligence. In traditional communication, the transmission of information sources only considers syntactic information. Semantic communication, by extracting and measuring semantic information from information sources such as images, transmits semantic information oriented towards downstream tasks, thus achieving intelligent and simplified communication.

[0003] Traditional communication systems have been designed and improved based on the Shannon separation principle, but achieving optimal performance in practical systems is difficult, thus driving the development of joint coding techniques. Source-channel joint coding is a classic topic in information theory and coding theory. Current joint coding methods, when the source dimension increases, suffer from insufficient model capacity in deep learning-based source-channel joint coding to fully perceive the source distribution, leading to reduced accuracy in semantic information transmission and a significant degrade in system performance.

[0004] Therefore, how to achieve global perception of semantic information of the information source and improve the transmission accuracy of semantic information has become an important research problem. Summary of the Invention

[0005] In view of this, the purpose of this disclosure is to propose a semantic communication encoding and transmission method, apparatus and device based on sliding window, so as to solve or partially solve the above problems.

[0006] To achieve the above objectives, a first aspect of this disclosure provides a semantic communication encoding and transmission method based on a sliding window, the method comprising:

[0007] A first source image is acquired, and the first source image is encoded using a sliding window encoder to obtain a first semantic level representation vector of the first source image.

[0008] The signal-to-noise ratio (SNR) of the target channel is obtained. Based on the SNR, the first semantic level representation vector is encoded using the channel modulation network to obtain the first channel modulation implicit representation vector. The channel modulation network is an adaptive network trained to encode the semantic level representation vector based on the SNR.

[0009] Obtain the target rate, and adjust the transmission rate of the first channel modulation implicit representation vector through the rate modulation network according to the target rate to obtain the first channel input vector. The rate modulation network is a network obtained by training an adaptive network to adjust the transmission rate of the channel modulation implicit representation vector according to the target rate.

[0010] The first channel input vector is subjected to power normalization processing to obtain a first power normalized channel input vector, and the first power normalized channel input vector is input to the target channel.

[0011] Based on the same inventive concept, a second aspect of this disclosure proposes a semantic communication encoding and transmission apparatus based on a sliding window, comprising:

[0012] The image acquisition module is configured to acquire a first source image, encode the first source image using a sliding window encoder, and obtain a first semantic level representation vector of the first source image.

[0013] The signal-to-noise ratio (SNR) acquisition module is configured to acquire the SNR of the target channel, and based on the SNR, encode the first semantic level representation vector using the channel modulation network to obtain the first channel modulation implicit representation vector, wherein the channel modulation network is an adaptive network trained to encode the semantic level representation vector based on the SNR.

[0014] The rate adjustment module is configured to acquire a target rate and adjust the transmission rate of the first channel modulation implicit representation vector through a rate modulation network according to the target rate to obtain a first channel input vector, wherein the rate modulation network is a network trained on an adaptive network to adjust the transmission rate of the channel modulation implicit representation vector according to the target rate.

[0015] The normalization processing module is configured to perform power normalization processing on the first channel input vector to obtain a first power normalized channel input vector, and input the first power normalized channel input vector to the target channel.

[0016] Based on the same inventive concept, a third aspect of this disclosure proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor, when executing the computer program, implements the semantic communication encoding and transmission method based on a sliding window as described above.

[0017] Based on the same inventive concept, a fourth aspect of this disclosure proposes a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the sliding window-based semantic communication encoding transmission method as described above.

[0018] As can be seen from the above, this disclosure proposes a semantic communication coding and transmission method, apparatus, and device based on a sliding window. By using a sliding window encoder to encode the acquired first source image, a corresponding first semantic level representation vector is obtained. This first semantic level representation vector represents the semantic features of the first source image, which is then used to combat the effects of channel noise and fading. Based on the noise ratio of the acquired target channel, channel coding is performed on the first semantic level representation vector using a channel modulation network to obtain a first channel modulation implicit representation vector. Then, a rate modulation network is used for rate adjustment to obtain a first channel input vector, which is adjusted in real time according to the target rate. The obtained first channel input vector is then subjected to power normalization to obtain a first power-normalized channel input vector, which is adapted to the target channel and input into the target channel for transmission. By passing the first semantic level representation vector through both the channel modulation network and the rate adjustment network, an adaptive signal-to-noise ratio strategy and a rate adaptive strategy for image transmission are provided, improving the perception of source semantic information and achieving global perception of source semantic information. By employing dual encoding of source semantic coding and channel coding, the encoded first channel transmission vector is transmitted into the channel so that other receiving ends can receive the first channel transmission vector, identify the corresponding semantic information, and improve the transmission accuracy of semantic information. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of a semantic communication encoding and transmission method based on a sliding window, according to an embodiment of the present disclosure.

[0021] Figure 2 This is a structural block diagram of a semantic communication encoding and transmission apparatus based on a sliding window according to an embodiment of the present disclosure;

[0022] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0024] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0025] The following are definitions of terms used in this disclosure:

[0026] JSCC: Joint source-channel coding (JSCC) is a classic topic in information theory and coding theory.

[0027] CNN: Convolutional Neural Network (CNN) has representation learning capabilities and can perform translation-invariant classification of input information according to its hierarchical structure.

[0028] SwinJSCC: Swin transformer for deep joint source-channel coding (SwinJSCC).

[0029] Swin Transformer: A sliding window transformer (Transformer using shifted windows, Swin Transformer).

[0030] Channel ModNet: Channel modulation network.

[0031] Rate ModNet; Rate Modulation Network.

[0032] CSI: Channel state information.

[0033] DIV2K: DIV2K is a single-image super-resolution dataset that can be used to reconstruct high-resolution images from low-resolution images.

[0034] Based on the above description, this embodiment proposes a semantic communication encoding and transmission method based on a sliding window, such as... Figure 1 As shown, the method includes:

[0035] Step 101: Obtain the first source image, and encode the first source image using a sliding window encoder to obtain the first semantic level representation vector of the first source image.

[0036] In specific implementation, a first source image that needs to be transmitted is acquired, and the first source image is encoded using a trained sliding window encoder to perform source encoding, thereby obtaining a first semantic level representation vector corresponding to the first source image. The first semantic level representation vector is a vector that represents the semantic features of the first source image.

[0037] A sliding window encoder is obtained by pre-training the initial encoder. In this embodiment, the training set uses the DIV2K public dataset, and the validation is performed using the Kodak and CLIC (Deep Learning Image Compression Challenge) 2021 public datasets. In this embodiment, the resolution of all images is a multiple of 128 to avoid padding of the neural encoder and decoder.

[0038] Step 102: Obtain the signal-to-noise ratio (SNR) of the target channel. Based on the SNR, use the channel modulation network to encode the first semantic level representation vector to obtain the first channel modulation implicit representation vector. The channel modulation network is an adaptive network trained to encode the semantic level representation vector based on the SNR.

[0039] In practice, an adaptive network is trained to obtain a channel modulation network, thereby enhancing its adaptability to the constantly changing characteristics of the transmission channel. The signal-to-noise ratio (SNR) of the input target channel is obtained, and the first semantic-level representation vector is encoded using the channel modulation network to obtain a first channel modulation implicit representation vector. This first channel modulation implicit representation vector is obtained by channel coding the first semantic-level representation vector, and it is adapted to the SNR of the target channel.

[0040] Step 103: Obtain the target rate, and adjust the transmission rate of the first channel modulation implicit representation vector according to the target rate through the rate modulation network to obtain the first channel input vector. The rate modulation network is a network obtained by training an adaptive network to adjust the transmission rate of the channel modulation implicit representation vector according to the target rate.

[0041] In practice, an adaptive network is trained to obtain a rate modulation network, which can automatically adjust in real time to adapt to any target rate. The target rate is obtained, and based on the target rate, the rate of the implicit representation vector of the first channel modulation is adjusted using the rate modulation network to obtain a first channel input vector, wherein the rate of the first input vector is adapted to the target rate.

[0042] Step 104: Perform power normalization processing on the first channel input vector to obtain a first power normalized channel input vector, and input the first power normalized channel input vector to the target channel.

[0043] In practice, the first channel input vector obtained from the above steps is power normalized to obtain the first power normalized channel input vector, so that different modulation methods can obtain the same average power, which facilitates the comparison of system performance.

[0044] After the first power-normalized channel input vector is transmitted to the target channel, it is processed by the target channel to obtain the first implicit noise representation vector. This fully considers channel fading and simulates real channel transmission conditions.

[0045] The first implicit noise representation vector is expressed by the formula:

[0046]

[0047] in, This is the first implicit noise representation vector. Let be the first power-normalized channel input vector, and n be the noise vector, where each element in the noise vector is an element independently sampled from a Gaussian distribution. in denoted as average noise power, and h as channel state information vector.

[0048] In the above scheme, when steps 101 to 104 are applied to the source as the transmitter, the acquired first source image is encoded using a sliding window encoder to obtain its corresponding first semantic level representation vector. This first semantic level representation vector represents the semantic features of the first source image, which is then used to combat the effects of channel noise and fading. Based on the noise ratio of the target channel, the first semantic level representation vector is channel-coded using a channel modulation network to obtain a first channel modulation implicit representation vector. Then, a rate modulation network is used for rate adjustment to obtain a first channel input vector, which is adjusted in real time according to the target rate. The obtained first channel input vector is then power-normalized to obtain a first power-normalized channel input vector, which is adapted to the target channel and input into the target channel for transmission. By passing the first semantic level representation vector through both the channel modulation network and the rate adjustment network, an adaptive signal-to-noise ratio strategy and a rate adaptive strategy for image transmission are provided, improving the perception of source semantic information and achieving global perception of source semantic information. By employing dual encoding of source semantic coding and channel coding, the encoded first channel transmission vector is transmitted into the channel so that other receiving ends can receive the first channel transmission vector, identify the corresponding semantic information, and improve the transmission accuracy of semantic information.

[0049] In some embodiments, the sliding window encoder includes a slice embedding layer, at least one sliding window transformer module, and a slice merging layer. Step 101 specifically includes:

[0050] Step 1011: Segment the first source image to obtain a first image slice.

[0051] In specific implementation, the sliding window encoder includes a slice embedding layer, at least one sliding window transformer module, and a slice merging layer. The acquired first source image is segmented; in this embodiment, the first source image is an RGB (red, green, blue) image.

[0052] For example, the first source image is x. Where H is the height of the pixels in the first source image, and W is the width of the pixels in the first source image. The first source image is segmented to obtain first image slices, the number of pixels in the first image slice being... The size is C1. The elements are sorted in image space from left to right and from top to bottom.

[0053] Step 1012: Input the first image slice into the slice embedding layer of the sliding window encoder to obtain the first embedding vector.

[0054] In specific implementation, the first image slice obtained from the above segmentation is input into the slice embedding layer of the sliding window encoder. Based on the above example, the first embedding vector is obtained.

[0055] Step 1013: Input the first embedding vector into at least one sliding window transformer module of the sliding window encoder to obtain the first mapping vector.

[0056] In specific implementation, based on the above example, the first embedding vector x is... e The input is fed into the sliding window transformer module. For example, the sliding window transformer module contains N1 sub-modules to obtain a first mapping vector, which is denoted as...

[0057] Step 1014: Input the first mapping vector into the slice merging layer of the sliding window encoder, and obtain the first semantic level representation vector after processing by the slice merging layer.

[0058] In specific implementation, the first mapping vector is sliced ​​in a subspace, and the slice merging layer is used to map the sliced ​​first mapping vector, merging them to obtain the first semantic level representation vector.

[0059] In some embodiments, step 1014 specifically includes:

[0060] Step 10141: The first mapping vector is downsampled using a slice merging layer to obtain the second embedding vector.

[0061] Step 10142: Input the second embedding vector into at least one sliding window transformer module of the sliding window encoder to obtain the second mapping vector.

[0062] In specific implementation, the first mapping vector is downsampled through the slice merging layer in the sliding window encoder to obtain the second embedding vector, wherein the second embedding vector is represented as... in The second embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the second mapping vector. For example, the sliding window converter module contains N2 sub-modules.

[0063] Step 10143: The second mapping vector is downsampled using the slice merging layer to obtain the third embedding vector;

[0064] Step 10144: Input the third embedding vector into at least one sliding window transformer module of the sliding window encoder to obtain the third mapping vector.

[0065] In specific implementation, the second mapping vector is downsampled through the slice merging layer in the sliding window encoder to obtain the third embedding vector, wherein the third embedding vector is represented as... in The third embedding vector is input into the sliding window transformer module to obtain the third mapping vector. For example, the sliding window converter module contains N3 sub-modules.

[0066] Step 10445: The third mapping vector is downsampled using a slice merging layer to obtain the target embedding vector;

[0067] Step 10446: Input the target embedding vector into at least one sliding window transformer module of the sliding window encoder to obtain the first semantic level representation vector.

[0068] In specific implementation, the third mapping vector is downsampled through the slice merging layer in the sliding window encoder to obtain the target embedding vector, wherein the target embedding vector is represented as... in The target embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the first semantic level representation vector y. ′ , For example, the sliding window converter module contains N4 sub-modules.

[0069] In some embodiments, the channel modulation network includes at least one fully connected channel layer, and step 102 specifically includes:

[0070] Step 1021: The first semantic level representation vector is sequentially input into the at least one fully connected channel layer for encoding processing to obtain the first target channel dot product result.

[0071] In specific implementation, the channel modulation network includes at least one signal-to-noise ratio (SNR) modulation module and at least one fully connected channel layer. Each SNR modulation module includes three fully connected network layers. The first semantic level representation vector is sequentially input into each fully connected layer of the channel modulation network, and a dot product is performed to obtain the first target channel dot product result.

[0072] The adaptive network is pre-trained to obtain the channel modulation network. In this embodiment, the training set uses data from the DIV2K public dataset, and the validation is performed using data from the Kodak and CLIC2021 public datasets.

[0073] Step 1022: Input the first target channel dot product result into the first activation function to obtain the first activation value.

[0074] Step 1023: Perform a dot product between the first activation value and the first semantic level representation vector to obtain the first channel modulation implicit representation vector.

[0075] In specific implementation, the first target channel dot product result is input into the first activation function, and the first activation function performs calculations to obtain the first activation value. The first activation value is then multiplied by the first semantic level representation vector to obtain the first channel modulation implicit representation vector.

[0076] In some embodiments, when the number of fully connected channel layers is multiple, step 1021 specifically includes:

[0077] Step 10211: Determine the signal-to-noise ratio modulation vector based on the signal-to-noise ratio, input the first semantic level representation vector into the first channel fully connected layer, and perform a dot product with the signal-to-noise ratio modulation vector to obtain the first channel dot product result.

[0078] In specific implementation, the signal-to-noise ratio (SNR) modulation module in the channel modulation network uses the channel's SNR as input to calculate the SNR modulation vector. The first semantic level representation vector is then input to the first fully connected channel layer and multiplied with the SNR modulation vector to obtain the first channel dot product result. This first channel dot product result is then used as input to the second fully connected channel layer for subsequent processing.

[0079] Step 10212: Use the first channel dot product result as the input of the next channel fully connected layer, and repeat the above dot product calculation process until all channel fully connected layers are passed to obtain the first target channel dot product result.

[0080] In a specific implementation, for example, the channel modulation network includes eight fully connected layers. The first channel dot product result is input to the second fully connected layer and multiplied with the signal-to-noise ratio vector to obtain the second channel dot product result. The above steps are repeated until the eighth channel dot product result is obtained, and the eighth channel dot product result is used as the first target channel dot product result.

[0081] The above scheme enables semantic communication coding transmission to adapt to the signal-to-noise ratio of the channel, allowing transmission to adapt to various channel conditions and ensuring normal transmission.

[0082] In some embodiments, the rate modulation network includes at least one rate fully connected layer and a coding mask module, and step 103 specifically includes:

[0083] Step 1031: Obtain the target rate, and input the first channel modulation implicit representation vector into the at least one rate fully connected layer for adjustment processing to obtain the target rate dot product result.

[0084] In practice, the rate modulation network includes at least one rate fully connected layer, at least one rate modulation module, and a coding mask module. Each rate modulation module includes three fully connected layers. The implicit representation vector of the first channel modulation is sequentially input into each rate fully connected layer to obtain the target rate dot product result.

[0085] The rate modulation network is obtained by pre-training the adaptive network. In this embodiment, the training set uses data from the DIV2K public dataset, and the validation is performed using data from the Kodak and CLIC2021 public datasets.

[0086] Step 1032: Input the target rate dot product result into the second activation function to obtain the second activation value, and use the second activation value as the rate characterization value.

[0087] Step 1033: Perform a dot product between the rate representation value and the first channel modulation implicit representation vector to obtain the relevant modulation embedding vector.

[0088] In specific implementation, the target rate dot product result is input into a second activation function, and a second activation value is obtained through calculation by the second activation function. This second activation value is used as the rate representation value. The rate representation value is then multiplied by the first channel modulation implicit representation vector to obtain a related modulation embedding vector, which is used for subsequent calculations to obtain the first channel input vector for the target channel.

[0089] Step 1034: Input the rate characterization value into the encoding mask module to obtain the first initial auxiliary information mask;

[0090] Step 1035: Perform entropy encoding on the first initial auxiliary information mask to generate and output the first auxiliary information mask, so that other information sources can receive the first auxiliary information mask and reconstruct the first information source image based on the first auxiliary information mask.

[0091] In specific implementation, the rate characterization value is input to the coding mask module in the rate modulation network to obtain a first initial auxiliary information mask. Entropy coding is performed on the first initial auxiliary information mask to generate a first auxiliary information mask. The first auxiliary information mask is output so that other information sources can receive it and reconstruct the first source image based on it.

[0092] Step 1036: Perform a dot product between the first initial auxiliary information mask and the related modulation embedding vector to obtain the first channel input vector.

[0093] In specific implementation, the first initial auxiliary information mask is multiplied by the relevant modulation embedding vector to obtain the first channel input vector. Subsequently, the first channel input vector is input into the target channel so that the receiving end can receive the first channel input vector and obtain the reconstructed source image based on the first channel input vector.

[0094] In some embodiments, step 1031 specifically includes:

[0095] Step 10311: Input the first channel modulation implicit representation vector into the first rate fully connected layer and perform a dot product with the target rate to obtain the first rate dot product result.

[0096] In practice, the rate modulation module in the rate modulation network uses the target rate as input to calculate the rate vector. The first channel modulation implicit representation vector is then input to the first rate fully connected layer, where it is multiplied by the rate vector to obtain a first rate multiplication result. This first rate multiplication result is then used as input to the second rate fully connected layer for subsequent processing.

[0097] Step 10312: Use the first rate dot product result as the input of the next rate fully connected layer, and repeat the above dot product calculation process until the target rate dot product result is obtained after passing through all rate fully connected layers.

[0098] In a specific implementation, for example, the rate modulation network includes eight fully connected layers. The first rate dot product result is input into the second rate fully connected layer, and a dot product is performed with the rate vector to obtain the second rate dot product result. The above steps are repeated until the eighth rate dot product result is obtained, and the eighth rate dot product result is used as the target rate dot product result.

[0099] The above scheme adapts semantic communication encoding and transmission to the available bandwidth of the channel, adjusts the transmission rate in real time, avoids transmission problems caused by excessive or insufficient available bandwidth, ensures reliable and consistent transmission, and improves robustness in dynamic wireless communication scenarios.

[0100] In some embodiments, the method further includes:

[0101] Step 10A: Receive the second noise implicit representation vector and the second auxiliary information mask sent by other information sources in the target channel.

[0102] In specific implementation, the information source, which is the transmitting end, can also be the receiving end to receive information. When the information source is the receiving end, it receives the second noise implicit representation vector and the second auxiliary information mask sent by other information sources in the target channel, so as to perform image reconstruction based on the second noise implicit representation vector and the second auxiliary information mask.

[0103] Step 10B: Perform entropy decoding on the second auxiliary information mask to obtain the second initial auxiliary information mask.

[0104] In practice, since the second auxiliary information mask is first entropy encoded before transmission, after receiving the second auxiliary information mask, entropy decoding is performed to obtain the second initial auxiliary information mask corresponding to the second auxiliary information mask.

[0105] Step 10C: Use the second initial auxiliary information mask to perform feature reconstruction on the second noise implicit representation vector to obtain the reconstructed second semantic level representation vector.

[0106] In specific implementation, the second noise implicit representation vector is masked by zero-value filling using the second initial auxiliary information mask, so as to realize the reconstruction of the second semantic level representation vector and ensure the consistency between the reconstructed second semantic level representation vector and the second noise implicit representation vector.

[0107] Step 10D: Input the reconstructed second semantic level representation vector into the channel modulation network, and use the channel modulation network to demodulate the second semantic level representation vector to obtain the second channel modulation implicit representation vector.

[0108] Step 10E: The second channel modulation implicit representation vector is decoded using a sliding window decoder to obtain the reconstructed second source image.

[0109] In practice, the reconstructed second semantic level representation vector is input into the channel modulation network for channel decoding to obtain the second channel modulation implicit representation vector. Then, a sliding window decoder is used to perform source decoding on the second channel modulation implicit representation vector to obtain the reconstructed second source image.

[0110] In some embodiments, step 10D specifically includes:

[0111] Step 10D1: The reconstructed second semantic level representation vector is sequentially input into the at least one fully connected channel layer for encoding processing to obtain the second target channel dot product result.

[0112] In specific implementation, the channel modulation network includes at least one signal-to-noise ratio (SNR) modulation module and at least one fully connected channel layer. Each SNR modulation module includes three fully connected network layers. The reconstructed second semantic level representation vector is sequentially input into each fully connected layer of the channel modulation network, and a dot product is performed to obtain the second target channel dot product result.

[0113] Step 10D2: Input the dot product result of the second target channel into the third activation function to obtain the third activation value.

[0114] Step 10D3: Perform a dot product between the third activation value and the reconstructed second semantic level representation vector to obtain the second channel modulation implicit representation vector.

[0115] In specific implementation, the dot product result of the second target channel is input into the third activation function, and the third activation value is obtained by the operation of the third activation function. The third activation value is then multiplied by the reconstructed second semantic level representation vector to obtain the second channel modulation implicit representation vector.

[0116] In some embodiments, when the number of fully connected channel layers is multiple, step 10D1 specifically includes:

[0117] Step 10D11: Determine the signal-to-noise ratio modulation vector based on the signal-to-noise ratio, input the reconstructed second semantic level representation vector into the first channel fully connected layer, and perform a dot product with the signal-to-noise ratio modulation vector to obtain the first channel reconstruction dot product result.

[0118] In specific implementation, the signal-to-noise ratio (SNR) modulation module in the channel modulation network uses the channel's SNR as input to calculate the SNR modulation vector. The reconstructed second semantic level representation vector is then input to the first fully connected channel layer and multiplied with the SNR modulation vector to obtain the first channel reconstruction dot product result. This first channel reconstruction dot product result is then used as input to the second fully connected channel layer for subsequent processing.

[0119] Step 10D12: Use the first channel reconstruction dot product result as the input of the next channel fully connected layer, and repeat the above dot product calculation process until all channel fully connected layers are passed to obtain the second target channel dot product result.

[0120] In a specific implementation, for example, the channel modulation network includes eight fully connected layers. The first channel reconstruction dot product result is input into the second fully connected channel layer, and a dot product is performed with the signal-to-noise ratio vector to obtain the second channel reconstruction dot product result. The above steps are repeated until the eighth channel reconstruction dot product result is obtained, and the eighth channel reconstruction dot product result is used as the second target channel dot product result.

[0121] The above scheme enables semantic communication coding transmission to adapt to the signal-to-noise ratio of the channel, allowing transmission to adapt to various channel conditions and ensuring normal transmission.

[0122] In some embodiments, the sliding window decoder includes a slice segmentation layer and at least one sliding window transformer module, and step 10E specifically includes:

[0123] Step 10E1: Input the second channel modulation implicit representation vector into the sliding window transformer module of the sliding window decoder to obtain the fourth mapping vector.

[0124] Step 10E2: Input the fourth mapping vector into the slice segmentation layer of the sliding window decoder, and use the slice segmentation layer to upsample the fourth mapping vector to obtain the target output vector.

[0125] Step 10E3: Merge the image slices corresponding to the target output vector to obtain the reconstructed second source image.

[0126] In practical implementation, the implicit representation vector of the second channel modulation is: in The second channel modulation implicit representation vector is input into the sliding window transformer module of the sliding window decoder to obtain the fourth mapping vector. For example, the sliding window converter module contains N4 sub-modules.

[0127] The fourth mapping vector is upsampled using a slice separation layer to obtain the target output vector, denoted as x″. The image slices corresponding to the target output vector are merged to obtain the reconstructed second source image.

[0128] In some embodiments, step 10E2 specifically includes:

[0129] Step 10E21: Upsample the fourth mapping vector using a slice segmentation layer to obtain the fifth embedding vector.

[0130] Step 10E22: Input the fifth embedding vector into the sliding window transformer module of the sliding window decoder to obtain the fifth mapping vector.

[0131] Step 10E23: Upsample the fifth mapping vector using the slice segmentation layer to obtain the sixth embedding vector.

[0132] In specific implementation, the fourth mapping vector is upsampled using a slice separation layer to obtain a fifth embedding vector, wherein the fifth embedding vector is represented as... The fifth embedding vector is input into the sliding window transformer module of the sliding window decoder to obtain the fifth mapping vector, wherein the fifth mapping vector is represented as... For example, the sliding window converter module contains N3 sub-modules.

[0133] The fifth mapping vector is upsampled using a slice segmentation layer to obtain a sixth embedding vector, wherein the sixth embedding vector is represented as...

[0134] Step 10E24: Input the sixth embedding vector into the sliding window transformer module of the sliding window decoder to obtain the sixth mapping vector.

[0135] Step 10E25: Upsample the sixth mapping vector using a slice segmentation layer to obtain the target output vector.

[0136] In specific implementation, the sixth embedding vector is input into the sliding window transformer module of the sliding window decoder to obtain the sixth mapping vector, wherein the sixth mapping vector is represented as For example, the sliding window converter module contains N2 sub-modules.

[0137] The sixth mapping vector is upsampled using a slice segmentation layer to obtain the target output vector.

[0138] In some embodiments, a loss function is used to improve the quality of image reconstruction during the training of the initial encoder, channel modulation network, and rate modulation network. The loss function is expressed by the formula:

[0139]

[0140] Where φ and θ respectively contain f e with f d All network parameters, f e For encoder, f e It consists of a sliding window encoder, a channel modulation network, and a rate tuning network, f d For decoder, f d It consists of a sliding window decoder and a channel modulation network.

[0141] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0142] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0143] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a semantic communication encoding and transmission device based on a sliding window.

[0144] refer to Figure 2 , Figure 2 The sliding window-based semantic communication encoding and transmission apparatus, as described in this embodiment, includes:

[0145] The image acquisition module 201 is configured to acquire a first source image, encode the first source image using a sliding window encoder, and obtain a first semantic level representation vector of the first source image.

[0146] The signal-to-noise ratio (SNR) acquisition module 202 is configured to acquire the SNR of the target channel, and based on the SNR, encode the first semantic level representation vector using the channel modulation network to obtain a first channel modulation implicit representation vector, wherein the channel modulation network is an adaptive network trained to encode the semantic level representation vector based on the SNR.

[0147] The rate adjustment module 203 is configured to acquire a target rate and adjust the transmission rate of the first channel modulation implicit representation vector through a rate modulation network according to the target rate to obtain a first channel input vector, wherein the rate modulation network is a network trained on an adaptive network to adjust the transmission rate of the channel modulation implicit representation vector according to the target rate.

[0148] The normalization processing module 204 is configured to perform power normalization processing on the first channel input vector to obtain a first power normalized channel input vector, and input the first power normalized channel input vector to the target channel.

[0149] In some embodiments, the sliding window encoder includes a slice embedding layer, at least one sliding window transformer module, and a slice merging layer, and the image acquisition module 201 specifically includes:

[0150] An image segmentation unit is configured to segment the first source image to obtain a first image slice;

[0151] The first embedding vector determination unit is configured to input the first image slice into the slice embedding layer of the sliding window encoder to obtain the first embedding vector;

[0152] The first mapping vector determination unit is configured to input the first embedded vector into at least one sliding window transformer module of the sliding window encoder to obtain the first mapping vector;

[0153] The first semantic level representation vector determination unit is configured to input the first mapping vector into the slice merging layer of the sliding window encoder, and obtain the first semantic level representation vector through the slice merging layer.

[0154] In some embodiments, the first semantic hierarchy representation vector determination unit is specifically configured as follows:

[0155] The first mapping vector is downsampled using a slice merging layer to obtain the second embedding vector;

[0156] The second embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the second mapping vector;

[0157] The second mapping vector is downsampled using the slice merging layer to obtain the third embedding vector;

[0158] The third embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the third mapping vector;

[0159] The third mapping vector is downsampled using a slice merging layer to obtain the target embedding vector;

[0160] The target embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the first semantic level representation vector.

[0161] In some embodiments, the channel modulation network includes at least one fully connected channel layer, and the signal-to-noise ratio acquisition module 202 specifically includes:

[0162] The channel coding unit is configured to sequentially input the first semantic level representation vector into the at least one fully connected channel layer for encoding processing to obtain the first target channel dot product result.

[0163] The activation value determination unit is configured to input the dot product result of the first target channel into the first activation function to obtain the first activation value;

[0164] The dot product calculation unit is configured to perform a dot product operation on the first activation value and the first semantic level representation vector to obtain the first channel modulation implicit representation vector.

[0165] In some embodiments, in response to the presence of multiple fully connected channel layers, the channel coding unit is specifically configured as follows:

[0166] Based on the signal-to-noise ratio, a signal-to-noise ratio modulation vector is determined. The first semantic level representation vector is input to the first channel fully connected layer and multiplied with the signal-to-noise ratio modulation vector to obtain the first channel dot product result.

[0167] The first channel dot product result is used as the input to the next fully connected channel layer. The above dot product calculation process is repeated until all fully connected channels are passed to obtain the target channel dot product result.

[0168] In some embodiments, the rate modulation network includes at least one rate fully connected layer and a coding mask module, and the rate adjustment module 203 specifically includes:

[0169] A rate adjustment unit is configured to acquire a target rate by sequentially inputting the first channel modulation implicit representation vector into the at least one rate fully connected layer for adjustment processing to obtain the target rate dot product result.

[0170] The rate characterization value determination unit is configured to input the target rate dot product result into a second activation function to obtain a second activation value, and use the second activation value as the rate characterization value.

[0171] The dot product calculation unit is configured to perform a dot product operation on the rate representation value and the first channel modulation implicit representation vector to obtain the relevant modulation embedding vector.

[0172] The encoding mask unit is configured to input the rate characterization value into the encoding mask module to obtain a first initial auxiliary information mask;

[0173] The entropy coding unit is configured to perform entropy coding processing on the first initial auxiliary information mask, generate and output the first auxiliary information mask, so that other information sources can receive the first auxiliary information mask and reconstruct the first information source image based on the first auxiliary information mask;

[0174] The channel input vector determination unit is configured to perform a dot product of the first initial auxiliary information mask and the related modulation embedding vector to obtain the first channel input vector.

[0175] In some embodiments, the apparatus further includes a source reconstruction module, which specifically includes:

[0176] The vector receiving unit is configured to receive a second noise implicit representation vector and a second auxiliary information mask transmitted by other information sources in the target channel;

[0177] The entropy decoding unit is configured to perform entropy decoding processing on the second auxiliary information mask to obtain the second initial auxiliary information mask;

[0178] The feature reconstruction unit is configured to use the second initial auxiliary information mask to perform feature reconstruction on the second noise implicit representation vector to obtain the reconstructed second semantic level representation vector.

[0179] The demodulation unit is configured to input the reconstructed second semantic level representation vector into the channel modulation network, and use the channel modulation network to demodulate the second semantic level representation vector to obtain the second channel modulation implicit representation vector.

[0180] The decoding unit is configured to use a sliding window decoder to decode the implicit representation vector of the second channel modulation to obtain the reconstructed second source image.

[0181] In some embodiments, the sliding window decoder includes a slice segmentation layer and at least one sliding window transformer module, and the decoding unit is specifically configured as follows:

[0182] The second channel modulation implicit representation vector is input into the sliding window transformer module of the sliding window decoder to obtain the fourth mapping vector;

[0183] The fourth mapping vector is input to the slice segmentation layer of the sliding window decoder, and the slice segmentation layer is used to upsample the fourth mapping vector to obtain the target output vector.

[0184] The image slices corresponding to the target output vector are merged to obtain the reconstructed second source image.

[0185] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0186] The apparatus of the above embodiments is used to implement the corresponding semantic communication encoding and transmission method based on sliding window in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0187] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the semantic communication encoding and transmission method based on a sliding window as described in any of the above embodiments.

[0188] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0189] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0190] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0191] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0192] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0193] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0194] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0195] The electronic devices described above are used to implement the corresponding sliding window-based semantic communication encoding and transmission methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0196] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the semantic communication encoding and transmission method based on a sliding window as described in any of the above embodiments.

[0197] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0198] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the semantic communication encoding and transmission method based on sliding windows as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0199] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0200] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0201] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0202] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A semantic communication encoding and transmission method based on a sliding window, characterized in that, include: A first source image is acquired, and the first source image is encoded using a sliding window encoder to obtain a first semantic level representation vector of the first source image. The signal-to-noise ratio (SNR) of the target channel is obtained. Based on the SNR, the first semantic level representation vector is encoded using a channel modulation network to obtain a first channel modulation implicit representation vector. The channel modulation network is an adaptive network trained to encode the semantic level representation vector based on the SNR. Obtain the target rate, and adjust the transmission rate of the first channel modulation implicit representation vector through the rate modulation network according to the target rate to obtain the first channel input vector. The rate modulation network is a network obtained by training an adaptive network to adjust the transmission rate of the channel modulation implicit representation vector according to the target rate. The first channel input vector is subjected to power normalization processing to obtain a first power normalized channel input vector, and the first power normalized channel input vector is input to the target channel; The channel modulation network includes at least one fully connected channel layer. The step of encoding the first semantic level representation vector using a channel modulation network based on the signal-to-noise ratio to obtain the first channel modulation implicit representation vector includes: The first semantic level representation vector is sequentially input into the at least one fully connected channel layer for encoding processing to obtain the first target channel dot product result; The result of the first target channel dot product is input into the first activation function to obtain the first activation value; The first activation value is multiplied by the first semantic level representation vector to obtain the first channel modulation implicit representation vector. In response to the fact that there are multiple fully connected layers in the channel, The step of sequentially inputting the first semantic level representation vector into the at least one fully connected channel layer for encoding processing to obtain the target channel dot product result includes: Based on the signal-to-noise ratio, a signal-to-noise ratio modulation vector is determined. The first semantic level representation vector is input to the first channel fully connected layer and multiplied with the signal-to-noise ratio modulation vector to obtain the first channel multiplication result. The first channel dot product result is used as the input to the next fully connected channel layer. The above dot product calculation process is repeated until all fully connected channel layers are passed to obtain the target channel dot product result. The rate modulation network includes at least one rate fully connected layer and a coding mask module. The process of obtaining the target rate, which involves adjusting the transmission rate of the first channel modulation implicit representation vector through a rate modulation network based on the target rate to obtain the first channel input vector, includes: To obtain the target rate, the first channel modulation implicit representation vector is sequentially input into the at least one rate fully connected layer for adjustment processing to obtain the target rate dot product result. The result of the dot product of the target rate is input into the second activation function to obtain the second activation value, and the second activation value is used as the rate characterization value; The rate representation value is multiplied by the first channel modulation implicit representation vector to obtain the relevant modulation embedding vector. The rate characterization value is input into the encoding mask module to obtain the first initial auxiliary information mask; The first initial auxiliary information mask is entropy encoded to generate and output a first auxiliary information mask for other information sources to receive. The first information source image is then reconstructed based on the first auxiliary information mask. The first initial auxiliary information mask is multiplied by the relevant modulation embedding vector to obtain the first channel input vector. The sliding window encoder includes a slice embedding layer, at least one sliding window transformer module, and a slice merging layer. The step of encoding the first source image using a sliding window encoder to obtain a first semantic level representation vector of the first source image includes: The first source image is segmented to obtain a first image slice; The first image slice is input into the slice embedding layer of the sliding window encoder to obtain the first embedding vector; The first embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the first mapping vector; The first mapping vector is input to the slice merging layer of the sliding window encoder, and the first semantic level representation vector is obtained after processing by the slice merging layer. The step of inputting the first mapping vector into the slice merging layer of the sliding window encoder, and obtaining the first semantic level representation vector after processing by the slice merging layer, includes: The first mapping vector is downsampled using a slice merging layer to obtain the second embedding vector; The second embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the second mapping vector; The second mapping vector is downsampled using the slice merging layer to obtain the third embedding vector; The third embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the third mapping vector; The third mapping vector is downsampled using a slice merging layer to obtain the target embedding vector; The target embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the first semantic level representation vector.

2. The method according to claim 1, characterized in that, Also includes: Receive the second implicit noise representation vector and the second auxiliary information mask sent by other information sources in the target channel; The second auxiliary information mask is subjected to entropy decoding to obtain the second initial auxiliary information mask; The second noise implicit representation vector is reconstructed using the second initial auxiliary information mask to obtain the reconstructed second semantic level representation vector. The reconstructed second semantic level representation vector is input into the channel modulation network, and the channel modulation network is used to demodulate the second semantic level representation vector to obtain the second channel modulation implicit representation vector. The second channel modulation implicit representation vector is decoded using a sliding window decoder to obtain the reconstructed second source image.

3. The method according to claim 2, characterized in that, The sliding window decoder includes a slice segmentation layer and at least one sliding window transformer module. The step of decoding the implicit representation vector of the second channel modulation using a sliding window decoder to obtain the reconstructed second source image includes: The second channel modulation implicit representation vector is input into the sliding window transformer module of the sliding window decoder to obtain the fourth mapping vector; The fourth mapping vector is input to the slice segmentation layer of the sliding window decoder, and the slice segmentation layer is used to upsample the fourth mapping vector to obtain the target output vector. The image slices corresponding to the target output vector are merged to obtain the reconstructed second source image.

4. A semantic communication encoding and transmission device based on a sliding window, characterized in that, include: The image acquisition module is configured to acquire a first source image, encode the first source image using a sliding window encoder, and obtain a first semantic level representation vector of the first source image. The signal-to-noise ratio (SNR) acquisition module is configured to acquire the SNR of the target channel, and based on the SNR, encode the first semantic level representation vector using a channel modulation network to obtain a first channel modulation implicit representation vector, wherein the channel modulation network is an adaptive network trained to encode the semantic level representation vector based on the SNR. The rate adjustment module is configured to acquire a target rate and adjust the transmission rate of the first channel modulation implicit representation vector through a rate modulation network according to the target rate to obtain a first channel input vector, wherein the rate modulation network is a network trained on an adaptive network to adjust the transmission rate of the channel modulation implicit representation vector according to the target rate. The normalization processing module is configured to perform power normalization processing on the first channel input vector to obtain a first power normalized channel input vector, and input the first power normalized channel input vector to the target channel; The channel modulation network includes at least one fully connected channel layer. The step of encoding the first semantic level representation vector using a channel modulation network based on the signal-to-noise ratio to obtain the first channel modulation implicit representation vector includes: The first semantic level representation vector is sequentially input into the at least one fully connected channel layer for encoding processing to obtain the first target channel dot product result; The result of the first target channel dot product is input into the first activation function to obtain the first activation value; The first activation value is multiplied by the first semantic level representation vector to obtain the first channel modulation implicit representation vector. In response to the fact that there are multiple fully connected layers in the channel, The step of sequentially inputting the first semantic level representation vector into the at least one fully connected channel layer for encoding processing to obtain the target channel dot product result includes: Based on the signal-to-noise ratio, a signal-to-noise ratio modulation vector is determined. The first semantic level representation vector is input to the first channel fully connected layer and multiplied with the signal-to-noise ratio modulation vector to obtain the first channel multiplication result. The first channel dot product result is used as the input to the next fully connected channel layer. The above dot product calculation process is repeated until all fully connected channel layers are passed to obtain the target channel dot product result. The rate modulation network includes at least one rate fully connected layer and a coding mask module. The process of obtaining the target rate, which involves adjusting the transmission rate of the first channel modulation implicit representation vector through a rate modulation network based on the target rate to obtain the first channel input vector, includes: To obtain the target rate, the first channel modulation implicit representation vector is sequentially input into the at least one rate fully connected layer for adjustment processing to obtain the target rate dot product result. The result of the dot product of the target rate is input into the second activation function to obtain the second activation value, and the second activation value is used as the rate characterization value; The rate representation value is multiplied by the first channel modulation implicit representation vector to obtain the relevant modulation embedding vector. The rate characterization value is input into the encoding mask module to obtain the first initial auxiliary information mask; The first initial auxiliary information mask is entropy encoded to generate and output a first auxiliary information mask for other information sources to receive. The first information source image is then reconstructed based on the first auxiliary information mask. The first initial auxiliary information mask is multiplied by the relevant modulation embedding vector to obtain the first channel input vector. The sliding window encoder includes a slice embedding layer, at least one sliding window transformer module, and a slice merging layer. The step of encoding the first source image using a sliding window encoder to obtain a first semantic level representation vector of the first source image includes: The first source image is segmented to obtain a first image slice; The first image slice is input into the slice embedding layer of the sliding window encoder to obtain the first embedding vector; The first embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the first mapping vector; The first mapping vector is input to the slice merging layer of the sliding window encoder, and the first semantic level representation vector is obtained after processing by the slice merging layer. The step of inputting the first mapping vector into the slice merging layer of the sliding window encoder, and obtaining the first semantic level representation vector after processing by the slice merging layer, includes: The first mapping vector is downsampled using a slice merging layer to obtain the second embedding vector; The second embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the second mapping vector; The second mapping vector is downsampled using the slice merging layer to obtain the third embedding vector; The third embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the third mapping vector; The third mapping vector is downsampled using a slice merging layer to obtain the target embedding vector; The target embedding vector is input into at least one sliding window transformer module of the sliding window encoder to obtain the first semantic level representation vector.

5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the semantic communication encoding and transmission method based on a sliding window as described in any one of claims 1 to 3.