Semantic communication method and apparatus, device and storage medium
By comparing learning, training semantic encoder and decoder, the problem of semantic communication with large semantic distance in fading channels is solved, and higher semantic communication effect and semantic accuracy are achieved.
Patent Information
- Application Number
- PCT/CN2023/136428
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-12-05
- Publication Date
- 2025-05-08
AI Technical Summary
The existing semantic communication methods cannot effectively reduce the semantic distance between the original model and the reconstruction model in the fading channel, which affects the effect of semantic communication.
The semantic encoder and semantic decoder are trained using contrast learning. By selecting the target training sample and its enhanced sample in the training sample set, the semantic contrast loss function is determined to reduce the semantic distance between the original model and the reconstructed model.
It effectively reduces the semantic distance between the original model and the reconstruction model, ensures that the transmitted information can better maintain its semantic accuracy in downstream tasks, and greatly improves the effect of semantic communication.
Smart Images

Figure CN2023136428_08052025_PF_FP_ABST
Abstract
Description
A semantic communication method, device, equipment and storage medium Technical Field
[0001] The present invention relates to the field of semantic communication technology, and in particular to a semantic communication method, apparatus, device and storage medium. Background Art
[0002] In recent years, semantic communication has attracted increasing research interest as a promising approach to efficient transmission in wireless network models. Compared to the traditional communication paradigm based on Shannon theory, semantic communication prioritizes the preservation of meaningful semantic information over the accuracy of transmitted symbols. This can significantly reduce the amount of transmitted data and improve communication efficiency.
[0003] In semantic communication involving model transmission, the primary challenge is how to effectively extract the model's semantic information at the transmitter and accurately reconstruct it at the receiver under limited communication conditions. Existing semantic communication methods extract the semantic information of the input model through deep learning-based semantic encoding at the transmitter, followed by corresponding deep semantic decoding at the receiver to reconstruct the model's semantic information. This significantly reduces the amount of data communicated without compromising the communication of semantic information.
[0004] During model transmission, since the original model at the transmitter and the reconstructed model at the receiver share the same semantic information, the semantic distance (also known as semantic similarity) between two nearly identical models sharing the same semantic information should be sufficiently small. However, existing semantic communication methods, especially over fading channels, often suffer from a large semantic distance between two nearly identical models sharing the same semantic information. This makes it impossible to effectively reduce the semantic distance between the original and reconstructed models, preventing the transmitted information from maintaining its semantic accuracy in downstream tasks, thus hindering the effectiveness of semantic communication.
[0005] Summary of the Invention
[0006] The present invention provides a semantic communication method, apparatus, device and storage medium to solve the technical problem in existing semantic communication methods that the semantic distance between the original model and the reconstructed model is large and the semantic accuracy of the reconstructed model cannot be ensured.
[0007] In order to solve the above technical problems, an embodiment of the present invention provides a semantic communication method, including:
[0008] Inputting the target model to be transmitted into a preset semantic encoder, so that the semantic encoder extracts the semantic information of the target model and outputs the semantic information of the target model;
[0009] Transmitting the semantic information of the target model to a preset semantic decoder through a wireless channel, so that the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model;
[0010] The semantic encoder and the semantic decoder are pre-trained by contrastive learning;
[0011] When training the semantic encoder and the semantic decoder, obtaining a training sample set, and selecting a training sample from the training sample set as a target training sample;
[0012] Using the model damage of the training sample in the wireless channel as data enhancement, determining a first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample, and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample;
[0013] Determining a semantic contrast loss function for the training sample set with the goal of minimizing the first semantic distance and maximizing the second semantic distance;
[0014] The semantic encoder and the semantic decoder are trained according to the training sample set and the semantic contrast loss function.
[0015] As a preferred solution, extracting the semantic information of the target model and outputting the semantic information of the target model includes:
[0016] Extracting semantic information of the target model and performing nonlinear mapping on the extracted semantic information to generate a k-dim complex-valued vector;
[0017] Performing power normalization processing on the k-dim complex-valued vector, and outputting semantic information of the target model for transmission in a wireless channel;
[0018] The extracted semantic information is nonlinearly mapped using the following formula:
[0019] in, is a k-dim complex-valued vector, represents the semantic encoding operation of parameter θ1, and x is the target model;
[0020] The k-dim complex value vector is power normalized by the following formula:
[0021] Where s is the semantic information of the target model used for transmission in the wireless channel, k is the channel bandwidth of the wireless channel, P is the average power of the semantic information transmitter, and * represents the conjugate transpose.
[0022] As a preferred solution, the target model is reconstructed according to the semantic information of the target model, and the reconstructed model corresponding to the target model is output, including:
[0023] Perform semantic decoding on the semantic information of the target model to obtain the reconstructed model corresponding to the target model;
[0024] The semantic information of the target model is semantically decoded using the following formula:
[0025] in, is the reconstructed model corresponding to the target model, represents the semantic decoding operation of parameter θ2, It is the semantic information of the target model received by the semantic information receiving end.
[0026] As a preferred solution, the semantic encoder is a convolutional neural network model;
[0027] The convolutional neural network model includes: a convolution, several downsampling modules, and a channel coding module; wherein the downsampling module includes: a ResBolck module and a convolution for downsampling the target model; the ResBolck module is the basic module in ResNet.
[0028] As a preferred solution, the semantic decoder is a convolutional neural network model;
[0029] The convolutional neural network model includes: a convolution head, several upsampling modules, and a re-encoding module; wherein, the upsampling module includes: a ResBolck module and a Pixel-Shuffle module for upsampling the target model; the ResBolck module is the basic module in ResNet.
[0030] As a preferred solution, determining the first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample, and determining the second semantic distance between the remaining training samples in the training sample set and the target training sample include:
[0031] Inputting a target training sample into the semantic encoder so that the semantic encoder extracts semantic information of the target training sample and outputs the semantic information of the target training sample;
[0032] Transmitting the semantic information of the target training sample to the semantic decoder via a wireless channel, so that the semantic decoder reconstructs the target training sample according to the semantic information of the target training sample, outputs a reconstructed sample model corresponding to the target training sample, and uses the reconstructed sample model as an enhanced sample of the target training sample;
[0033] Generate a feature map of the target training sample, a feature map of the enhanced sample, and feature maps of the remaining training samples in the training sample set;
[0034] Mapping the feature map of the target training sample into a preset semantic space through a preset projection network to obtain a first projection result, mapping the feature map of the enhanced sample into the semantic space to obtain a second projection result, and mapping the feature maps of the remaining training samples in the training sample set into the semantic space to obtain a third projection result;
[0035] Determining a first semantic distance between a target training sample and an enhanced sample corresponding to the target training sample according to a cosine similarity between the first projection result and the second projection result;
[0036] A second semantic distance between the remaining training samples in the training sample set and the target training sample is determined according to the cosine similarity between the first projection result and the third projection result.
[0037] As a preferred solution, the training of the semantic encoder and the semantic decoder according to the training sample set and the semantic contrast loss function includes:
[0038] Performing first-stage training and second-stage training on the semantic encoder and the semantic decoder according to the training sample set and the semantic contrast loss function;
[0039] The first phase of training includes:
[0040] Determining a model reconstruction loss function according to a reconstruction loss between a training sample in the training sample set and a reconstruction sample model corresponding to each training sample;
[0041] Determining a first loss function according to the model reconstruction loss function and the semantic contrast loss function, and then performing a first-stage training on the semantic encoder and the semantic decoder according to the first loss function;
[0042] The first loss function is:
[0043] Among them, α1∈[0,1] is a hyperparameter that controls the trade-off between the model reconstruction loss function and the semantic contrast loss function. Reconstruct the loss function for the model, is the semantic contrast loss function;
[0044] The second phase of training includes:
[0045] Determine the downstream task loss function based on the model damage in the wireless channel;
[0046] Determining a second loss function according to the downstream task loss function and the model reconstruction loss function, and then performing a second-stage training on the semantic encoder and the semantic decoder according to the second loss function;
[0047] The second loss function is:
[0048] Among them, α2∈[0,1] is a hyperparameter that controls the trade-off between the downstream task loss function and the model reconstruction loss function. Reconstruct the loss function for the model, is the loss function for the downstream task.
[0049] Based on the above embodiment, another embodiment of the present invention provides a semantic communication device, including: a semantic encoding module, a semantic decoding module and a model training module;
[0050] The semantic encoding module is used to input the target model to be transmitted into a preset semantic encoder, so that the semantic encoder extracts the semantic information of the target model and outputs the semantic information of the target model;
[0051] The semantic decoding module is used to transmit the semantic information of the target model to a preset semantic decoder through a wireless channel, so that the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model;
[0052] The model training module is used to train the semantic encoder and semantic decoder using contrastive learning;
[0053] When training the semantic encoder and the semantic decoder, obtaining a training sample set, and selecting a training sample from the training sample set as a target training sample;
[0054] Using the model damage of the training sample in the wireless channel as data enhancement, determining a first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample, and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample;
[0055] Determining a semantic contrast loss function for the training sample set with the goal of minimizing the first semantic distance and maximizing the second semantic distance;
[0056] The semantic encoder and the semantic decoder are trained according to the training sample set and the semantic contrast loss function.
[0057] Based on the above embodiments, another embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the semantic communication method based on contrastive learning described in the above invention embodiment.
[0058] Based on the above embodiments, another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the semantic communication method based on contrastive learning described in the above invention embodiment.
[0059] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0060] In the present invention, a target model to be transmitted is input into a preset semantic encoder so that the semantic encoder outputs the semantic information of the target model; the semantic information of the target model is transmitted to a preset semantic decoder through a wireless channel so that the semantic decoder outputs a reconstructed model corresponding to the target model; wherein the semantic encoder and the semantic decoder are pre-trained by contrastive learning; when training the semantic encoder and the semantic decoder, a training sample is selected from the training sample set as the target training sample; the model damage of the training sample in the wireless channel is used as data enhancement, and the first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample is determined, and the second semantic distance between the remaining training samples and the target training sample is determined; with the first semantic distance being minimized and the second semantic distance being maximized as the goal, the semantic contrast loss function of the training sample set is determined; according to the semantic contrast loss function, the semantic encoder and the semantic decoder are trained.
[0061] The present invention trains a semantic encoder and a semantic decoder by contrastive learning. According to the target, the first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample is minimized, and the second semantic distance between the remaining training samples and the target training sample is maximized. The semantic encoder and the semantic decoder are then trained according to the semantic contrastive loss function. This can effectively reduce the semantic distance between the original model and the reconstructed model, that is, the semantic similarity, to ensure that the transmitted information can better maintain its semantic accuracy in downstream tasks, thereby greatly improving the effect of semantic communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] FIG1 is a flow chart of a semantic communication method provided by one embodiment of the present invention;
[0063] FIG2 is a model architecture diagram of a semantic encoder and a semantic decoder of the present invention;
[0064] FIG3 is a framework diagram of semantic contrast coding of the present invention;
[0065] FIG4 is a comparison chart of the accuracy of the present invention and other methods when the bandwidth compression ratio is different but other parameters are the same;
[0066] FIG5 is a comparison diagram of the peak signal-to-noise ratio of the present invention and other methods when the bandwidth compression ratio is different but other parameters are the same;
[0067] FIG6 is a comparison chart of the accuracy of the present invention and other methods when the bandwidth compression ratio is different and under low channel conditions and other parameters are consistent;
[0068] FIG7 is a comparison diagram of PSNR of the present invention and other methods when bandwidth compression ratios are different and under low channel conditions and other parameter conditions are the same;
[0069] FIG8 is a visual comparison diagram of the present invention and other methods provided on the Kodak data set;
[0070] FIG9 is a schematic structural diagram of a semantic communication device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0072] Example 1
[0073] Please refer to FIG1 , which is a flowchart of a semantic communication method provided by an embodiment of the present invention, including the following specific steps:
[0074] S1. Inputting the target model to be transmitted into a preset semantic encoder, so that the semantic encoder extracts the semantic information of the target model and outputs the semantic information of the target model;
[0075] The present invention first deploys a semantic encoder based on contrastive learning (CL) and convolutional neural network (CNN) at the semantic information sending end, and a semantic decoder based on contrastive learning (CL) and convolutional neural network (CNN) at the semantic information receiving end. The semantic encoder extracts the semantic information of the input target model and outputs the semantic information of the target model.
[0076] S2. Transmitting the semantic information of the target model to a preset semantic decoder via a wireless channel, so that the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model;
[0077] The semantic encoder and the semantic decoder are pre-trained by contrastive learning. When training the semantic encoder and the semantic decoder, a training sample set is obtained, and a training sample is selected from the training sample set as a target training sample. Model damage of the training sample in the wireless channel is used as data enhancement to determine a first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample, and a second semantic distance between the remaining training samples in the training sample set and the target training sample is determined. With the goal of minimizing the first semantic distance and maximizing the second semantic distance, a semantic contrast loss function of the training sample set is determined. According to the training sample set and the semantic contrast loss function, the semantic encoder and the semantic decoder are trained.
[0078] After the semantic encoder outputs the semantic information of the target model, wireless transmission is performed through a wireless fading channel to transmit the semantic information of the target model to the semantic decoder, and then the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model.
[0079] However, in the above-mentioned process of model semantic information extraction, model transmission and model reconstruction, the semantic distance between the original model and the reconstructed model in the prior art is often very large, and the semantic accuracy of the reconstructed model cannot be ensured, thereby affecting the effect of semantic communication.
[0080] To reduce the semantic distance between the original model and the reconstructed model and improve the semantic accuracy of the reconstructed model, the present invention establishes a semantic communication system for wireless large-scale model transmission, in which a semantic encoder and a semantic decoder based on a convolutional neural network (CNN) are deployed in the transmitter and receiver, respectively. The deployment of the semantic encoder and semantic decoder and the construction of the semantic communication system are completed through the following main parts:
[0081] Preferably, extracting the semantic information of the target model and outputting the semantic information of the target model includes: extracting the semantic information of the target model and performing nonlinear mapping on the extracted semantic information to generate a k-dim complex-valued vector; performing power normalization processing on the k-dim complex-valued vector and outputting the semantic information of the target model for transmission in a wireless channel;
[0082] The extracted semantic information is nonlinearly mapped using the following formula:
[0083] in, is a k-dim complex-valued vector, represents the semantic encoding operation of parameter θ1, and x is the target model;
[0084] The k-dim complex value vector is power normalized by the following formula:
[0085] Where s is the semantic information of the target model used for transmission in the wireless channel, k is the channel bandwidth of the wireless channel, P is the average power of the semantic information transmitter, and * represents the conjugate transpose.
[0086] Preferably, reconstructing the target model according to the semantic information of the target model and outputting the reconstructed model corresponding to the target model includes: semantically decoding the semantic information of the target model to obtain the reconstructed model corresponding to the target model;
[0087] The semantic information of the target model is semantically decoded using the following formula:
[0088] in, is the reconstructed model corresponding to the target model, represents the semantic decoding operation of parameter θ2, It is the semantic information of the target model received by the semantic information receiving end.
[0089] 1. Deployment of Semantic Encoder and Semantic Decoder
[0090] (1) Semantic decoder: The semantic encoder is used to extract the input model x∈R c×h×wits semantic information and directly implement the non-linear mapping from the semantic information to a k-dim complex-valued vector which is given by the following formula:
[0091] where represents the semantic encoding operation of parameter θ1 is a k-dim complex-valued vector, x is the target model, and c, h, and w represent the number of channels, height, and width of the model, respectively.
[0092] For simplicity, we use n = c×h×w to represent the dimension of x. Usually, k < n should be satisfied to meet the bandwidth constraint, and k / n is called the bandwidth compression ratio. In particular, a large bandwidth compression ratio indicates good communication conditions, while a small bandwidth compression ratio indicates limited bandwidth usage. In addition, a power normalization layer is required at the end of the semantic encoding network to meet the average power constraint at the transmitter The power normalization operation can be written as:
[0093] where s is the channel input signal that satisfies the power constraint, is the semantic information of the target model for transmission in the wireless channel, k is the channel bandwidth of the wireless channel, P is the average power at the semantic information transmitter, and * represents the conjugate transpose.
[0094] Next, s will be transmitted through an additive white Gaussian (AWGN) channel, which is given by the following formula:
[0095] where is the received signal, ∈∈C k represents independent and identically distributed (IID) channel noise samples, following a symmetric complex Gaussian distribution with mean and variance σ 2 .
[0096] (2) Semantic encoder: The semantic decoder deployed at the receiver will reconstruct the original model from The reconstruction process is as follows: where
[0097] where is the reconstructed model corresponding to the target model8] represents the semantic decoding operation of parameter θ2 is the semantic information of the target model received at the semantic information receiver.
[0098] Subsequently will be used to perform downstream tasks and obtain the inference result through the following process:
[0099] in, The parameter φ1 is used to represent the feature extraction operation performed by the CNN backbone of the downstream task, f x ={f (1) ,f (2) ,…f (C)} is the output feature map with C channels. x Passed to the classifier You can get the inference result With parameter φ2, it can be expressed as:
[0100] Maintaining semantic information in the reconstructed model is crucial for model inference performance, especially when the channel bandwidth is limited. Therefore, designing the semantic encoder and semantic decoder according to the above process and subsequently training them is crucial in the present invention.
[0101] Preferably, the semantic encoder is a convolutional neural network model; the convolutional neural network model includes: a convolution, several downsampling modules, and a channel coding module; wherein the downsampling module includes: a ResBolck module and a convolution for downsampling the target model; the ResBolck module is the basic module in ResNet.
[0102] Preferably, the semantic decoder is a convolutional neural network model; the convolutional neural network model includes: a convolution head, several upsampling modules, and a re-encoding module; wherein the upsampling module includes: a ResBolck module and a Pixel-Shuffle module for upsampling the target model; the ResBolck module is the basic module in ResNet.
[0103] 2. Construction of semantic communication framework:
[0104] In this paper, we propose a semantic communication framework based on CL. Because the network architecture of the semantic encoder and semantic decoder plays a crucial role in extracting semantic information, we do not use the existing method of stacking convolutional layers because this simple architecture lacks this capability.
[0105] Please refer to Figure 2, which shows the model architecture of the semantic encoder and semantic decoder of the present invention. The semantic encoder includes a 5×5 convolutional layer, two downsampling modules, and a channel coding module. Each downsampling module contains a basic block from ResNet (which we call ResBolck) to capture the spatial features of the model, and a 4×4 convolution with a stride of 2 to downsample the model. The channel coding module is used to mitigate channel impairments and output a k-dim complex-valued channel input that meets bandwidth and power constraints.
[0106] In the semantic decoder, we use a symmetric architecture consisting of a 5×5 convolutional layer, two upsampling modules, and a re-encoding module. In the up-sampling module, ResBolcks are used, as in the encoder. Here, we employ the pixel-shuffle technique to upsample the model, as it offers a more efficient computational paradigm and better reconstruction performance than transposed convolutions. The re-encoding module consists of a 3×3 convolution with a sigmoid activation function to generate the reconstructed model. Note that all convolutions are subjected to batch normalization and the parametric rectified linear unit (PReLU) activation function, unless otherwise specified.
[0107] Preferably, the determining of the first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample, and the determining of the second semantic distance between the remaining training samples in the training sample set and the target training sample include: inputting the target training sample into the semantic encoder so that the semantic encoder extracts the semantic information of the target training sample and outputs the semantic information of the target training sample; transmitting the semantic information of the target training sample to the semantic decoder through a wireless channel so that the semantic decoder reconstructs the target training sample according to the semantic information of the target training sample, outputs a reconstructed sample model corresponding to the target training sample, and uses the reconstructed sample model as the enhanced sample of the target training sample; generating a feature map of the target training sample, the enhanced sample model, and the like; The method comprises the following steps: mapping the feature map of the target training sample to a preset semantic space through a preset projection network to obtain a first projection result, mapping the feature map of the enhanced sample to the semantic space to obtain a second projection result, and mapping the feature maps of the remaining training samples in the training sample set to the semantic space to obtain a third projection result; determining a first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample based on the cosine similarity between the first projection result and the second projection result; and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample based on the cosine similarity between the first projection result and the third projection result.
[0108] 3. Determination of semantic contrast loss function
[0109] The key design of semantic contrastive encoding is inspired by the success of CL, which uses data augmentation to generate samples with similar visual representations and minimize the distance between them to pre-train the backbone. Therefore, in this paper, the CL process can be modified to adapt it to the semantic communication system. We replace data augmentation with the wireless transmission process because the model corruption that occurs during the transmission process can be regarded as a form of data augmentation. The semantic distance between the original model and the reconstructed model should be kept small for an efficient semantic communication system.
[0110] In addition, we also utilize the pre-trained backbone to extract features and use a learnable projection network to map these features into the semantic space. By incorporating the contrastive loss in the semantic space, we jointly optimize the semantic encoder and decoder instead of pre-training the backbone in CL.
[0111] Please refer to Figure 3, which is a framework diagram of the semantic contrast coding of the present invention. The process first obtains a training sample set exist Select a target training sample model x for semantic encoding and decoding, from which we can obtain the reconstructed It is used as the enhanced sample model of the target training sample. Backbone network Applied to x and Generate corresponding feature maps respectively and
[0112] After that, we introduce a projection head The above feature maps are mapped to the semantic space, which is achieved through a learnable multi-layer perception. Next, a fully connected projection network with a learnable parameter ψ is constructed. And the subsequent normalization operation maps the features into a semantic space defined as a hypersphere. During the training phase, you can update To enhance the understanding of features, and thus learn the mapping from features to semantics. Specifically, f x and The projection result can be expressed as and where q x called anchor, v + We can use the cosine similarity between anchor and positive to define x and The semantic distance between them.
[0113] The present invention is for training sample set The rest of the training samples in The same process will be followed. We obtain its corresponding feature map by inputting m into the backbone network Then use f m Projected into the semantic space, where we have v m Similarly, the semantic distance between x and m can be defined as the cosine similarity between the anchor and the negative.
[0114] The goal of semantic contrast encoding is to minimize the semantic distance between the original model and the reconstructed model, while maximizing the semantic distance between the original model and the disrespectful model. Therefore, we can use the InfoNCE function to define the training sample set The semantic contrast loss can be expressed as
[0115] where τ>0 is the temperature coefficient used to smooth the probability distribution.
[0116] Preferably, the training of the semantic encoder and the semantic decoder according to the training sample set and the semantic contrast loss function includes: performing first-stage training and second-stage training on the semantic encoder and the semantic decoder according to the training sample set and the semantic contrast loss function; wherein the first-stage training includes: determining a model reconstruction loss function according to the reconstruction loss between the training samples in the training sample set and the reconstruction sample models corresponding to each training sample; determining a first loss function according to the model reconstruction loss function and the semantic contrast loss function, and then performing first-stage training on the semantic encoder and the semantic decoder according to the first loss function;
[0117] The first loss function is:
[0118] Among them, α1∈[0,1] is a hyperparameter that controls the trade-off between the model reconstruction loss function and the semantic contrast loss function. Reconstruct the loss function for the model, is the semantic contrast loss function;
[0119] The second stage training includes: determining a downstream task loss function based on model damage in the wireless channel; determining a second loss function based on the downstream task loss function and the model reconstruction loss function, and then performing a second stage training on the semantic encoder and the semantic decoder based on the second loss function;
[0120] The second loss function is:
[0121] Among them, α2∈[0,1] is a hyperparameter that controls the trade-off between the downstream task loss function and the model reconstruction loss function. Reconstruct the loss function for the model, is the loss function for the downstream task.
[0122] 4. Training process of semantic encoder and semantic decoder
[0123] Next, we introduce how to consider semantic contrastive encoding and semantic contrastive loss to design the loss function and training process. Based on semantic contrastive encoding, we design a two-stage training strategy to optimize the semantic encoder and decoder.
[0124] (1) First training stage: The first stage is pre-training. We use the semantic contrastive encoding method to train the weights of the encoder θ1, decoder θ2 and item network ψ at the same time. However, when we only optimize the semantic contrastive loss, it is difficult to achieve a fast convergence speed. Therefore, we combine it with x and Combined with the reconstruction loss between, reducing the reconstruction loss can help improve the convergence speed of early training. Specifically, we can apply the mean square error (MSE) function to evaluate the training sample set The reconstruction loss can be expressed as
[0125] Therefore, the loss function of the first training stage can be summarized as a linear combination given by:
[0126] Among them, α1∈[0,1] is a hyperparameter that controls the trade-off between the model reconstruction loss function and the semantic contrast loss function. Reconstruct the loss function for the model, is the semantic contrast loss function. For example, in a real semantic communication system, we can set α = k / n. In this case, when the bandwidth compression is small, the system prioritizes the preservation of semantic information over the quality of reconstruction; on the contrary, as the bandwidth compression increases, the system shifts its focus to maintaining the quality of reconstruction.
[0127] (1) Second training stage: In the second training stage, our goal is to further optimize the performance of the semantic communication system by jointly fine-tuning the encoder, decoder, and classifier with a small learning rate to achieve considerable inference performance and reconstruction model quality. One of the reasons for fine-tuning the classifier is that the weights of the backbone network and the classifier are usually trained without considering channel impairments, which causes the output of the backbone network to be the reconstructed model after adding the input instead of the original model. This may lead to performance degradation. Therefore, fine-tuning the classifier using the semantic encoder and decoder can alleviate this problem and help enhance semantic transmission. The loss function of this stage can be expressed as:
[0128] Where α2∈[0,1] is a hyperparameter that controls the trade-off between the downstream task loss function and the model reconstruction loss function. Reconstruct the loss function for the model, is the downstream task loss function. Specifically, when the downstream task is a classification problem, the cross entropy function can be used to model the loss, which is given by
[0129] Among them, y i and Represents the true value and predicted probability of the i-th category respectively. Symbol N cls Indicates the number of classes in the dataset.
[0130] Figure 2 shows the model architecture of the semantic encoder and semantic decoder proposed in this invention. In a specific embodiment, to verify the effectiveness of the proposed framework, we conducted experiments on CIFAR-10, which contains 60,000 32×32 color models divided into 10 categories. The training set contains 50,000 models, while the test set contains 10,000 models. The projection network adopts a two-layer fully connected structure with an output size of 32. The number of training epochs for the two replaced stages is set to 200 and 100, respectively, with a batch size of 128. In addition, we use the Adam optimizer, with a learning rate of 0.01 in the first pre-training stage and 0.0001 in the second fine-tuning stage. These learning rates are adjusted every 50 epochs with a decay factor of 0.5.
[0131] In order to verify the performance of the semantic communication framework proposed in this invention, the following comparative experiments are conducted to illustrate:
[0132] (1) Please refer to Figure 4, which is a comparison chart of the accuracy of the proposed method and other methods in the Python simulation environment when the bandwidth compression ratio is different and other parameter conditions are consistent. As can be seen from Figure 4, the accuracy of the proposed method is always better than or matches other existing methods. These results show that the semantic communication framework proposed in the present invention can effectively extract semantic information to meet the requirements of downstream tasks and remove irrelevant redundant information to ensure that semantic information can be successfully transmitted, especially when the channel bandwidth is limited.
[0133] (2) Please refer to Figure 5, which is a comparison chart of the peak signal-to-noise ratio (PSNR) of the proposed method and other methods in the Python simulation environment when the bandwidth compression ratio k / n is different and other parameter conditions are consistent. From Figure 5, we can see that with the increase of the bandwidth compression ratio, the PSNR of all methods is improved. Although the method proposed in the present invention sacrifices some model quality to prioritize semantic information when the bandwidth compression ratio is low, it can quickly catch up with the PSNR of DeepJSCC at higher compression ratios. These results show that the method can prioritize the transmission of semantic information rather than irrelevant background information to ensure the performance of downstream tasks in bandwidth-limited scenarios, while transmitting enough background information to obtain good model quality when bandwidth is not an obstacle, which further proves the effectiveness of the method.
[0134] (3) Please refer to Figure 6, which is a comparison chart of the accuracy of the proposed method and other methods in the Python simulation environment, when the bandwidth compression ratio is different and the channel condition is low, and other parameter conditions are consistent. It describes the performance comparison of these methods in terms of accuracy under poor channel conditions. It considers a low SNR of 5dB and the bandwidth compression ratio varies from 1 / 48 to 1 / 2.5. From Figure 6, we can observe that compared with the other three methods, the method proposed in the present invention still shows an advantage in accuracy, indicating its robustness in low SNR scenarios.
[0135] (4) Please refer to Figure 7, which is a PSNR comparison chart of the proposed method and other methods in the Python simulation environment, when the bandwidth compression ratio is different and the channel condition is low, and other parameter conditions are consistent. From Figure 7, we can see that when the bandwidth compression ratio is low, the framework proposed by the present invention can adaptively sacrifice global information to obtain comparable semantic performance, while achieving sufficient reconstruction quality in terms of PSNR as the bandwidth compression ratio decreases. These results in Figure 7 further verify the effectiveness and robustness of the proposed method in low SNR scenarios.
[0136] (5) Please refer to Figure 8, which is a visual comparison diagram of the proposed method and several other methods provided by the present invention on the Kodak dataset in a Python simulation environment, where the encoder and decoder are trained on the STL10 dataset. From this figure, we can observe that the proposed method removes redundant background information while retaining the main semantic information compared to the competing methods, thereby reducing model damage in semantic areas (e.g., the macaw and rafters in the figure). In addition, the method achieves similar PSNR and multi-scale structural similarity (MS-SSIM) performance as DeepJSCC and DeepJSCC-ft, indicating the effectiveness of the method in reconstructing semantic information. These results further demonstrate the superiority of the proposed method in achieving leading accuracy in downstream tasks compared to the comparison methods.
[0137] It can be seen that the present invention provides a semantic communication method. Through the present invention, the semantic distance between the original model and the reconstructed model, that is, the semantic similarity, can be effectively reduced, ensuring that the transmitted information can better maintain its semantic accuracy in downstream tasks, thereby greatly improving the effect of semantic communication. Through the contrastive learning method, better model inference performance is provided while maintaining communication efficiency. The problem of semantic information loss is overcome by contrastive learning to ensure that the transmitted information can better maintain its semantic accuracy in downstream tasks. A two-stage training process and corresponding loss function are adopted to achieve a good trade-off between model recognition performance and reconstruction quality in downstream tasks, thereby providing a more flexible communication solution.
[0138] Example 2
[0139] Please refer to FIG9 , which is a schematic diagram of the structure of a semantic communication device provided by an embodiment of the present invention, wherein the device includes: a semantic encoding module, a semantic decoding module, and a model training module;
[0140] The semantic encoding module is used to input the target model to be transmitted into a preset semantic encoder, so that the semantic encoder extracts the semantic information of the target model and outputs the semantic information of the target model;
[0141] The semantic decoding module is used to transmit the semantic information of the target model to a preset semantic decoder through a wireless channel, so that the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model;
[0142] The model training module is used to train the semantic encoder and semantic decoder using contrastive learning;
[0143] When training the semantic encoder and the semantic decoder, obtaining a training sample set, and selecting a training sample from the training sample set as a target training sample;
[0144] Using the model damage of the training sample in the wireless channel as data enhancement, determining a first semantic distance between the target training sample and the enhanced sample corresponding to the target training sample, and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample;
[0145] Determining a semantic contrast loss function for the training sample set with the goal of minimizing the first semantic distance and maximizing the second semantic distance;
[0146] The semantic encoder and the semantic decoder are trained according to the training sample set and the semantic contrast loss function.
[0147] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive work.
[0148] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0149] Example 3
[0150] Accordingly, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the semantic communication method described in the above embodiment of the invention.
[0151] The electronic device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The device may include, but is not limited to, a processor and a memory.
[0152] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device and connects various parts of the entire device using various interfaces and lines.
[0153] Example 4
[0154] Accordingly, an embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the semantic communication method described in the above embodiment of the invention.
[0155] The memory can be used to store the computer program, and the processor realizes various functions of the device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0156] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0157] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A semantic communication method, characterized in that: include: Inputting the target model to be transmitted into a preset semantic encoder, so that the semantic encoder extracts the semantic information of the target model and outputs the semantic information of the target model; Transmitting the semantic information of the target model to a preset semantic decoder through a wireless channel, so that the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model; The semantic encoder and the semantic decoder are pre-trained by contrastive learning; When training the semantic encoder and the semantic decoder, a training sample set is obtained, and a training sample is selected from the training sample set as a target training sample; Using the model damage of the training sample in the wireless channel as data enhancement, determining a first semantic distance between a target training sample and an enhanced sample corresponding to the target training sample, and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample; Determine a semantic contrast loss function of a training sample set with the goal of minimizing the first semantic distance and maximizing the second semantic distance; The semantic encoder and the semantic decoder are trained according to the training sample set and the semantic contrast loss function.
2. The semantic communication method according to claim 1, characterized in that: The extracting the semantic information of the target model and outputting the semantic information of the target model includes: Extracting semantic information of the target model, and performing nonlinear mapping on the extracted semantic information to generate a k-dim complex value vector; Performing power normalization processing on the k-dim complex value vector, and outputting semantic information of the target model for transmission in the wireless channel; The extracted semantic information is nonlinearly mapped using the following formula: in, is a k-dim complex-valued vector, represents the semantic encoding operation of parameter θ1, and x is the target model; The k-dim complex value vector is power normalized by the following formula: Wherein, s is the semantic information of the target model used for transmission in the wireless channel, k is the channel bandwidth of the wireless channel, P is the average power of the semantic information transmitter, and * represents the conjugate transpose.
3. The semantic communication method according to claim 1, characterized in that: The step of reconstructing the target model according to the semantic information of the target model and outputting the reconstructed model corresponding to the target model includes: Semantically decode the semantic information of the target model to obtain a reconstructed model corresponding to the target model; The semantic information of the target model is semantically decoded by the following formula: in, is the reconstructed model corresponding to the target model, represents the semantic decoding operation of parameter θ2, It is the semantic information of the target model received by the semantic information receiving end.
4. The semantic communication method according to claim 1, characterized in that: The semantic encoder is a convolutional neural network model; The convolutional neural network model includes: a convolution, several downsampling modules, and a channel coding module; wherein the downsampling module includes: a ResBolck module and a convolution for downsampling the target model; the ResBolck module is a basic module in ResNet.
5. The semantic communication method according to claim 1, characterized in that: The semantic decoder is a convolutional neural network model; The convolutional neural network model includes: a convolution, several upsampling modules, and a re-encoding module; wherein the upsampling module includes: a ResBolck module and a Pixel-Shuffle module for upsampling the target model; the ResBolck module is the basic module in ResNet.
6. The semantic communication method according to claim 1, characterized in that: The step of determining a first semantic distance between a target training sample and an enhanced sample corresponding to the target training sample, and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample comprises: Inputting the target training sample into the semantic encoder so that the semantic encoder extracts the semantic information of the target training sample and outputs the semantic information of the target training sample; Transmitting the semantic information of the target training sample to the semantic decoder through a wireless channel, so that the semantic decoder reconstructs the target training sample according to the semantic information of the target training sample, outputs a reconstructed sample model corresponding to the target training sample, and uses the reconstructed sample model as an enhanced sample of the target training sample; Generate a feature map of the target training sample, a feature map of the enhanced sample, and feature maps of the remaining training samples in the training sample set; By using a preset projection network, the feature map of the target training sample is mapped into a preset semantic space to obtain a first projection result, the feature map of the enhanced sample is mapped into the semantic space to obtain a second projection result, and the feature maps of the remaining training samples in the training sample set are mapped into the semantic space to obtain a third projection result; Determining a first semantic distance between a target training sample and an enhanced sample corresponding to the target training sample according to a cosine similarity between the first projection result and the second projection result; According to the cosine similarity between the first projection result and the third projection result, a second semantic distance between the remaining training samples in the training sample set and the target training sample is determined.
7. The semantic communication method according to claim 1, characterized in that: The step of training the semantic encoder and the semantic decoder according to the training sample set and the semantic contrast loss function includes: According to the training sample set and the semantic contrast loss function, performing first-stage training and second-stage training on the semantic encoder and the semantic decoder; The first stage of training includes: Determining a model reconstruction loss function according to a reconstruction loss between a training sample in the training sample set and a reconstruction sample model corresponding to each training sample; Determine a first loss function according to the model reconstruction loss function and the semantic contrast loss function, and then perform a first-stage training on the semantic encoder and the semantic decoder according to the first loss function; The first loss function is: Among them, α1∈[0,1] is a hyperparameter that controls the trade-off between the model reconstruction loss function and the semantic contrast loss function. Reconstruct the loss function for the model, is the semantic contrast loss function; The second phase of training includes: Determine the downstream task loss function based on the model corruption in the wireless channel; Determine a second loss function according to the downstream task loss function and the model reconstruction loss function, and then perform a second-stage training on the semantic encoder and the semantic decoder according to the second loss function; The second loss function is: Among them, α2∈[0,1] is a hyperparameter that controls the trade-off between the downstream task loss function and the model reconstruction loss function. Reconstruct the loss function for the model, is the loss function for the downstream task.
8. A semantic communication device, characterized in that: include: Semantic encoding module, semantic decoding module and model training module; The semantic encoding module is used to input the target model to be transmitted into a preset semantic encoder, so that the semantic encoder extracts the semantic information of the target model and outputs the semantic information of the target model; The semantic decoding module is used to transmit the semantic information of the target model to a preset semantic decoder through a wireless channel, so that the semantic decoder reconstructs the target model according to the semantic information of the target model and outputs a reconstructed model corresponding to the target model; The model training module is used to train the semantic encoder and the semantic decoder by contrastive learning; When training the semantic encoder and the semantic decoder, a training sample set is obtained, and a training sample is selected from the training sample set as a target training sample; Using the model damage of the training sample in the wireless channel as data enhancement, determining a first semantic distance between a target training sample and an enhanced sample corresponding to the target training sample, and determining a second semantic distance between the remaining training samples in the training sample set and the target training sample; Determine a semantic contrast loss function of a training sample set with the goal of minimizing the first semantic distance and maximizing the second semantic distance; The semantic encoder and the semantic decoder are trained according to the training sample set and the semantic contrast loss function.
9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the semantic communication method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the semantic communication method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Title generation method, computer equipment, storage medium and computer program product
CN116208824A
Semantic communication transmission method and system for substation inspection images
CN116320538A
Semantic communication method and system for collaborative learning of transceiver resisting mismatching of background knowledge base
CN116488748A
MIMO voice transmission method and system based on semantic perception network
CN116543769A
Multitask-oriented voice semantic communication method, device and system
CN116884404A
Cited By
Semantic communication method, system and equipment based on personalized federal learning
CN120378929A