Semantic communication method based on semantic perception mask and vector quantization codebook
By adopting semantic aware compression and vector quantization codebook methods in semantic communication, the problems of low transmission efficiency and poor interference resistance in the prior art are solved, and more efficient semantic information processing and transmission are achieved, which is suitable for tasks such as image classification.
Patent Information
- Application Number
- CN202510155162.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
The existing semantic communication methods have low transmission efficiency and poor anti-interference in tasks such as image classification. The random mask requires a lot of experimental adjustments, making it difficult to effectively compress and transmit semantic information.
The semantic communication method based on semantic sensing compression and vector quantization codebook is adopted, and the semantic sensing mask is performed through the CLIP pre-trained model, the image block features with high semantic importance are extracted, and the image block features with high semantic importance are converted into coded feature vectors, and the nearest neighbor search is mapped to the basis vector in the vector quantization codebook, and the transmission is carried out using a robust basis vector index.
It significantly improves the processing efficiency of semantic information, enhances the model's ability to learn semantic information, improves transmission efficiency and anti-interference performance, and shows better performance in low signal-to-noise ratio environments.
Smart Images

Figure CN120088548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and particularly to a semantic communication method based on semantic-aware compression and vector quantization codebooks. Background Art
[0002] In recent years, semantic communication, as a key technology to meet the task requirements of the intelligent communication era, can overcome the challenges associated with traditional communication systems and has received extensive attention in the industry. Based on advanced artificial intelligence technologies, it preprocesses data so that all communication participants can reduce the network burden by transmitting the most relevant information for the receiver or the goal of the communication task. By only extracting and transmitting the semantic features of information, semantic communication effectively solves the limitations of traditional systems in processing massive data, thereby reducing data traffic and improving communication efficiency.
[0003] Considering that images and videos constitute a large part of the data stream and often contain a large amount of redundant information, task-oriented semantic communication is introduced. In applications focusing on specific tasks, only the task-related semantic information is extracted and sent, and then used for decision-making. This method can achieve efficient communication in low-bandwidth environments and ensure the accurate and effective completion of tasks.
[0004] In particular, in tasks such as image classification, mask techniques are usually adopted to reduce the amount of data to be transmitted, thereby improving the effect of data compression. However, in previous works, most input image segments were usually randomly masked, which may blur key features and hinder the model's ability to learn transferable representations. In addition, random masking requires adjusting the mask ratio for different tasks and datasets, so a large number of experiments are needed.
[0005] In addition, the output of the semantic encoder is usually a continuous feature vector, and existing semantic communication methods often directly map this source feature vector to channel symbols for transmission, resulting in low transmission efficiency and poor anti-interference ability. Summary of the Invention
[0006] Object of the Invention: The object of the present invention is to provide a semantic communication method based on semantic-aware compression and vector quantization codebooks, and through the designed semantic communication architecture based on semantic-aware compression and vector quantization codebooks, to achieve target tasks such as image classification with minimal distortion and loss.
[0007] Technical Solution: To achieve the above object of the invention, a semantic communication method based on semantic-aware masking and vector quantization codebooks provided by the present invention includes the following steps:
[0008] Step 1: Load the contrastive language-image pre-trained CLIP pre-trained model, and perform semantic-aware masking on the input image according to the semantic importance represented by the attention vectors of the last self-attention layer of the model's image encoder;
[0009] Step 2: Extract the semantic information of the image patches retained in the semantic-aware masking operation through the semantic encoder, and convert it into an encoded feature vector;
[0010] Step 3: Map the encoded feature vector to the corresponding basis vector in the vector quantization codebook by the method of nearest neighbor search;
[0011] Step 4: Modulate the basis vector index obtained in Step 3 onto the symbols in the finite constellation diagram and transmit it through the channel;
[0012] Step 5: Recover the transmitted data from the received signal through the semantic decoder;
[0013] Step 6: According to the joint training method, use the neural network to train the semantic encoder, decoder and vector quantization codebook to minimize the loss function of the target task.
[0014] Further, in Step 1, the input image is segmented into multiple image patches, and the semantic importance of each image patch is quantified by the attention weights of the last self-attention layer of the CLIP image encoder, and the semantic importance of the image patch is converted into the probability of being retained in the masking process.
[0015] Further, in Step 1, the input image is segmented into L image patches, and the semantic importance of each image patch is quantified by the attention weights of the last self-attention layer of the CLIP image encoder, specifically as follows:
[0016]
[0017] Among them, The attention vector representing the class token, Q = s class W Q The query vector representing the class token, K = [s class ; s 1 ;...; s L W K Represents the key matrix of the attention mechanism; Represents the attention weight of the last self-attention layer in the CLIP image encoder, D represents the embedding dimension, Represents a real number, s class Represents the embedding vector of the class token, s 1 ;...; s L Represents the feature vector corresponding to each image patch.
[0018] Further, when it is a multi-head attention mechanism, α is obtained by averaging all the heads.
[0019] Further, in step 2, the semantic encoder adopts a Vision Transformer (ViT), which divides the image into patches through multiple Transformer modules, and then serializes the image patches to extract features.
[0020] Further, in step 4, the indices of the encoded features are mapped into binary bits, and these binary bits are mapped to finite constellation symbols for transmission over a wireless channel. During the transmission, only the indices of all the basis vectors in the codebook need to be transmitted.
[0021] Further, in step 6, the loss function includes a semantic feature reconstruction loss and a codebook quantization loss. The reconstruction loss is defined as the mean square error between the normalized reconstructed features and the target features; the codebook quantization loss introduces a new term to increase the distance between the basis vectors by reducing the semantic similarity between the basis vectors and promoting the orthogonality between the basis vectors, which is specifically expressed as:
[0022]
[0023] where represents the normalized codebook, and the (i, j)-th element in the matrix represents the semantic similarity between the basis vectors c i and c j ||·|| 2 represents the L2 norm of the matrix, β represents a hyperparameter, and sg[·] represents the stop gradient, which is used to distinguish between forward propagation and backward propagation; during forward propagation, sg[·] does not affect the numerical calculation, while during backward propagation, the gradient does not pass through the encoder and decoder; the second term in the formula represents the quantization loss, which represents the Euclidean distance between the feature vector x e output by the encoder during non-gradient calculation and the basis vector c k in the codebook that is closest to it, and the third term represents the commitment loss, which calculates the Euclidean distance between x e and the non-gradient version of c k .
[0024] Further, the entire training process in step 6 specifically includes:
[0025] Step 6-1, for each training epoch, the sender uniformly and randomly extracts a number of samples from the input data for batch training;
[0026] Step 6-2: For each sample, first determine the masked image patch part of the image based on the attention vector of the last self-attention layer of the CLIP image encoder;
[0027] Step 6-3: Extract semantic information through the semantic encoder to obtain an encoded feature vector and map it to a basis vector in the vector quantization codebook;
[0028] Step 6-4: Randomly generate channel noise where σ 2 represents the noise variance and I represents the identity matrix;
[0029] Step 6-5: Modulate the basis vector index and pass it through the AWGN physical channel;
[0030] Step 6-6: Obtain the decoded signal through the semantic decoder;
[0031] Step 6-7: Train according to the loss function ;
[0032] Step 6-8: Repeat steps 6-2 to 6-7, and use the adaptive moment estimation algorithm to optimize the model parameters;
[0033] Step 6-9: Repeat steps 6-1 to 6-8 until all training rounds are completed to obtain an approximately optimal solution for the parameters.
[0034] The present invention also provides a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the semantic communication method based on the semantic-aware mask and the vector quantization codebook.
[0035] The present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor, it implements the steps of the semantic communication method based on the semantic-aware mask and the vector quantization codebook.
[0036] Beneficial effects: The semantic communication method based on the semantic-aware mask and the vector quantization codebook provided by the present invention proposes a semantic-aware compression strategy, which can selectively mask the image patches with lower semantic importance instead of randomly masking, thereby enhancing the model's ability to learn semantic information and improving the training efficiency. In addition, an improved robust vector quantization codebook shared at the sending end and the receiving end is designed. This codebook uses orthogonal and trainable basis vector indices to represent the encoded features during transmission, thereby enhancing the robustness of the system and reducing the amount of transmitted data. Experiments show that the method proposed by the present invention can significantly improve the processing efficiency of semantic information and exhibits better performance in challenging low signal-to-noise ratio situations. Description of the Drawings
[0037] Figure 1 This is the overall flowchart of the semantic communication method according to the embodiments of the present invention.
[0038] Figure 2 This is the semantic communication model diagram according to the embodiments of the present invention.
[0039] Figure 3 This is the comparison diagram of the experimental effects according to the embodiments of the present invention. Detailed implementation manners
[0040] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments described with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.
[0041] As Figure 1 shown, a semantic communication method based on semantic perception mask and vector quantization codebook disclosed in the embodiments of the present invention includes the following steps:
[0042] Step 1: Load a contrastive language–image pre-training (CLIP) pre-trained model, and perform semantic perception masking on the input image according to the semantic importance represented by the attention vectors of the last self-attention layer of the image encoder of the model.
[0043] In some embodiments, in order to implement a masking strategy based on semantic importance, the attention weights of the last self-attention layer of the CLIP image encoder are used. First, the input image is segmented into L image patches, and then all the image patches and class tokens are used as the input of the CLIP model. The input of the last self-attention layer of CLIP can be expressed as
[0044]
[0045] where s class represents the embedding vector of the class token, s 1 ;...; s L represents the feature vector corresponding to each image patch, and D represents the embedding dimension. Therefore, the semantic importance of each image patch can be quantified by the attention weights of the last self-attention layer of CLIP, specifically as follows:
[0046]
[0047] where represents the attention vector of the class token, Q = s class W Q represents the query vector of the class token, K = [s class ; s 1;...; s L W K The key matrix representing the attention mechanism. The attention weights representing the last self-attention layer in the CLIP image encoder. The above equation illustrates the interaction between the class token and other features in the single-head attention mechanism. When it comes to the multi-head attention mechanism, i.e., multiple attention modules stacked together, α is obtained by averaging over all heads.
[0048] The class label of the last layer of the CLIP image encoder is used to align with the text embedding of the text encoder, so the vector α can be used to quantify the contribution degree of each image patch to the final output feature. Specifically, when i = 0, α 0 represents the comprehensive attention degree of the class token to the entire image, while when 1 ≤ i ≤ L, the i-th element α i represents the semantic importance of the i-th image patch, and this vector is the basis for establishing the non-uniform sampling distribution. α can be transformed into a probability distribution through the softmax function, i.e., α i represents the probability that the i-th image patch is retained during the masking process, thus realizing the masking compression of semantic-aware sampling.
[0049] Therefore, according to the attention vector α, a non-uniform sampling distribution is adopted to determine which image segments should be masked. Image patches with higher semantic importance (i.e., larger α i values) are more likely to be retained, while image patches with lower semantic importance are more likely to be masked.
[0050] Step 2: Extract the semantic information of the image patches retained in the semantic-aware masking operation through the semantic encoder and convert it into an encoded feature vector.
[0051] Specifically, the image patches S′ retained in the masking operation are input into the semantic encoder and encoded into a transmissible feature vector. In some embodiments, the semantic encoder adopts a Vision Transformer (ViT), which is a Transformer-based vision model that divides an image into patches through multiple Transformer modules and then serializes the patches to extract features. The ViT model has achieved good results in tasks such as image classification and object detection. Applying the ViT model to semantic communication can extract the semantic features of images, thus realizing efficient semantic communication. Its output can be expressed as:
[0052] X = S e (S′; θ 1 )
[0053] where, S e (·) represents the semantic encoder with parameters θ1 。
[0054] Step 3: Map the encoded feature vectors to the corresponding basis vectors in the vector quantization codebook by the method of nearest neighbor search.
[0055] Specifically, the transmissible feature vectors output in Step 2 are usually continuous. Existing semantic communication methods often directly map them to channel symbols for transmission, resulting in low transmission efficiency and high complexity. In some embodiments, a discrete codebook shared between the transmitter and the receiver is used, where the basis vectors in the codebook are used to represent the encoded features. These vectors are configured as trainable parameters and trained together with the model parameter θ. In the training phase, the continuous features generated by the encoder are mapped to the discrete indices corresponding to the basis vectors. Map the output of the semantic encoder to the discrete index corresponding to the basis vector in the discrete codebook shared between the transmitter and the receiver, thereby reducing the amount of data transmitted. The specific implementation is as follows:
[0056] First, the designed discrete codebook for the encoded features consists of K basis vectors, and the dimension of the basis vectors is H. Then the discrete codebook can be expressed as
[0057] Next, find the basis vector c e that is most similar to the feature vector x output by the encoder k , which is specifically represented as follows:
[0058]
[0059] where c k is the k-th basis vector in the codebook. Through the above method, the encoded feature vectors can be mapped to the basis vector c in the codebook that is closest to them k . This forward calculation is conceptualized as a specific DNN layer that maps the encoded vectors to the corresponding basis vectors.
[0060] However, the process shown in the above formula is non-differentiable. Therefore, a straight-through estimator is used to approximate the gradient in backpropagation, allowing the gradient information to be passed to the encoder for training. At the same time, in the forward pass, the basis vectors without gradients are directly passed to the decoder.
[0061] Step 4: Modulate the basis vector indices obtained in Step 3 onto the symbols in the finite constellation and transmit them through the channel.
[0062] Specifically, the indices of the basis vectors in the discrete codebook are transmitted over the wireless channel to obtain the feature vectors. First, the indices of the encoded features are mapped into binary bits, and these binary bits are mapped to finite constellation symbols for transmission over the wireless channel. During the transmission, only the indices of all the basis vectors need to be transmitted. Generally, the size of the indices is much smaller than the basis vectors in the codebook, so the amount of data to be transmitted can be significantly reduced during the transmission. The specific wireless channel transmission process is simplified as follows:
[0063] Y = HX + n
[0064] Among them, the input Y of the semantic decoder represents the feature vector after passing through the physical channel, and H represents the channel gain.
[0065] Step 5: Recover the transmitted data from the received signal through the semantic decoder.
[0066] Pass the feature vector Y through the semantic reconstruction decoder to recover all the features of the input image according to the retained important semantic features. The output of the decoder can be expressed as:
[0067]
[0068] Among them, S d (·) represents the semantic decoder with parameter θ 2 .
[0069] Step 6: According to the joint training method, use the neural network to train the semantic encoder, decoder, and vector quantization codebook to minimize the loss function of the target task.
[0070] In Step 6, train by calculating the semantic feature reconstruction loss and the codebook quantization loss until convergence.
[0071] First, introduce the reconstruction loss function of the semantic encoder-decoder. Specifically, the reconstruction loss is defined as the mean square error between the normalized reconstructed features and the target features, which can be expressed as follows:
[0072]
[0073] Among them, p and t represent the output of the semantic decoder and the target features respectively, and L represents the number of image blocks. This reconstruction loss function is mainly used to train the parameters of the semantic encoder-decoder to extract and reconstruct image features, so as to better implement the compression operation based on the semantic perception mask.
[0074] Subsequently, we define a loss function for the discrete codebook to train the basis vectors of the codebook to closely match the output features of the encoder. Since the discrete codebook is a non-differentiable function, a straight-through estimator is used to approximate the gradient during backpropagation, allowing gradient information to be passed to the encoder for training. During the forward pass, the basis vectors without gradients are directly passed to the decoder. Additionally, to further enhance the robustness of the discrete codebook, the normalized codebook can be expressed as:
[0075]
[0076] where ||·|| represents the L2 norm of the basis vectors. Since the cosine similarity between the codebook vectors increases with the increase in semantic similarity, we define the loss function of the discrete codebook by calculating the Euclidean distance between the basis vectors and the output features of the encoder and the cosine similarity between the basis vectors. Specifically, we introduce a new term into the differentiable loss function where the (i, j)-th element in the matrix represents the semantic similarity between the basis vectors c i and c j , and ||·|| 2 represents the L2 norm of the matrix. This term can increase the distance between the basis vectors by reducing the semantic similarity between the basis vectors and promoting the orthogonality between the basis vectors. The loss function is specifically expressed as
[0077]
[0078] where β represents a hyperparameter, and sg[·] represents stop gradient, which is used to distinguish between the forward pass and the backpropagation. Specifically, during the forward pass, sg[·] does not affect the numerical calculation, while during the backpropagation, the gradient does not pass through the encoder and the decoder. The second term in the formula represents the quantization loss, which represents the Euclidean distance between x e and c k when non-gradient calculation is performed. The third term represents the commitment loss, which calculates the Euclidean distance between x e and the non-gradient version of c k .
[0079] The total loss function of the entire training process can be expressed as:
[0080]
[0081] Exemplarily, the entire semantic communication method framework based on the semantic-aware mask and the vector quantization codebook is specifically as Figure 2 shown. In step 6, it specifically includes the following training process:
[0082] Step 6-1, for each training round \(t\in\{1,2,\ldots,T\}\), the sender randomly extracts \(B\) samples from the input data for batch training. Here, \(T\) is the total number of training rounds.
[0083] Step 6-2, for each sample, first determine the masked image patch part of the image based on the attention vector \(\alpha\).
[0084] Step 6-3, extract semantic information through the semantic encoder:
[0085] \(X = S e (S';\theta 1 )
[0086] Map \(X\) to the basis vectors in the vector quantization codebook
[0087] Step 6-4, randomly generate noise
[0088] Step 6-5, through the AWGN physical channel, the input \(Y\) of the receiver's semantic decoder can be simplified as:
[0089] \(Y = HX + n
[0090] Step 6-6, through the semantic decoder, obtain the decoded signal:
[0091]
[0092] Step 6-7, perform training according to the loss function:
[0093]
[0094] Step 6-8, repeat steps 6-2 to 6-7, and use the Adaptive Moment Estimation (ADAM) algorithm to optimize the model parameters.
[0095] Step 6-9, repeat steps 6-1 to 6-8 until all training rounds are completed to obtain an approximate optimal solution of the parameters.
[0096] According to the semantic communication method based on semantic-aware masking and vector quantization codebook proposed in the embodiments of the present invention, a semantic-aware compression strategy is proposed. It can selectively mask the image patches with lower semantic importance instead of random masking, thereby enhancing the model's ability to learn semantic information and improving the training efficiency. In addition, an improved robust vector quantization codebook shared at the sender and receiver is designed. This codebook uses orthogonal and trainable basis vector indices to represent the encoded features during transmission, thereby enhancing the robustness of the system and reducing the amount of transmitted data.
[0097] Figure 3Illustrates the comparison of experimental effects of embodiments of the present invention, where SAMDC is the semantic communication method based on semantic perception mask and vector quantization codebook proposed in the embodiments of the present invention, Masked VQ-VAE is the semantic communication method using a random mask strategy and a vector quantization codebook, Masked VAE is the MAE architecture using a random mask strategy, and JSCC and JPEG are traditional semantic communication methods. As Figure 3 can be seen, the method proposed by the present invention can significantly improve the processing efficiency of semantic information, has higher accuracy in classification tasks, and exhibits better performance in challenging low signal-to-noise ratio situations.
[0098] Another embodiment of the present invention discloses a computer system, including a memory, a processor, and a computer program / instructions stored on the memory and executable on the processor. When the computer program / instructions are executed by the processor, the steps of the foregoing semantic communication method based on semantic perception mask and vector quantization codebook are implemented.
[0099] Another embodiment of the present invention discloses a computer program product, including a computer program / instructions. When the computer program / instructions are executed by the processor, the steps of the foregoing semantic communication method based on semantic perception mask and vector quantization codebook are implemented.
[0100] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0101] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
Claims
1. A semantic communication method based on semantic-aware mask and vector quantization codebook, characterized in that: The following steps are involved: Step 1: Load the contrastive language-image pre-training CLIP pre-trained model and perform semantic-aware masking on the input image according to the semantic importance represented by the attention vector of the last self-attention layer of the model's image encoder; Step 2: Extract the semantic information of the image block retained in the semantic-aware mask operation through the semantic encoder and convert it into a coded feature vector; Step 3: Map the coded feature vector to the corresponding basis vector in the vector quantization codebook by the nearest neighbor search method; Step 4: modulate the basis vector index obtained in step 3 onto a symbol in a finite constellation diagram and transmit it through a channel; Step 5: Recover the transmitted data from the received signal through the semantic decoder; Step 6: According to the joint training method, the semantic encoder, decoder and vector quantization codebook are trained using a neural network to minimize the loss function of the target task.
2. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 1, characterized in that: In step 1, the input image is divided into multiple image blocks, and the semantic importance of each image block is quantified by the attention weight of the last self-attention layer of the CLIP image encoder, and the semantic importance of the image block is converted into the probability of the image block being retained during the masking process.
3. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 2, characterized in that: In step 1, the input image is divided into L image blocks, and the semantic importance of each image block is quantified by the attention weight of the last self-attention layer of the CLIP image encoder, as follows: in, The attention vector representing the category label, Q = s class W Q The query vector representing the category label, K = [s class ;s1;...;s L ]W K The key matrix representing the attention mechanism; represents the attention weight of the last self-attention layer in the CLIP image encoder, D represents the embedding dimension, represents a real number, s class Embedding vectors representing class labels, s1; ...; s L Represents the feature vector corresponding to each image block.
4. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 3, characterized in that: When it is a multi-head attention mechanism, α is obtained by averaging all heads.
5. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 1, characterized in that: In step 2, the semantic encoder uses a visual transformer ViT to segment the image into blocks through a multi-layer Transformer module, and then serializes the image blocks to extract features.
6. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 1, characterized in that: In step 4, the index of the coding feature is mapped into binary bits, and these binary bits are mapped to finite constellation symbols for transmission on a wireless channel. During the transmission process, only the indexes of all basis vectors in the codebook need to be transmitted.
7. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 1, characterized in that: The loss function in step 6 includes semantic feature reconstruction loss and codebook quantization loss. Defined as the mean square error between the normalized reconstructed feature and the target feature; the codebook quantization loss introduces a new term The distance between basis vectors is increased by reducing the semantic similarity between basis vectors and promoting the orthogonality between basis vectors, which can be specifically expressed as: in, represents the normalized codebook, the matrix The (i,j)th element in represents the basis vector c i and c j , ||·||2 represents the L2 norm of the matrix, β represents the hyperparameter, sg[·] represents the stop gradient, which is used to distinguish forward propagation from backward propagation; during forward propagation, sg[·] does not affect the numerical calculation, while during backward propagation, the gradient does not pass through the encoder and decoder; the second term in the formula represents the quantization loss, which represents the feature vector x output by the encoder during non-gradient calculation e and the basis vector c in the codebook closest to it k The Euclidean distance between them, the third term represents the commitment loss, calculate x e and c k The Euclidean distance between the non-gradient versions of .
8. The semantic communication method based on semantic-aware mask and vector quantization codebook according to claim 7, characterized in that: The entire training process in step 6 specifically includes: Step 6-1: For each training round, the sender uniformly and randomly selects a number of samples from the input data for batch training; Step 6-2, for each sample, first determine the mask image block part of the image based on the attention vector of the last self-attention layer of the CLIP image encoder; Step 6-3, extracting semantic information through a semantic encoder, obtaining a coded feature vector and mapping it to a basis vector in a vector quantization codebook; Step 6-4, randomly generate channel noise where σ 2 represents the noise variance, I represents the identity matrix; Step 6-5, modulate the basis vector index and pass it through the AWGN physical channel; Step 6-6, obtaining a decoded signal through a semantic decoder; Step 6-7, according to the loss function Conduct training; Step 6-8, repeating steps 6-2 to 6-7, and using an adaptive moment estimation algorithm to optimize model parameters; Step 6-9, repeat steps 6-1 to 6-8 until all training rounds are completed and the approximate optimal solution of the parameters is obtained.
9. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the semantic communication method based on semantic-aware mask and vector quantization codebook according to any one of claims 1 to 8 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the semantic communication method based on semantic-aware mask and vector quantization codebook according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Mine video semantic communication transmission method based on feature compression
CN120475172A
Aircraft control audio coding method and system based on dynamic acoustic masking
CN120636422A
Signal feature compression method based on codebook discrete quantization and multi-task learning
CN120724134A
A signal feature compression method based on codebook discrete quantization and multi-task learning
CN120724134B
Robust semantic communication method, device and equipment based on sparse vector coding
CN120768503A