Semantic information transmission communication method and system based on VQGAN optimization

Through the VQGAN-based semantic information transmission method, the VQGAN model is used to extract image features and transmit indexes, which solves the problems of efficient image transmission and device compatibility of semantic communication systems in bandwidth-limited scenarios and achieves high-quality image reconstruction.

CN120785478APending Publication Date: 2025-10-14TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510924141.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing semantic communication systems have difficulty efficiently transmitting high-resolution images in bandwidth-constrained scenarios, and have poor interoperability between multiple devices and platforms, making them incompatible with existing communication systems.

Method used

A semantic information transmission method based on VQGAN optimization is adopted. By establishing a VQGAN model, the encoder and decoder are used to extract image features, find the index of the discrete latent space, perform channel coding and modulation transmission, and reconstruct the image at the receiving end. The shared codebook is used to achieve device compatibility.

Benefits of technology

Efficiently transmit high-resolution images under bandwidth-constrained conditions, be compatible with different devices, reduce redundant data transmission, and ensure image reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120785478A_ABST
    Figure CN120785478A_ABST
Patent Text Reader

Abstract

The invention provides a semantic information transmission communication method and system based on VQGAN optimization. The method comprises the following steps: establishing a VQGAN model for an image; inputting an image to be transmitted into an encoder of the VQGAN model, and extracting image features corresponding to the image by the encoder; searching the index represented by the discrete potential space with the nearest distance according to the feature vector of the image feature; the index of the embedded vector is recorded, channel coding and modulation are carried out, and transmission is carried out through a channel; after channel transmission, the index is recovered through demodulation and channel decoding; and finding the spatial representation in the codebook according to the index, and performing image reconstruction by using a decoder of a receiving end to obtain a reconstructed image corresponding to the to-be-transmitted image. According to the method, a discrete codebook is established in a vector quantization mode, transmission content is reduced by only transmitting an index of the codebook, and high-resolution image reconstruction is performed based on a generative adversarial network and dynamic change of a confrontation channel and compatibility among different devices based on traditional channel coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a semantic information transmission communication method and system based on VQGAN optimization. Background Art

[0002] With the recent increase in computing power and continuous innovation in artificial intelligence (AI), many intelligent services, such as smart factories, vehicle-road collaboration, and virtual reality, are also developing. From a communications perspective, these intelligent services are widely faced with the challenge of quickly and efficiently transmitting large amounts of data to meet user needs. Traditional communication systems focus primarily on the accuracy of data symbols, requiring the transmission of large amounts of redundant data to ensure information integrity. Semantic communication emphasizes the transmission of meaningful information, enabling recipients to correctly understand and effectively utilize the data.

[0003] The core of semantic communication systems lies in the encoding and decoding method. Currently, semantic communication generally employs a joint source-channel coding architecture. During training, this architecture maps source data directly to channel symbols, sets fixed channel conditions, and designs a loss function based on task performance, leading to end-to-end training. In semantic communication, ideal channel conditions are often assumed, ignoring the dynamic nature of channels in real wireless communication environments. The time-varying nature of channel conditions can significantly reduce the effectiveness of joint coding strategies. Secondly, constellations in traditional communication systems are standardized, allowing communication equipment and protocols to be designed for interoperability based on these standards. However, semantic communication systems do not require signals to be mapped to fixed constellation points, making it difficult to achieve a unified standard for interoperability across multiple devices and platforms. This makes it difficult to ensure compatibility and efficient operation between systems, and it is also incompatible with existing communication systems. Finally, due to the large amount of data required to transmit high-resolution images, redundant data may be added to ensure reliability, increasing the transmission burden. Insufficient channel bandwidth can result in the inability to recover high-frequency details or edge information in the image. Summary of the Invention

[0004] In order to solve the above problems, the purpose of the present invention is to provide a semantic information transmission communication method and system based on VQGAN optimization, which can be applied to the semantic communication of high-resolution images in bandwidth-limited scenarios.

[0005] The present invention provides a semantic information transmission and communication method based on VQGAN optimization, the method comprising: Establishing a VQGAN model for an image; the VQGAN model includes an encoder and a decoder; Inputting the image to be transmitted into the encoder of the VQGAN model, and extracting image features corresponding to the image by the encoder; finding an index of a discrete latent space representation closest to the feature vector of the image feature in a codebook; record the index of the embedding vector and perform channel coding and modulation, and transmit through the channel; recover the index through demodulation and channel decoding after channel transmission; combine the embedding vectors in order, find the embedding vectors in the codebook according to the index, and reconstruct the image using the decoder at the receiving end to obtain the reconstructed image corresponding to the image to be transmitted.

[0006] Optionally, the VQGAN model is a generative model combining a generative adversarial network (GAN) and a vector quantization (VQ) technique. After building the VQGAN model, train the encoder and the decoder, and the related parameters of the training are: the image size for training is 3x256x256, the dimension of the latent space is set to 256, the learning rate is , Adam is used as the optimizer, the weight of the perceptual loss is 0.5, the weight of the reconstruction loss is 1, the training batch is 6, and the training round is 100 rounds.

[0007] Optionally, the sending end and the receiving end share the same codebook. The codebook is pre-trained by the VQGAN model and essentially is a fixed-size multidimensional array containing a fixed number of embedding vectors.

[0008] Optionally, the recording of the index of the embedding vector and the channel coding and modulation to generate a wireless signal for transmission through the channel include: arrange the index values of the index matrix representing the image to be transmitted in order to generate an index sequence, wherein each index is represented by 13-bit binary. Hamming code is used for channel coding, QPSK mapping is used for modulation to generate a complex symbol for modulation to generate a wireless signal, and the wireless signal is transmitted through the channel.

[0009] Optionally, the wireless signal transmitted through the channel is demodulated and channel decoded at the receiving end to recover the index, which includes: The receiving end performs QPSK demodulation on the received wireless signal to obtain a bit stream, decodes the demodulated bit stream using the error correction capability of Hamming code, and recovers the original bit stream to generate an index matrix.

[0010] The application also provides a semantic information transmission communication system based on VQGAN optimization, characterized in that the system comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor. The processor is configured to execute the VQGAN optimization-based semantic information transmission communication method according to the instructions in the program code.

[0011] The memory is configured to store program code and transmit the program code to the processor. The processor is configured to execute the VQGAN optimization-based semantic information transmission communication method according to the instructions in the program code.

[0012] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the VQGAN optimization-based semantic information transmission communication method.

[0013] The application further provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps of the VQGAN optimization-based semantic information transmission communication method.

[0014] The above and other objects, advantages and features of the application will become more apparent from the following detailed description of specific embodiments thereof, when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0015] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are intended to illustrate preferred embodiments of the application, and should not be considered limiting of the scope of the application. Indeed, those of ordinary skill in the art will appreciate that they can possibly devise other embodiments and adaptations that are readily obtained from the content of this disclosure, without departing from the purview and scope of the application. In the drawings: Figure 1 is a flowchart of a VQGAN optimization-based semantic information transmission communication method according to an embodiment of the application; Figure 2 is a schematic diagram of a VQ-GAN network structure according to an embodiment of the application; Figure 3 is a schematic diagram of a network structure of an encoder according to an embodiment of the application; Figure 4 is a schematic diagram of a network structure of a decoder according to an embodiment of the application; Figure 5 is a schematic diagram of a process of semantic communication according to an embodiment of the application. DETAILED DESCRIPTION

[0016] Embodiments of the application will be described below with reference to the accompanying drawings, which are intended to explain the application, not to limit it.

[0017] As Figure 1As shown, the embodiment of the present application provides a semantic information transmission communication method based on VQGAN optimization, which comprises the following steps: S1, establishing a VQGAN model for images at the sending end; the VQGAN model comprises an encoder and a decoder.

[0018] Firstly, a VQGAN model optimized for chest radiograph image data needs to be constructed. The model contains an encoder and a decoder. The encoder is responsible for compressing the input high-dimensional image into a low-dimensional feature representation; the decoder is responsible for reconstructing the original image according to the feature representation.

[0019] S2, inputting the image to be transmitted into the encoder of the VQGAN model, and extracting the image features corresponding to the image by the encoder.

[0020] For example, a chest X-ray image with a size of 256x256x3 is input into the encoder, and the encoder captures the key features of the image such as chest texture and skeletal outline, etc. After processing by the encoder, a feature vector with a dimension of 64x64x256 is output, which contains a highly abstracted representation of the key visual information of the original image.

[0021] S3, finding the index of the nearest discrete latent space representation in the codebook according to the feature vector of the image features.

[0022] The VQGAN model pre-trains a discrete codebook, which is essentially a fixed-size multidimensional array containing a fixed number of "embedding vectors". The receiving end has exactly the same pre-trained codebook as the sending end. For each 256-dimensional feature vector at each 64x64 position in the feature vector, the distance between it and all embedding vectors in the codebook is calculated. Assuming that the 42nd embedding vector is the closest in position (i, j), the index value 42 is recorded. Finally, the feature tensor of the entire picture is quantized into a 64x64 index matrix.

[0023] S4, record the index of the embedding vector and perform channel encoding and modulation to generate a wireless signal for transmission through the channel.

[0024] The 64x64 index matrix representing the picture, which has 4096 index values, is arranged in order to form an index sequence, each index is represented by 13 bits of binary, and the sequence length is 4096x13 bits. The bit stream is channel encoded using (31, 26) Hamming code, modulated into complex symbols using QPSK, and finally sent through the wireless channel.

[0025] S5, the wireless signal after channel transmission is demodulated and channel decoded by the receiving end to recover the index.

[0026] The receiving end receives the wireless signal and performs QPSK demodulation to obtain a bit stream. The demodulated bit stream is decoded by using the error correction capability of the (31, 26) Hamming code to correct transmission errors. After decoding, the original bit stream with a length of 4096*13 bits is recovered, and is parsed into a 64*64 index matrix again.

[0027] S6, find the embedding vectors in the codebook according to the index.

[0028] For example, for the index value 42 of the position (i, j) in the recovered index matrix, the 42nd embedding vector is searched in the shared codebook. All the found embedding vectors are recombined according to their positions in the index matrix to form a 64*64*256 feature tensor. The tensor is an approximate representation of the continuous feature tensor output by the encoder of the sending end after quantization, transmission and recovery.

[0029] S7, combine the embedding vectors in order, use the decoder of the receiving end to perform image reconstruction, and obtain a reconstructed image corresponding to the image to be transmitted.

[0030] The recovered 64*64*256 feature tensor is input into the decoder. The decoder processes layer by layer, gradually improves the resolution and the number of channels, and finally outputs a 256*256*3 reconstructed chest radiograph. The reconstructed image should be as close as possible to the original picture in vision, and key information should be retained.

[0031] The embodiment of the application is mainly based on the VQGAN algorithm, proposes to use the vector quantization method to establish a discrete codebook, to reduce the transmission content by only transmitting the index of the codebook, and to resist the dynamic changes of the channel and the compatibility between different devices based on the traditional channel coding, and to perform high-resolution image reconstruction based on the generative adversarial network.

[0032] The structure of VQGAN (Vector Quantized Generative Adversarial Network) is as shown in Figure 2 The core of the architecture includes an encoder (Encoder), a codebook (Codebook), a decoder (Decoder) and a discriminator (Discriminator). It is a combination of generative adversarial network (GAN) and vector quantization (VQ) technology, the core of which is to train a codebook (Codebook) to represent data, introduce discreteness in the latent space, and use these codes to generate new data, which enhances the controllability of the generation process and has better stability in optimization.

[0033] After the model is built, the encoder and decoder are trained using the following key parameters: the image size for training is 3x256x256 (channel x height x width), the dimension of the latent space is set to 256, the learning rate is 2x10 -5 -4, Adam is used as the optimizer, the loss function consists of a perceptual loss and a reconstruction loss, the weight of the perceptual loss is 0.5, the weight of the reconstruction loss is 1, the training is performed with a batch size of 6, and a total of 100 training rounds are performed.

[0034] As shown in Figure 3 , the encoder is a deep convolutional neural network that gradually increases the number of channels through convolution layers (Conv2d) to extract image features, enhances the training ability of the network through residual blocks (Residual), and uses attention mechanisms (Attention) at specific resolutions to enhance the modeling of global information. The resolution gradually decreases, and more abstract representations are obtained through downsampling (DownSample), and finally the features are mapped to a continuous latent space.

[0035] As shown in Figure 4 , the decoder input is a discrete latent representation obtained by quantizing the codebook, and its processing flow is opposite to that of the encoder: through multiple convolutions (Conv2d) and residual blocks (Residual), the latent representation is gradually mapped to a high-dimensional space, the resolution is improved through upsampling (UpSample) to restore the spatial dimensions of the image and reconstruct the geometric structure of the image; at a specific resolution, an attention block (Attention) is used to capture global features, and finally a convolution layer is used to convert the last feature map to the original image channel number, outputting a reconstructed image tensor with dimensions 3xHxW.

[0036] As shown in Figure 5 , the images generated by the information source are first encoded by the transmitter to obtain continuous latent space representations . Subsequently, the features are mapped to the nearest latent space vector index through quantization. The transmitter and receiver share the same codebook, so only the index of the codebook needs to be transmitted. These indices are channel-encoded and modulated through the physical channel, and the receiving end receives discrete indices , and then maps the discrete representation back to the latent space with the help of the codebook. The decoder will reconstruct an image similar to the original input. This method ultimately only needs a fixed-size codebook to transmit higher-resolution images. The reconstructed image can be further analyzed and understood semantically, such as using deep learning models for classification, detection, etc.

[0037] The embodiment of the present application also provides a semantic information transmission communication system based on VQGAN optimization, the system comprising a processor and a memory: the memory is used for storing program code and transmitting the program code to the processor; the processor is used for implementing the semantic information transmission communication method based on VQGAN optimization of the above embodiment according to the instructions in the program code.

[0038] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the semantic information transmission communication method based on VQGAN optimization of any one of the above.

[0039] The embodiment of the present application also provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of the semantic information transmission communication method based on VQGAN optimization of any one of the above.

[0040] The above is only the preferred embodiment of the present application, and the protection scope of the present application is not limited to the above embodiment only, and any technical scheme falling within the idea of the present application belongs to the protection scope of the present application. It should be noted that, for ordinary skilled in the art, some improvements and refinements without departing from the principle of the present application are also considered as the protection scope of the present application.

Claims

1. A semantic information transmission communication method based on VQGAN optimization, characterized in that: The method comprises: Establishing a VQGAN model for the image at a transmitting end; the VQGAN model includes an encoder and a decoder; Inputting the image to be transmitted into the encoder of the VQGAN model, and extracting image features corresponding to the image by the encoder; Searching for the index of the nearest discrete latent space representation in the codebook according to the feature vector of the image feature; Record the index of the embedded vector and perform channel coding and modulation to generate a wireless signal for transmission through the channel; The wireless signal after channel transmission is demodulated and channel decoded by the receiving end to recover the index; Find the embedding vector in the codebook based on the index; The embedded vectors are combined in sequence, and an image is reconstructed using a decoder at the receiving end to obtain a reconstructed image corresponding to the image to be transmitted.

2. The method according to claim 1, characterized in that The VQGAN model is a generative model that combines the generative adversarial network (GAN) and vector quantization (VQ) technology. After building the VQGAN model, the encoder and decoder are trained. The training parameters are: the training image size is 3×256×256, the dimension of the latent space is set to 256, and the learning rate is , Adam is used as the optimizer, the weight of the perception loss is 0.5, the weight of the reconstruction loss is 1, the training batch is 6, and the number of training rounds is 100 rounds.

3. The method according to claim 2, characterized in that The transmitting end and the receiving end share the same codebook; The codebook is generated by VQGAN model pre-training and is essentially a multidimensional array of fixed size containing a fixed number of embedding vectors.

4. The method according to claim 1, wherein Recording the index of the embedding vector and performing channel coding and modulation to generate a wireless signal for transmission through the channel includes: Arranging the index values ​​of the index matrix representing the image to be transmitted in sequence to generate an index sequence, wherein each index is represented by 13 bits of binary; Hamming code is used for channel coding, and QPSK mapping is used to modulate complex symbols to generate wireless signals, which are then transmitted through the channel.

5. The method according to claim 4, characterized in that After the wireless signal is transmitted through the channel, it is demodulated and decoded by the receiving end to recover the index, including: The receiving end performs QPSK demodulation on the received wireless signal to obtain a bit stream, uses the error correction capability of the Hamming code to decode the demodulated bit stream, restores the original bit stream, and re-parses to generate an index matrix.

6. A semantic information transmission communication system based on VQGAN optimization, characterized in that: The system includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the semantic information transmission and communication method based on VQGAN optimization according to any one of claims 1 to 5 according to the instructions in the program code.

7. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the semantic information transmission and communication method based on VQGAN optimization according to any one of claims 1 to 5 are implemented.

8. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the semantic information transmission and communication method based on VQGAN optimization according to any one of claims 1 to 5.

Citation Information

Cited By

  • Image adaptive quantization communication method and system

    CN121442092A

  • An image adaptive quantization communication method and system

    CN121442092B