Semantic communication method combining variable length coding and constellation probability shaping

By combining variable-length coding with constellation probability shaping for semantic communication, the coding length is dynamically adjusted and the semantic features are mapped to constellation points in the same distribution. This solves the problems in existing systems where the modulation module cannot perceive the distribution of semantic features and fixed-length coding cannot dynamically adjust the code rate. It improves the system's spectral efficiency and anti-interference capability and is suitable for resource-constrained devices.

CN121750158APending Publication Date: 2026-03-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing semantic communication systems, the separation of quantization and modulation processes makes it impossible for the modulation module to perceive the distribution of semantic features, making it difficult to achieve probabilistic shaping optimization of constellation points. Furthermore, the fixed-length coding mechanism cannot dynamically adjust the transmission code rate, thus failing to fully leverage the efficiency advantages of variable-length coding in time-varying channels.

Method used

We employ a joint variable-length coding and constellation probability shaping method. We dynamically adjust the coding length through a mask rate control function, perform weighted sampling using the maximum entropy criterion, use layer normalization to make the semantic feature distribution approximate a Gaussian prior, and achieve the same distribution mapping between semantic features and constellation points through finite scalar quantization and identically distributed mapping. We also perform end-to-end optimization by combining gradient pass-through estimation.

Benefits of technology

It improves the system's spectral efficiency and anti-interference capability, enhances the quality of source reconstruction, reduces storage and computing overhead, and is suitable for resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750158A_ABST
    Figure CN121750158A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of communication, and particularly relates to a semantic communication method combining variable length coding and constellation probabilistic shaping, which comprises the following steps of: firstly, constructing a semantic encoder based on a mask auto-encoder framework, and designing a mask rate control function through channel state information and information source characteristics so as to dynamically adjust an information source mask rate; a self-adaptive variable-length coding mechanism is realized; according to the obtained mask rate, sampling a non-masked information source unit by adopting a maximum entropy criterion so as to reserve the information amount of the information source to the maximum extent; the constellation point optimal probability is shaped and approximated to discrete two-dimensional Gaussian distribution, and an end-to-end optimization target is deduced based on variational inference; semantic feature coding distribution is constrained through layer normalization operation to approximate Gaussian prior, so that an objective function is effectively simplified; and finally, designing a quantization modulator of semantic feature-constellation point identical distribution mapping based on finite scalar quantization, and adopting a gradient straight-through estimator to solve the non-differentiable problem of the quantization process, and finally realizing end-to-end joint optimization of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, in particular to a semantic communication method combining variable length coding and constellation probability shaping. BACKGROUND

[0002] With the large-scale deployment of 5G networks and the in-depth research of 6G technologies, wireless communication systems are facing increasingly severe challenges in terms of transmission rate, latency, and spectral efficiency. Traditional communication systems usually rely on increasing antenna size, expanding spectrum bandwidth, or increasing transmission power to enhance channel capacity. However, these methods are costly and energy-consuming. As the network scales, the feasibility of relying solely on hardware upgrades to improve system performance gradually decreases. How to meet the demands of future networks for large-scale data transmission, high-bandwidth applications, and low-latency communication under the constraints of limited resources and costs has become a pressing problem.

[0003] Semantic communication, as a new communication method, utilizes the feature extraction capability of deep learning and the end-to-end joint optimization mechanism to realize the transition from "bit transmission" to "semantic communication". Semantic communication systems compress information content through semantic compression, converting raw data into low-dimensional, compact semantic feature vectors, while preserving key information and reducing redundant transmission. Semantic communication also has good adaptive ability, and through joint training, it can dynamically adjust semantic coding and transmission strategies according to real-time channel conditions, network load, and device capabilities, demonstrating excellent flexibility in complex network environments.

[0004] In the process of integrating semantic communication with existing communication systems, the system architecture gradually evolves from analog transmission to digital modulation. Early semantic communication systems usually output floating-point data generated by neural networks and use analog modulation for continuous signal transmission. Although they have high signal fidelity in theory, they require high-precision control of signal amplitude and phase, making hardware implementation complex and costly, which makes it difficult to support large-scale practical deployment. To improve system practicality and compatibility, researchers have turned to digital semantic communication solutions, such as using scalar quantization or vector quantization to discretize semantic features into codebook indexes, or mapping features to digital constellation points based on variational autoencoders.

[0005] However, existing digital semantic communication schemes still have obvious limitations: a. The separation of quantization and modulation processes makes it difficult for the modulation module to perceive the semantic feature distribution, making it difficult to optimize the probability shaping of constellation points; b. Vector quantization requires maintaining a large-scale codebook, resulting in storage and update overheads; c. The variational autoencoder architecture supports joint learning of semantic coding and constellation mapping, but the KL divergence term in the objective function can easily cause reconstruction ambiguity and posterior collapse, restricting system stability and reconstruction quality.

[0006] Furthermore, existing semantic communication systems mostly employ fixed-length coding mechanisms, making it difficult to dynamically adjust the transmission rate according to channel conditions and thus failing to fully leverage the efficiency advantages of variable-length coding in time-varying channels. Although recent research has attempted to achieve rate adaptation through structures such as nested codebooks, a systematic solution is still lacking in how to effectively balance the semantic importance of the source with the dynamic changes in channel state, becoming a key bottleneck restricting further improvements in semantic communication performance.

[0007] Therefore, there is an urgent need in this field for a new semantic communication method to systematically solve the above problems. Summary of the Invention

[0008] In view of this, the present invention aims to provide a semantic communication method that combines variable-length coding and constellation probability shaping. This method can dynamically adjust the coding length according to the channel state and source characteristics, and achieve co-distribution mapping between semantic features and constellation points, thereby improving the anti-interference capability and reconstruction quality of the system while improving spectral efficiency.

[0009] To achieve the above objectives, this invention provides a semantic communication method combining joint variable-length coding and constellation probability shaping, comprising the following steps: S1: Decompose the input signal into several source units and reassemble them; S2: Design a mask rate control function based on channel state information and source characteristics. Change the code length ), Round down; S3: Based on the masking rate, the unmasked source units are weighted and sampled using the maximum entropy criterion to retain the maximum amount of source information, and then masked to obtain a masked source containing the maximum source information. S4: Input the masked information source into the semantic encoder to extract continuous semantic feature encoding; S5: Perform layer normalization processing on the semantic feature encoding to make its distribution approximate Gaussian prior, so as to align the constellation probability to the optimal distribution; S6: Use a finite scalar quantizer A quantization modulator, which is composed of a "semantic feature-constellation point" co-distributed modulator, encodes and maps the semantic features into discrete constellation symbols, thereby achieving a co-distributed mapping between semantic features and constellation points; S7: At the receiving end, the demodulated constellation symbols are quantized and decoded, and the semantic features are reconstructed by combining the mask information. Finally, the source is reconstructed through the semantic decoder. S8: Calculate the loss function and use gradient pass-through estimation to solve the non-differentiability of the scalar quantization modulation process, thereby achieving end-to-end joint optimization of the system.

[0010] Further, in step S1, the input signal First it is broken down into indivual Source unit of size Each source unit corresponds to an index. Sources are arranged in index order Reorganized into:

[0011] Further, in step S2, the mask rate The base mask rate is based on channel state information. With source feature-based adjustment factors Joint control.

[0012] To simplify system design, channel state information is expressed using signal-to-noise ratio. This indicates the mask rate. The error rate decreases as the signal-to-noise ratio (SNR) decreases. Meanwhile, because the bit error rate curve of digital modulation increases sharply below a certain SNR, [the following is a possible interpretation:] ... The function is fitted, and the base mask rate is expressed as:

[0013] in, Indicates the upper and lower bounds of the selectable signal-to-noise ratio. This indicates the upper and lower bounds of the selectable mask rate. This represents the sigmoid function.

[0014] Simultaneously, adjustment factors are designed based on source entropy. To fine-tune the base mask rate, the value range of elements in the image source is divided into equal parts. For each source unit, there are intervals. The corresponding entropy value It can be represented as: ,in, Representative source unit The Middle The histogram probability of each pixel value in the entire information source. .

[0015] The higher the average entropy of the total source units in the image, the richer the information content, and the higher the mask adjustment factor. The mask rate adjustment factor will increase accordingly. Represented as:

[0016] The final mask rate is expressed as: ,in, This indicates that the mask rate value is limited to a range within the interval. middle; Finally, the encoded code length is obtained. ), Round down to the nearest integer to determine the number of unmasked source units.

[0017] Furthermore, in step S3, for an entropy value of source unit Its sampling weight is expressed as:

[0018] The index of the unmasked source unit is obtained by using a weighted sampler based on Gumbel-Softmax, ensuring that the weights are large. Priority samples are retained:

[0019] in, The sampled and retained set of indices is used to generate a mask control matrix with elements of 0 and 1 in index order. ,like for The element in the matrix M is the first element. One element was assigned the value 0.

[0020] Source By Mask controller with parameters This will yield the masked source containing the maximum source information: The mask controller's processing procedure is as follows: , The Hadama product represents element-wise multiplication. Representative shape A matrix of all 1s.

[0021] Furthermore, in step S4, the masked source... It is fed into the semantic encoder:

[0022] in, Therefore A semantic compression encoder with learnable parameters, which includes a feature extractor based on a VisionTransformer architecture for learning masked information sources. The system provides an efficient semantic representation, along with a linear compressor to further reduce information redundancy in the semantic representation, outputting a continuous semantic feature encoding. , The length of the compressed source feature.

[0023] Furthermore, in step S5, the constellation probability shaping optimal distribution follows a Maxwell-Boltzmann distribution, which can be approximated by a two-dimensional discretization of a Gaussian distribution. The optimization objective of the semantic communication system, with a Gaussian distribution prior, is obtained through variational inference, specifically expressed as maximizing the marginal log-likelihood. :

[0024] We introduce an approximate prior Gaussian q(s|x) to replace the true prior p(s|x), and then... Represented as:

[0025] in, For the refactoring item, This is a distribution constraint term.

[0026] Furthermore, in step S5, the semantic feature encoding... Let its mean and standard deviation be denoted as . and Perform layer normalization on the elements:

[0027] Thus, semantic features are encoded. Transform it into an approximate standard normal distribution, as the objective function. The distribution constraint term is handled with soft constraints, and the optimization objective simplifies to minimizing the end-to-end reconstruction loss. .

[0028] Furthermore, in step S6, the semantic feature encoding after layer normalization is performed. Sent to Quantization modulator for regular parameters ,make Achieving a co-distributive mapping with the modulation constellation symbol set, the processed modulation constellation symbol set is represented as: .

[0029] Among them, quantization modulator Finite scalar quantizer It consists of two parts: a "semantic feature-constellation point" co-distributed modulator and a "semantic feature-constellation point" modulator.

[0030] Encoding continuous semantic features Quantified into A set of elements The closest value among them For quantization interval; The modulator contains two orthogonal carriers. Taking a co-distributed modulator as an example, using Corresponding quantification level To control the carrier amplitude:

[0031] Where G represents the gain factor, Corresponding to the minimum carrier amplitude, , The reference amplitude unit representing the constellation symbol coordinates in square QAM modulation.

[0032] Therefore, The mapping of two feature elements in the matrix yields a length of... The constellation symbol of the bit, at this time modulating the constellation symbol set. The distribution is regarded as discrete semantic features A two-dimensional extension, including one constellation symbol. The coordinates are represented as:

[0033] Using additive Gaussian white noise wireless channels Then, the constellation symbol set at the receiving end can be represented as:

[0034] One of the constellation symbols The coordinates are ( This is indicated. If the receiver uses minimum Euclidean distance demodulation, this step is also performed using a scalar quantizer. simplify:

[0035] in, The coordinates representing the demodulation constellation symbols are determined based on the carrier amplitude corresponding to the receiving constellation point. and The corresponding quantitative level can be obtained. and According to the set Decoding semantic feature encoding:

[0036] in, ) is a quantitative demodulator with β as the transformation rule parameter.

[0037] Furthermore, in step S7, the decoded semantic feature vector Based on the source element mask matrix that arrives at the receiver in advance Fill in:

[0038] in, This indicates that the decoded semantic vector will be processed according to the index signal. Fill in with the source Same shape, The Hadama product represents element-wise multiplication. Representative shape A matrix of all 1s Representative shape A matrix of all zero elements, index signal from The position of the 0 element is obtained.

[0039] Finally, the padded features are fed into the semantic decoder to reconstruct the source:

[0040] in, Therefore A semantic decoder with learnable parameters.

[0041] Furthermore, in step S8, the loss function includes the reconstruction term. Reconstruction loss - Expressed using the minimum mean square error of reconstruction (MSE):

[0042] Encoding the semantic features To address the non-differentiability introduced by scalar quantization, pass-through estimation is introduced to ensure backpropagation of the gradient:

[0043] in, This signifies stopping gradient calculation; that is, when the model performs forward inference, for any input, It can be used as a signal hold; and when the model backpropagates, for any input, The output of all of them is 0.

[0044] This invention also provides a semantic communication system combining joint variable-length coding and constellation probabilistic shaping, comprising: A source decomposition and masking control module is used to implement steps S1 to S3 as described in claim 1; A semantic encoding and layer normalization module is used to implement steps S4 to S5 as described in claim 1; A quantization modulation module is used to implement step S6 as described in claim 1; A semantic decoding and reconstruction module is used to implement step S7 as described in claim 1; An end-to-end training module is used to implement step S8 as described in claim 1.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Based on the mask autoencoder architecture, the present invention dynamically adjusts the mask rate by combining the channel state and source characteristics to achieve adaptive transmission, thus overcoming the problem of low efficiency of traditional fixed-length coding in time-varying channels.

[0046] (2) This invention constrains the semantic feature distribution to Gaussian prior by layer normalization, aligning it with the constellation probability shaping optimal distribution, and simplifies the optimization objective to a single reconstruction loss, thereby avoiding posterior collapse and training instability caused by KL divergence in variational autoencoders.

[0047] (3) This invention uses finite scalar quantization and identically distributed mapping modulation to realize the identically distributed mapping of continuous semantic features to discrete constellation symbols without the need to maintain a large-scale codebook; combined with gradient pass-through estimation, it supports the joint optimization of semantic encoding, shaping, modulation and decoding.

[0048] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart of the semantic communication system supporting variable-length coding and constellation joint modulation according to the present invention; Figure 2 This is a system framework diagram of the present invention; Figure 3 A schematic diagram of the working process of a modem for the same distribution mapping of "semantic features-constellation points". Detailed Implementation

[0050] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0051] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0052] This invention provides a semantic communication method that supports variable-length coding and constellation point joint modulation, combined with the appendix. Figures 1-3 To explain, Figure 1 This is a flowchart of the semantic communication method supporting variable-length coding and constellation joint modulation according to the present invention; Taking image sources as an example, Figure 2 This is a framework diagram of the semantic communication method used in this embodiment, which specifically includes the following: Step S1: Image segmentation Input image First it is broken down into indivual Image patches of various sizes Each image patch corresponds to an index. Images in index order Reorganized into:

[0053] In step S2: Mask rate calculation First, the calculation is based on the signal-to-noise ratio. base mask rate and adopt Fitting the function:

[0054] in, For the selectable upper and lower bounds of the signal-to-noise ratio, For selectable upper and lower bounds of the mask rate, The sigmoid function.

[0055] Next, the entropy value of the image patch is calculated. The value range of elements in the image source is divided into equal parts. For each interval, The corresponding entropy value is:

[0056] in, Represents image blocks The Middle The histogram probability of each pixel value in the entire information source. .

[0057] Finally, the adjustment factor was designed using the source entropy. Fine-tuning the base mask rate yields the final mask rate. :

[0058]

[0059] in, This indicates that the mask rate value is limited to a range within the interval. In the end, the final code length is )( (Rounded down) is the number of unmasked source units.

[0060] Step S3: Maximum Entropy Weighted Sampling and Masking Calculate the sampling weight of each image patch based on the entropy value:

[0061] The index of the unmasked source unit is obtained using a Gumbel-Softmax-based weighted sampler, ensuring that the weights are large. Prioritized samples are retained to generate the mask control matrix. The reserved block has a corresponding position of 0, and the mask block has a position of 1.

[0062] in, For the sampled and retained set of indices, generate a mask control matrix with elements of 0 and 1, in index order. ,like for The element in the matrix M is the first element. One element was assigned the value 0.

[0063] image By Mask controller with parameters To obtain the masked source containing the maximum source information. :

[0064] The processing procedure is as follows , The Hadama product represents element-wise multiplication. For the shape of A matrix of all 1s.

[0065] Step S4: Semantic Feature Extraction Masked source Feed into a semantic encoder that includes a feature extractor and a linear compressor based on the Vision Transformer architecture:

[0066] Learning masked sources Effective semantic representation, and output continuous semantic feature encoding. , This represents the length of the compressed source feature.

[0067] Step S5: Layer Normalization The mean is The standard deviation is semantic feature encoding Perform layer normalization:

[0068] To make its distribution approximate a standard normal distribution, after normalization As in the objective function The soft constraint handling of the distribution constraint term simplifies the optimization objective to minimizing the end-to-end reconstruction loss. To match the Gaussian prior required for constellation probability shaping, specifically, it is expressed as maximizing the marginal log-likelihood. :

[0069] Step S6: Quantization modulation, channel transmission and reception demodulation (1) Quantization modulation: Send in Quantization modulator for regular parameters , making Achieving a co-distributive mapping with the modulation constellation symbol set, the processed modulation constellation symbol set is represented as: .

[0070] The quantization modulator consists of a finite scalar quantizer. It consists of two parts: a "semantic feature-constellation point" co-distributed modulator and a "semantic feature-constellation point" modulator.

[0071] Will Quantified into A set of elements The closest value among them This is the quantization interval. Figure 3 This is a schematic diagram of the operation of a quantization modem with "semantic feature-constellation point" co-distributed mapping, preferably using two orthogonal carriers. Co-distributed modulator, utilizing Corresponding quantification level To control the carrier amplitude: ,in, Corresponding to the minimum carrier amplitude G represents the gain factor.

[0072] The above modulation process will After mapping the two feature elements in the data, we obtain a length of... Bit constellation symbols, modulation constellation symbol set The distribution is regarded as discrete semantic features A two-dimensional extension, including one constellation symbol.

[0073]

[0074] (2) Channel transmission: Constellation symbols pass through additive white Gaussian noise. wireless channels The signal is transmitted to the receiving end, where it is demodulated using the minimum Euclidean distance and a scalar quantizer. simplify:

[0075]

[0076] in, The coordinates representing the demodulation constellation symbols are determined based on the carrier amplitude corresponding to the receiving constellation point. and The corresponding quantitative level can be obtained. and According to the set Decoding semantic feature encoding:

[0077] in, It is a quantitative demodulator with β as the transformation rule parameter.

[0078] Step S7: Semantic Reconstruction Based on the source element mask matrix that arrives at the receiver in advance Fill in the decoded semantic feature vector :

[0079] in, This indicates that the decoded semantic vector will be processed according to the index signal. Fill in with the source Same shape, The Hadama product represents element-wise multiplication. Representative shape A matrix of all 1s Representative shape A matrix of all zero elements, index signal Available from The position of the 0 element is obtained.

[0080] Finally, the padded features are fed into a semantic decoder to reconstruct the image:

[0081] in, Therefore A semantic decoder with learnable parameters.

[0082] Step S8: Training Reconstruct the loss function - Expressed using the minimum mean square error of reconstruction (MSE):

[0083] Introducing pass-through estimation ensures backpropagation of gradients, thus addressing semantic feature encoding. The non-differentiability resulting from scalar quantization:

[0084] in, This signifies stopping gradient calculation; that is, when the model performs forward inference, for any input, It can be used as a signal hold; and when the model backpropagates, for any input, The output of all of them is 0.

[0085] Example of effect description: The semantic communication method based on joint variable-length coding and constellation probability shaping was simulated and tested using the CIFAR-10 dataset under additive white Gaussian noise (AWGN) and Rayleigh fading channel conditions.

[0086] Compared with a fixed-rate semantic communication system based on variational autoencoders, the structural similarity index (SSIM) of the reconstructed source is improved by about 14.5% in an AWGN channel with a 0dB signal-to-noise ratio (SNR) and by about 60% in a Rayleigh channel with a 0dB signal-to-noise ratio (SNR) ratio, making it more suitable for efficient data transmission and reconstruction in low SNR environments.

[0087] Compared with a semantic communication system scheme using uniform distribution modulation without constellation probability shaping, the root mean square error (MSE) of the semantic feature encoding obtained by demodulation and decoding is reduced by about 75% under an AWGN channel with a resolution of 0dB. This indicates that the constellation probability shaping joint optimization method of the present invention effectively improves the anti-interference capability of the system.

[0088] Compared to vector quantization and transmission of codebook index (codebook size) Compared to semantic communication system schemes that use 9-bit encoding for each index, this invention reduces the code length by about 27% while maintaining comparable reconstruction quality, making it more suitable for application scenarios with high bandwidth requirements.

[0089] Under 16-QAM modulation, the model parameters proposed in this invention are only 0.862M, while a fixed-rate semantic communication system based on variational autoencoders with comparable reconstruction quality requires 33.0M parameters under the same modulation scheme. This significantly reduced number of model parameters gives this invention a clear advantage in terms of storage and computational overhead, making it more suitable for resource-constrained devices such as edge devices, embedded systems, and mobile terminals.

[0090] In summary, this invention is based on a masked autoencoder architecture, which jointly adjusts the mask rate based on channel state and source features to achieve adaptive transmission; it simplifies the optimization objective by constraining semantic features to approximate a Gaussian distribution through layer normalization; and it achieves end-to-end joint optimization by employing finite scalar quantization and gradient pass-through estimation, effectively balancing the semantic importance of the source with the dynamic changes in channel state, thereby improving semantic communication performance.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A semantic communication method combining joint variable-length coding and constellation probabilistic shaping, characterized in that, Includes the following steps: S1: Decompose the input signal into several source units and reassemble them; S2: Design a mask rate control function based on channel state information and source characteristics. Change the code length ), Round down: in, This indicates that the mask rate value is limited to a range within the interval. middle, The base mask rate is based on channel state information. This is a regulation factor based on source characteristics; S3: Based on the code length, the unmasked source units are weighted and sampled using the maximum entropy criterion to retain the maximum amount of source information, and then masked to obtain a masked source containing the maximum source information. S4: Input the masked information source into the semantic encoder to extract continuous semantic feature encoding; S5: Perform layer normalization on the semantic feature encoding to make its distribution approximate a Gaussian prior, aligning it with the constellation probability shape-optimal distribution. This constellation probability shape-optimal distribution follows a Maxwell-Boltzmann distribution, approximated by a two-dimensional discretization of the Gaussian distribution. An approximate prior Gaussian q(s|x) is introduced to replace the true prior p(s|x), maximizing the marginal log-likelihood. Represented as: in, For the refactoring item, For distribution constraint terms; S6: Use a finite scalar quantizer A quantization modulator, consisting of a "semantic feature-constellation point" co-distributed modulator, encodes and maps the semantic features into discrete constellation symbols, thereby achieving a co-distributed mapping between semantic features and constellation points. S7: At the receiving end, the demodulated constellation symbols are quantized and decoded, and the semantic features are reconstructed by combining the mask information. Finally, the source is reconstructed through the semantic decoder. S8: Calculate the loss function and use gradient pass-through estimation to solve the non-differentiability of the scalar quantization modulation process, thereby achieving end-to-end joint optimization of the system.

2. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, In step S1, the input signal Decomposed into indivual Source unit of size Each source unit corresponds to an index. Sources are arranged in index order Reorganized into: .

3. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, In step S2, the base mask rate and mask rate adjustment factor for: in, Indicates the upper and lower bounds of the optional mask rate. Indicates the signal-to-noise ratio, reflecting channel state information. Indicates the upper and lower bounds of the selectable signal-to-noise ratio. This represents the sigmoid function. For each source unit The corresponding entropy value, This represents the number of evenly divided intervals within the range of values ​​for the source element.

4. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, In step S3, a weighted sampler based on Gumbel-Softmax is used to obtain the index of the unmasked source units, ensuring that the weighted units are large. It is sampled first, and its sampling weight is: in, The entropy value is The source unit; The mask source y is: ,in, Represented as , The Hadama product represents element-wise multiplication. Represents a matrix containing only 1s; This represents a mask control matrix containing elements 0 and 1.

5. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 3, characterized in that, In step S4, the semantic encoder is: in, Therefore A semantic compression encoder with learnable parameters, containing a feature extractor based on a VisionTransformer architecture for learning masked information sources. An efficient semantic representation, and a linear compressor to further reduce information redundancy in the semantic representation; The semantic features are encoded as follows: , The length of the compressed source feature.

6. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, In step S5, the layer normalization operation is as follows: in, , Encode the semantic features The mean and standard deviation.

7. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, In step S6, the finite scalar quantizer Encoding continuous semantic features Rounding to the nearest whole number. A set of elements The closest value, The quantization interval is defined as follows: the co-distributed modulator maps the quantization levels to QAM constellation points, thereby aligning semantic features with constellation symbols in terms of distribution.

8. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, Step S7 is as follows: (1) Decode the semantic feature vector Source element mask matrix sent to the receiver in advance Fill in: in, This indicates that the decoded semantic vector will be processed according to the index signal. Fill in with the source Same shape, The Hadama product represents element-wise multiplication. This is a mask control matrix containing elements 0 and 1. Represents a matrix containing only 1s; For the shape of A matrix of all 1s For the shape of A matrix of all zero elements; (2) The padded features are fed into the semantic decoder to reconstruct the information source: ,in, Therefore A semantic decoder with learnable parameters.

9. The semantic communication method based on joint variable-length coding and constellation probabilistic shaping according to claim 1, characterized in that, In step S8, the loss function includes a reconstruction term. Reconstruction loss - The minimum mean square error (MSE) is used to represent the gradient pass-through estimation, which is used to solve the problem of semantic feature encoding. The non-differentiability introduced by scalar quantization ensures gradient backpropagation. in, This signifies stopping gradient calculation; that is, when the model performs forward inference, for any input, It can be used as a signal hold; and when the model backpropagates, for any input, The output of all of them is 0.

10. A semantic communication system combining joint variable-length coding and constellation probabilistic shaping, characterized in that, include: A source decomposition and masking control module is used to implement steps S1 to S3 as described in claim 1; A semantic encoding and layer normalization module is used to implement steps S4 to S5 as described in claim 1; A quantization modulation module is used to implement step S6 as described in claim 1; A semantic decoding and reconstruction module is used to implement step S7 as described in claim 1; An end-to-end training module is used to implement step S8 as described in claim 1.