Dynamic environment adaptive semantic communication system and method

Through the combination of non-uniform modulation and hybrid density network, the transmission compatibility and channel adaptation problems of task-oriented semantic communication systems in dynamic environments are solved, and efficient and reliable semantic communication is achieved, meeting the needs of high transmission rates and low latency.

CN120378053APending Publication Date: 2025-07-25SHANDONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510519327.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing task-oriented semantic communication systems have transmission compatibility problems in dynamic environments, insufficient information redundancy and channel adaptability, especially in multimodal semantic transmission and context-dependent decision-making scenarios.

Method used

A dynamic codebook design with non-uniform modulation and a hybrid density network (MDN) is used to build a channel adaptive decoder, map continuous semantic features to discrete codebook space through vector quantization, and quickly update channel feature distribution with Gaussian hybrid model (GMM), design the adaptive layer to compensate channel changes, and achieve efficient transmission and task adaptation.

Benefits of technology

It significantly improves transmission efficiency and task relevance, reduces communication overhead, ensures high transmission reliability and efficient edge inference tasks in dynamic channel environments, and meets the needs of high transmission rates and low latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378053A_ABST
    Figure CN120378053A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic environment adaptive semantic communication system and method, which are used for realizing the whole task-oriented semantic communication, the system comprises a transmitting terminal network, a codebook space, a wireless channel and a receiving terminal network, the transmitting terminal network comprises a feature extractor and a joint source channel encoder, namely a JSC encoder; the feature extractor is used for identifying and extracting features related to a specific task from the original input data; the JSC encoder is used for: mapping the continuous feature vector output by the feature extractor to a discrete symbol representation suitable for wireless transmission; the codebook space is used for mapping a continuous feature vector output by the feature extractor into a discrete codeword index through a vector quantization mechanism; the wireless channel is used for converting the discrete code word index into an electromagnetic signal capable of being physically transmitted and completing a transmission process of the electromagnetic signal in a noise and fading environment; according to the method, the comprehensive performance optimization of task-oriented semantic communication is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dynamic environment adaptive semantic communication system and method, belonging to the field of wireless communication technology. Background Art

[0002] Mobile communication systems have been transformed into infrastructure to support the digital needs of all industries, and the vision of the sixth-generation communication technology (6G) goes far beyond the purpose of communication. Against the backdrop of a significant increase in modern artificial intelligence (AI) terminal services, many new information entertainment devices are becoming the main computing platforms for many applications, generating an unprecedented amount of private data. However, for large-scale data scheduling, the transmission rate is approaching the Shannon limit again, posing severe challenges to the bandwidth resources, strict latency, and transmission reliability requirements of existing communication systems.

[0003] To this end, task-oriented semantic communication (ToSC) has been proposed to break through the "Shannon limit", which is considered a new communication paradigm for 6G mobile communication. Different from traditional bit-level communication systems, this system only extracts and transmits task-related information according to different task requirements and does not require precise bit recovery at the receiving end. On the one hand, the computing resources of the transceiver can be utilized to alleviate the problem of scarce communication resources. On the other hand, when tasks are offloaded to the edge network, edge inference tasks can be executed more flexibly, and efficient and secure interaction between the transceiver ends can be ensured, thereby improving the reliability and real-time performance of the communication system. So far, most of the task-related information extraction and transmission in task-oriented communication are realized based on joint source-channel coding (JSCC) under a deep neural network (DNN). Although the task-oriented JSCC scheme has achieved satisfactory results, there are still transmission compatibility problems when used in the context of digital communication systems. During the process of task-related information extraction and transmission, the data with continuous feature representation needs to be directly transmitted using methods such as analog modulation or constellation diagrams. However, this direct mapping of data to continuous channel input symbols brings a huge communication burden to the transmitter and places greater requirements on radio frequency technology. Therefore, introducing a discrete codebook to map continuous features and discretizing the existing system is an effective way to solve this problem.

[0004] It is worth noting that the training mechanisms of current ToSC systems are generally based on the assumptions of static data distribution and stationary channel models, implicitly requiring the channel state information (CSI) to maintain statistical consistency during the training and inference phases. This idealized assumption is often difficult to hold in an actual dynamic communication environment because the wireless channel has time-varying characteristics due to factors such as multipath fading, mobility interference, and weather changes, resulting in a significant decline in the performance of models trained based on a fixed distribution or even the failure of their functions.

[0005] To address the challenges posed by channel dynamics, Raghuram et al. (2021) proposed a domain adaptation framework based on autoencoders, which enables the communication system to adapt to dynamic changes in channel conditions online by introducing an adversarial training mechanism into the encoder-decoder architecture. By constructing a jointly optimized model that includes a channel feature extraction module and a task semantic recovery module, this method can theoretically reduce the need for frequent retraining in traditional methods due to channel fluctuations.

[0006] However, this task-agnostic end-to-end communication architecture still has key limitations: Although the one-hot encoding it uses simplifies the symbol mapping process, it fails to effectively model the hierarchical latent structure present in semantic information (such as the relevance between semantic elements, task-sensitive feature distributions, etc.), resulting in insufficient information representation ability in complex task scenarios (such as multi-modal semantic transmission, context-dependent decision-making). Specifically, when discretizing semantic symbols into orthogonal vectors, one-hot encoding destroys the inherent topological relationships in the semantic space, making it difficult for the system to capture task-related semantic dependencies, which is particularly prominent in scenarios where it is necessary to dynamically balance transmission efficiency and task completion. Summary of the Invention

[0007] In view of the deficiencies of the prior art, the present invention proposes a dynamically environment-adaptive semantic communication system.

[0008] Specifically, the present invention first considers the requirements for communication resources, transmission delay, and high transmission rate of existing systems in large-scale data transmission, and adopts a task-oriented semantic communication scheme to efficiently transmit content-aware and semantic-related information in a task-oriented manner. This approach can fundamentally solve the problem of information transmission redundancy existing in traditional information transmission-based communication protocols and combine the information bottleneck theory to implement inference tasks under the terminal network.

[0009] To solve the modulation compatibility problem faced by semantic communication in digital communication systems, the present invention innovatively constructs a dynamic codebook design method based on non-uniform modulation. Specifically, continuous semantic features are mapped to a discrete codebook space through vector quantization (VQ), where the probability distribution of the codebook basis vectors is non-uniformly optimized according to the task objective: higher quantization density is assigned to features related to high-frequency tasks, while sparse quantization intervals are used for redundant features. This design breaks through the limitation of the uniform distribution of constellation points in traditional uniform modulation (such as QAM), and makes the transmission energy consumption of discrete indexes negatively correlated with task relevance while ensuring low quantization distortion. When the digital modulation module uses a finite constellation diagram (such as 16-QAM) to transmit the codebook index, the anti-channel fading ability is further enhanced through a non-uniform constellation mapping strategy. The receiving end reconstructs the semantic features based on the quantization codebook and inputs them into the edge inference model to achieve the joint optimization of communication overhead and task performance.

[0010] To address the domain adaptation problem in a dynamic channel environment, the present invention introduces a mixture density network (MDN) to construct a channel adaptive decoder. This module probabilistically models the channel features through a Gaussian mixture model (GMM), where each Gaussian component corresponds to a specific channel state (such as Rayleigh fading, Rice fading, etc.). In the online deployment phase, the system quickly updates the GMM parameters using a small number of channel observation samples to capture the time-varying characteristics of the channel distribution. Based on the learned channel features, an optimal affine transformation layer is designed at the front end of the decoder to compensate for the distribution shift caused by channel changes, ensuring high decoding accuracy and transmission reliability. Experimental results show that this method can reliably transmit task-related features and has robust adaptability to different channel conditions.

[0011] In the present invention, the system first obtains the semantic features of the input data through a feature extractor, and then uses a task-aware joint source-channel (JSC) encoder to compress the features. On this basis, a vector quantization module is introduced to map the encoded features to a discrete semantic embedding space, improving the transmission efficiency through dimensionality reduction and redundancy removal. To enhance task adaptability, a dynamic codebook based on non-uniform modulation is innovatively designed. This codebook keeps the discrete semantic representation highly correlated with specific task objectives through an adaptive optimization mechanism.

[0012] For the transmission optimization process, a digital modulation module is designed to map discrete feature indices to a finite-point constellation diagram, significantly reducing the transmission information volume and communication overhead. At the channel processing level, to address the problem that the system needs to frequently retrain the encoder due to the dynamic changes in the wireless channel (traditional methods have defects such as high data dependence, long time consumption, and high repetition costs), a two-level solution is proposed: First, a Gaussian mixture model (GMM) is used to establish a channel feature expression framework, which can capture the characteristics of the new channel distribution with a small number of samples; at the same time, an Adaptation Layer is deployed at the decoding end, and an optimal inverse affine transformation is implemented based on the new channel-related features to effectively compensate for the feature distribution shift caused by channel fluctuations.

[0013] At the receiving end, the original semantic feature zd(x) is reconstructed through a vector quantization decoding module and the data index is restored. Finally, the task inference network performs edge computing tasks to ensure high-precision inference performance. Through the co-design of channel feature prediction and adaptive decoding, this architecture achieves the goal of quickly adapting to a new channel environment with only a small number of samples, avoiding repeated training of the encoder, significantly reducing system resource consumption, and improving the overall operation efficiency.

[0014] Term Explanation: 1. VQ-VAE (Vector Quantized Variational Autoencoder) training stage: It refers to the process of training a special autoencoder model. This model includes an encoder, a decoder, and a discrete codebook. During training, the encoder maps the input data (such as image patches) to a continuous latent vector, and then through the vector quantization (VQ) step, this continuous vector is replaced with the closest embedding vector (codeword) in the codebook to obtain a discrete latent representation (codebook index). The decoder then attempts to reconstruct the original input from this discrete representation. The goal of training is to minimize the reconstruction error, and the encoder, decoder, and embedding vectors in the codebook are optimized through a specific loss function (usually including reconstruction loss, codebook learning loss, and commitment loss). The key to this stage is to learn a compact discrete codebook that can effectively capture the main features of the data.

[0015] 2. One-hot Encoding: It is a technique that converts categorical variables into a numerical form that is easy for machine learning algorithms to process. For a categorical feature with N possible values (e.g., N indices in a codebook), one-hot encoding creates an N-dimensional vector. In this vector, the position representing the current value (such as the selected codebook index k) is set to 1, while all other N - 1 positions are set to 0. For example, if the codebook size is 5 and the index is 3, its one-hot encoding will be the vector [0, 0, 1, 0, 0]. This encoding method avoids the non-existent order relationships that may be introduced by mapping categories to ordered numerical values (such as 1, 2, 3...), making each category orthogonal in the vector space. In the present invention, it may be used to represent the selected discrete codeword index for input into subsequent neural network layers or for channel coding.

[0016] 3. Multilayer Perceptron (MLP): It is a feedforward artificial neural network model. It contains at least three layers of nodes: an input layer, one or more hidden layers, and an output layer. Except for the input nodes, each node is a neuron with a non-linear activation function (such as ReLU, Sigmoid, or Tanh). MLP conducts information transmission through full connections between layers, that is, each node is connected to all nodes in the next layer. By learning connection weights and biases, MLP can approximate complex non-linear functions and is commonly used in classification, regression, and feature extraction and transformation tasks. In the present invention, MLP can be used to implement the non-linear transformation in the JSC encoder to project the feature vector into a low-dimensional latent space.

[0017] 4. Convolutional Neural Network (CNN): It is a class of deep learning models particularly suitable for processing grid-structured data (such as images). Its core is the convolutional layer, which extracts spatial hierarchical features (such as edges, textures, shapes, etc.) by performing sliding convolution operations on the input data using learnable filters (convolution kernels). CNN usually also includes a pooling layer for reducing the spatial dimension of the feature map and enhancing the translational invariance of the model, as well as a fully connected layer for the final classification or regression output. Due to its characteristics of weight sharing and local connection, CNN is highly parameter-efficient when processing high-dimensional data and can effectively capture local patterns. In the present invention, CNN is used as a feature extractor to extract semantic features from the original input data (such as images).

[0018] The technical solution of the present invention is as follows: A dynamic environment adaptive semantic communication system for realizing the entire task-oriented semantic communication, including: a transmitting end network, a codebook space, a wireless channel, and a receiving end network, where the transmitting end network includes a feature extractor and a joint source-channel encoder, i.e., a JSC encoder; The feature extractor is used to: identify and extract features related to a specific task from the original input data; The JSC encoder is used to: map the continuous feature vectors output by the feature extractor to discrete symbol representations suitable for wireless transmission; The codebook space is used to: map the continuous feature vectors output by the feature extractor to discrete codeword indices through a vector quantization mechanism; The wireless channel is used to: convert the discrete codeword indices into physically transmittable electromagnetic signals and complete the transmission process of the electromagnetic signals in a noisy and fading environment; The receiving end network is used to: complete signal demodulation, semantic feature reconstruction, and task inference.

[0019] According to a preferred embodiment of the present invention, the receiving end network includes a vector quantization decoding module, an adaptive layer, and an edge task inference module; The vector quantization decoding module is used to: demodulate the received electromagnetic signals into discrete codeword indices and reconstruct the semantic feature vector zd(x) based on the mapping relationship of the codebook space; The adaptive layer is used to: dynamically compensate for the feature distribution offset caused by channel changes; based on the channel distribution characteristics predicted by the Gaussian mixture model, the adaptive layer quickly optimizes the feature transformation parameters through a small number of new channel samples to align the reconstructed semantic features with the task requirements under the current channel conditions; The edge task inference module is used to: execute terminal tasks on the reconstructed semantic features.

[0020] According to a preferred embodiment of the present invention, identifying and extracting features related to a specific task from the original input data; mapping the continuous feature vectors output by the feature extractor to discrete symbol representations suitable for wireless transmission; as shown in Equation (I): ; Wherein, is a jointly designed feature extractor and JSC encoder, with the original input data as the input , and the parameters are ; is a set of trainable parameters including all the weights and bias terms of the feature extractor and JSC encoder; is a vector, each element of which corresponds to an encoded symbol, represents the discrete symbol representation obtained after passing through the feature extractor and JSC encoder, and each vector dimension is D.

[0021] More preferably, features related to a specific task are identified and extracted from the original input data; including: The feature extractor is a convolutional neural network CNN; the convolutional neural network CNN includes a convolutional layer, a pooling layer, and a fully connected layer; The conversion from the original input data to continuous semantic features is achieved through the convolutional neural network CNN; specifically including: the convolutional layer extracts the spatial hierarchical features of the image through the local receptive field and the weight sharing mechanism, the pooling layer performs spatial downsampling, and finally the fully connected layer generates a continuous semantic feature vector with a fixed dimension.

[0022] According to the preference of the present invention, the continuous feature vector output by the feature extractor is mapped to a discrete symbol representation suitable for wireless transmission; including: First, the continuous feature vector is non-linearly transformed through a multi-layer perceptron or a convolutional network, and the continuous feature vector is projected into a low-dimensional latent space to obtain a continuous latent vector; the specific implementation process is: the continuous feature vector is input into the fully connected layer of the multi-layer perceptron or the convolutional layer of the convolutional network, and the high-order feature correlations are extracted layer by layer through a non-linear activation function, and a low-dimensional continuous latent vector is output; Secondly, a codebook-based vector quantization mechanism is adopted, and the continuous latent vector is mapped to a discrete symbol index through nearest neighbor search, and the codebook is trained by jointly optimizing the source compression rate and the channel noise resistance; Finally, one-hot encoding is performed on the discrete symbol index to obtain a discrete symbol representation.

[0023] According to the preference of the present invention, a codebook-based vector quantization mechanism is adopted, and the continuous latent vector is mapped to a discrete symbol index through nearest neighbor search; including: Calculate the Euclidean distance between the current continuous latent vector and all the embedded vectors in the pre-trained codebook. The embedded vectors in the pre-trained codebook refer to a set of fixed vectors learned through the VQ-VAE training stage, and each embedded vector represents a discrete codeword; Determine the codeword index with the highest matching degree through nearest neighbor search; including: Select the index of the embedded vector with the smallest Euclidean distance, which is the obtained discrete symbol index.

[0024] According to the preference of the present invention, the discrete codeword index is converted into an electromagnetic signal that can be physically transmitted, and the transmission process of the electromagnetic signal in a noise and fading environment is completed; including: Map the discrete codeword index to a specific constellation symbol through digital modulation; each constellation point is represented in complex form as: s = I + jQ, where I and Q are the real and imaginary parts respectively, corresponding to the in-phase component and the quadrature component; convert the discrete digital signal, i.e., the constellation point s, into an electromagnetic signal that can be physically transmitted; adopt Gray coding to design the constellation mapping rule to ensure that only one bit differs between adjacent constellation points. Complete signal demodulation to obtain the discrete index value; and retrieve the corresponding codeword from the locally stored codebook based on this discrete index value, and finally input the discrete feature representation into the edge inference module; the specific implementation process includes: 1) Compensate for signal attenuation and phase offset through channel estimation. 2) Calculate the bit probability of the received symbol based on the maximum likelihood criterion or soft demodulation algorithm. 3) Restore the discrete index through hard decision or soft input decoding. 4) Extract the corresponding compressed feature vector from the codebook according to the index.

[0025] Preferably according to the present invention, the adaptive layer quickly optimizes the feature transformation parameters through a small number of new channel samples based on the channel distribution characteristics predicted by the Gaussian mixture model, so that the reconstructed semantic features are aligned with the task requirements under the current channel conditions; including: Adopt a generative channel model to approximately simulate the probability density of the real channel conditions. ; is the parameter of the generative channel model, and the conditional density of the channel is modeled by a set of Gaussian mixture models, as follows: ; Among them, k is the number of components, is the mean vector, is the covariance matrix, and [0, 1] is the prior probability of component i, and the prior probability of the component is represented by the softmax function; refers to the conditional probability density function, which represents the probability density that the signal / feature actually observed at the receiving end is x under the condition that the discrete symbol / index sent by the transmitting end is z. is the parameter set of the generative channel model; x represents the observed signal or feature vector at the receiving end; z represents the discrete symbol or codebook index sent by the transmitting end. Adjust the mean, covariance, and mixing weights of each Gaussian component in the original GMM model through a set of specific transformation parameters to match the statistical characteristics of the target channel; these transformation parameters include: Affine transformation parameters of the mean: For the mean vector of each Gaussian component, the affine transformation is composed of a linear transformation matrix and an offset vector Composed of, the transformed mean value (z) is calculated by the following formula: (z) = (z) + ; Wherein, is the matrix for linearly transforming the mean value of the i-th Gaussian component; is the offset vector or translation vector for the mean value of the i-th Gaussian component; Affine transformation parameters of the covariance: For the covariance matrix of each Gaussian component, the affine transformation is implemented by a transformation matrix The transformed covariance is calculated by the following formula: = ; Wherein, is the covariance matrix of the i-th component in the original GMM model given z; is the new covariance matrix of the i-th component in the GMM model adapted to the new channel given z after transformation; is the matrix for linearly transforming the covariance of the i-th Gaussian component; Affine transformation parameters of the mixing weights: The mixing weight or prior probability of each Gaussian component is transformed by a scaling factor and an offset to obtain the adjusted weight; the transformed prior logit (z) is calculated by the following formula: (z) = (z) + ; (z) is the logit value corresponding to the mixing weight of the i-th component in the original GMM model; is the scaling factor for the logit value of the i-th component; is the offset for the logit value of the i-th component.

[0026] According to the preference of the present invention, the transmitted signal experiences fading on the wireless channel. Therefore, at the receiving end, it is expressed as: ; Wherein, represents the demodulated symbol, ; is the demodulator function; CN indicates following a complex Gaussian distribution, represents the covariance matrix of the distribution, is the variance, is the identity matrix.

[0027] Further preferably, the discrete feature representation reconstructed by the receiver network is input into the edge task inference module for edge inference tasks; as follows: ; wherein, is the edge task inference module with the parameter set , is the result label output by the edge task inference module.

[0028] A dynamic environment adaptive semantic communication method includes: Identifying and extracting features related to a specific task from the original input data; Mapping the continuous feature vector output by the feature extractor to a discrete symbol representation suitable for wireless transmission; Mapping the continuous feature vector output by the feature extractor to a discrete codeword index through a vector quantization mechanism; the wireless channel is used to: convert the discrete codeword index into an electromagnetic signal that can be physically transmitted and complete the transmission process of the electromagnetic signal in a noisy and fading environment; Completing signal demodulation, semantic feature reconstruction, and task inference.

[0029] The beneficial effects of the present invention are: Aiming at the deficiencies of task-oriented semantic communication in terms of compatibility, information redundancy, communication overhead, and channel adaptation ability in the existing technology, an innovative few-shot learning adaptive method, FA-NFM, is proposed, which can effectively balance the transmission utility and communication resource requirements. The present invention significantly improves the transmission efficiency and task relevance by introducing a codebook design with non-uniform modulation and dynamically optimizing the codebook to better adapt to specific task objectives. By designing a non-uniform feature codebook mapping and a vector quantization mechanism, data samples are encoded into discrete representations, and combined with constellation modulation and a digital modulation module with a finite-point constellation, the amount of information and transmission overhead in wireless communication are significantly reduced while ensuring data transmission compatibility. In addition, the present invention uses a mixture density network (MDN) combined with a Gaussian mixture model (GMM) to model channel characteristics, quickly captures changes in the channel distribution through few-shot learning, and designs an optimal feature transformation for channel changes to compensate for distribution offsets, thereby ensuring the efficiency and robustness of edge inference tasks. Based on the information bottleneck theory, the present invention can meet the requirements of high transmission rate and low latency while reducing information redundancy and communication burden, and ensure transmission reliability and decoding accuracy under different channel conditions through the ability to adapt to dynamic channel changes, ultimately achieving the comprehensive performance optimization of task-oriented semantic communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is an architecture diagram of a dynamic environment adaptive semantic communication system of the present invention; Figure 2 is a PSNR performance diagram under the condition of transitioning from flat fading to Rice fading; Figure 3 is a PSNR performance diagram under the condition of transitioning from Rice fading to flat fading; Figure 4 is a PSNR performance diagram under the condition of transitioning from Gaussian white noise to flat fading; Figure 5 is a PSNR performance diagram under different adaptive sample count conditions. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be further limited below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0032] Embodiment 1 A dynamic environment adaptive semantic communication system, as Figure 1 shown, for implementing the entire task-oriented semantic communication, includes: a transmitting end network, a codebook space, a wireless channel, and a receiving end network. The transmitting end network includes a feature extractor and a joint source-channel encoder, namely a JSC encoder; The feature extractor is used to: identify and extract features related to a specific task from the original input data; for example, in an image classification task, the feature extractor can be a Convolutional Neural Network (CNN), which can capture key visual features in the image, such as edges, textures, and shapes, etc.

[0033] The JSC encoder is used to: map the continuous feature vectors output by the feature extractor to discrete symbol representations suitable for wireless transmission; the JSC encoder not only considers the requirements of source coding (i.e., how to effectively compress information), but also takes into account the requirements of channel coding (i.e., how to reliably transmit information in a noisy channel). Therefore, the JSC encoder needs to be designed to optimize both the information compression efficiency and the transmission reliability simultaneously.

[0034] The codebook space is used to: map the continuous feature vectors output by the feature extractor to discrete codeword indices through a Vector Quantization (VQ) mechanism; its core role is to establish a bridge between semantic information and digital modulation symbols.

[0035] The wireless channel is used to: convert the discrete codeword indices into physically transmissible electromagnetic signals and complete the transmission process of the electromagnetic signals in a noisy and fading environment; The receiving - end network is used to: complete signal demodulation, semantic feature reconstruction, and task inference.

[0036] Embodiment 2 A dynamic - environment - adaptive semantic communication system, the receiving - end network includes a vector quantization decoding module, an adaptive layer, and an edge task inference module; The vector quantization decoding module is used to: demodulate the received electromagnetic signal into discrete codeword indices and reconstruct the semantic feature vector zd(x) based on the mapping relationship of the codebook space; specifically, it includes: First, the receiving - end receives and demodulates the wireless signal, and recovers the codeword index (a digital number) sent by the transmitting - end from it. Then, the vector quantization decoding module uses this recovered digital number to look up the corresponding specific feature vector in a pre - defined "codebook" (equivalent to an encoding dictionary) stored locally; finally, output this found feature vector as the reconstructed semantic information for subsequent task processing. The vector quantization decoding module restores the transmitted discrete symbols to continuous semantic representations through inverse vector quantization operations and corrects the symbol errors introduced by channel noise;

[0037] The Adaptation Layer is used for: dynamically compensating for the feature distribution shift caused by channel variations; based on the channel distribution characteristics predicted by the Gaussian Mixture Model (GMM), the Adaptation Layer quickly optimizes the feature transformation parameters through a small number of new channel samples, so that the reconstructed semantic features are aligned with the task requirements under the current channel conditions; for example, in the channel fading scenario, the Adaptation Layer can offset the impact of channel distortion on the classification features through the affine transformation in the feature space.

[0038] The Edge Task Inference Module is used for: performing terminal tasks (such as classification, detection, or prediction) on the reconstructed semantic features. This module adopts a lightweight neural network architecture and optimizes the inference logic in combination with task relevance constraints to ensure high-precision inference performance under feature compression and channel perturbation conditions. For example, in the image classification task, this module can directly output the classification probability distribution based on the reconstructed semantic features without fully reconstructing the original image data.

[0039] Identify and extract features related to a specific task from the original input data; map the continuous feature vectors output by the feature extractor to a discrete symbol representation suitable for wireless transmission; as follows: ; where, is the jointly designed feature extractor and JSC encoder, with the original input data as the input , and the parameters are ; is the set of trainable parameters including all the weights and bias terms of the feature extractor and JSC encoder; is a vector, each element of which corresponds to an encoded symbol, represents the discrete symbol representation obtained after passing through the feature extractor and JSC encoder, and each vector dimension is D.

[0040] The feature extractor and JSC encoder of the transmitter network are jointly optimized in a jointly designed form. The entire network will perform end-to-end learning according to task objectives (such as classification accuracy) and communication performance metrics (such as bit error rate). In this way, the network can automatically adjust its internal structure and parameters to achieve the best feature extraction and encoding effects. An effective conversion from the original input to the encoded symbols is achieved, and is the final output result of this conversion process.

[0041] Identify and extract features related to a specific task from the original input data; including: The feature extractor is a Convolutional Neural Network (CNN); the CNN includes a convolutional layer, a pooling layer, and a fully connected layer; The CNN is used to transform the original input data (such as image pixel data) into continuous semantic features; specifically: the convolutional layer extracts the spatial hierarchical features of the image (such as edges, textures, etc.) through local receptive fields and weight sharing mechanisms, the pooling layer performs spatial downsampling to enhance translational invariance, and finally the fully connected layer generates a continuous semantic feature vector of a fixed dimension. This design utilizes the efficient representation ability of CNN for image data, suppressing redundant background noise while retaining task-related semantic information.

[0042] Map the continuous feature vector output by the feature extractor to a discrete symbol representation suitable for wireless transmission; including: First, perform a non-linear transformation on the continuous feature vector through a multi-layer perceptron or a convolutional network, project the continuous feature vector into a low-dimensional latent space to obtain a continuous latent vector; the specific implementation process is: input the continuous feature vector into the fully connected layer of a multi-layer perceptron (MLP) or the convolutional layer of a convolutional network, and extract high-order feature correlations layer by layer through a non-linear activation function (such as ReLU), and output a low-dimensional continuous latent vector; this process compresses the high-dimensional features into a compact semantic representation by optimizing network parameters, while retaining key task-related information to improve transmission efficiency.

[0043] Second, adopt a codebook-based vector quantization mechanism to map the continuous latent vector to a discrete symbol index through nearest neighbor search, and this codebook is trained by jointly optimizing the source compression rate and channel noise resistance; introducing the vector quantization mechanism maps the encoded features to a discrete semantic embedding space, improving transmission efficiency through dimensionality reduction and redundancy removal.

[0044] Finally, use the principle of channel coding to perform one-hot encoding on the discrete symbol index to obtain a discrete symbol representation.

[0045] Adopt a codebook-based vector quantization mechanism to map the continuous latent vector to a discrete symbol index through nearest neighbor search; including: Calculate the Euclidean distance between the current continuous latent vector and all embedding vectors in the pre-trained codebook. The embedding vectors in the pre-trained codebook refer to a set of fixed vectors learned during the VQ-VAE training phase, and each embedding vector represents a discrete codeword; for example, in an image compression task, these vectors may correspond to typical feature patterns of image blocks; Determine the codeword index with the highest matching degree through Nearest Neighbor Search, including: Select the embedding vector index with the smallest Euclidean distance, which is the obtained discrete symbol index.

[0046] In low-complexity scenarios, it is implemented through exhaustive search, and in high-dimensional spaces, the Approximate Nearest Neighbor (ANN) algorithm can be used to accelerate. The embedding vectors in the codebook are optimized by backpropagation of quantization error during the training phase (such as using a straight-through estimator to solve the gradient propagation problem), and their distribution characteristics are closely related to the task requirements. For example, in an image classification task, the frequently occurring target contour features correspond to a higher density of codeword distribution, while the low-frequency detail features use sparse quantization.

[0047] To enhance task adaptability, a dynamic codebook based on non-uniform modulation is innovatively designed. Through an adaptive optimization mechanism, this codebook enables the discrete semantic representation to be highly correlated with specific task objectives.

[0048] This dynamic codebook scheme based on non-uniform modulation is mainly applied in the transmitter network, specifically inside or immediately after the JSC encoder. When the JSC encoder maps continuous feature vectors to the latent space, vector quantization is required to discretize them into codeword indices for transmission. This dynamic codebook is the codebook used for this quantization process.

[0049] Specific form ("dynamic codebook based on non-uniform modulation"): Non-uniformity: It means that the distribution of embedding vectors (codewords) in the latent space in the codebook is not uniform. Traditional VQ codebooks may make the codewords roughly evenly cover the data distribution area through methods such as K-means, or be learned in VQ-VAE. The "non-uniform" emphasizes that the distribution density of codewords is specifically designed according to task requirements and / or channel characteristics. For example, in the feature regions highly related to the task, the corresponding codewords will be denser and the quantization accuracy will be higher; while in the feature regions less related to the task or redundant, the codewords will be sparser and the quantization will be coarser.

[0050] Dynamic / Adaptive Optimization Mechanism: "Dynamic" may mean that the codebook can adapt online to changes (although this invention mainly emphasizes that a small number of samples adapt to the channel, but theoretically the codebook can also adapt to task changes). More importantly, its structure is formed through an "adaptive optimization mechanism". This mechanism means that during the training phase, the structure of the codebook (the position of the codewords \(e_k\) and their possible usage probabilities) is jointly optimized with the goal of the entire end-to-end system (e.g., maximizing task accuracy, minimizing transmission overhead or distortion). The optimization algorithm (such as gradient-based optimization) will adjust the position and distribution of the codewords so that the features that contribute greatly to the final task can be more accurately quantized and transmitted, while the features that are unimportant for the task use fewer bits or resources. This reflects the task-aware characteristic.

[0051] Association with Modulation: "Non-uniform modulation" may refer to two levels here: One is that the non-uniformity of the codebook itself affects the probability distribution of the symbols (indices) of the subsequent digital modulation input; the other is that it may be combined with geometric / probability constellation shaping technology so that the probability distribution of the transmitted constellation points is also non-uniform to approach the channel capacity. However, this part is described separately later in the original text. The focus here should be on the non-uniform structure of the codebook itself.

[0052] How to Make the Discrete Semantic Representation Highly Associated with the Specific Task Goal through the Adaptive Optimization Mechanism: During the end-to-end training process, the loss function directly reflects the final task performance (such as the classification error rate). When the feature region represented by a certain codeword contributes greatly to reducing the task loss, the optimization algorithm (through gradient backpropagation) will tend to: 1) map more data points near \(e_k\) (adjust the encoder); 2) may allocate more codewords around \(e_k\) to improve the quantization accuracy of this region; 3) ensure that the representation of \(e_k\) can be effectively utilized by the decoder and the task inference module. Conversely, if the feature represented by a certain codeword has little impact on the task performance, the optimization process may reduce the resources allocated to it (such as making it cover a larger region or reducing its usage probability). In this way, the codebook structure is adaptively associated with the task goal.

[0053] Example: Consider a traffic sign recognition task in autonomous driving. After the input image passes through the feature extractor, a feature vector containing information such as vehicles, roads, the sky, and traffic signs is obtained. Through the quantization of the adaptively optimized non-uniform codebook:

[0054] For the potential vector region representing the key features of the "stop" sign (such as the octagonal shape, red background, and white font), the codebook will allocate dense codewords to ensure that these key information are accurately quantized and transmitted.

[0055] For the feature regions representing background information such as the sky and ordinary road surface textures, the codebook will allocate sparse codewords for rough quantization because these information contribute less to the task of traffic sign recognition.

[0056] During the training process, if the system finds that the recognition rate of the "speed limit" sign is not high, the optimization mechanism may adjust the codebook to increase the codeword density in the region representing the digital features of the speed limit sign. The structure of the finally formed codebook (the density distribution of codewords) is highly correlated with the specific task objective of "traffic sign recognition".

[0057] Convert the discrete codeword index into an electromagnetic signal that can be physically transmitted and complete the transmission process of the electromagnetic signal in a noisy and fading environment; including: Transmitter: Source coding → Channel coding → Digital modulation → RF front-end → Wireless channel; Receiver: Wireless channel → RF front-end → Digital demodulation → Channel decoding → Source decoding; Map the discrete codeword index to specific constellation symbols through digital modulation; for example, the 16-QAM constellation diagram consists of 16 points, and each point corresponds to a 4-bit binary sequence (such as 0000 to 1111). Through Gray Coding, the binary sequences between adjacent constellation points differ by only one bit, thus reducing the bit error rate. Each constellation point is represented in complex form as: s = I + jQ, where I and Q are the real and imaginary parts respectively, corresponding to the in-phase component and the quadrature component; convert the discrete digital signal, that is, the constellation point s (belonging to the baseband digital signal) into an electromagnetic signal that can be physically transmitted through the RF front-end; to enhance the anti-channel interference ability, the system adopts Gray Coding to design the constellation mapping rule to ensure that adjacent constellation points differ by only one bit; for example: the natural binary code 0000 and 0001 correspond to adjacent points in the constellation diagram, but Gray Coding converts the binary code into Gray code through an exclusive OR operation (such as natural code 0000 → Gray code 0000, natural code 0001 → Gray code 0001), so that the Gray codes of adjacent constellation points differ by only 1 bit, thus reducing the multi-bit error probability when symbol misjudgment occurs; thereby reducing the impact of symbol misjudgment on semantic reconstruction. Design digital modulation to map the discrete feature index to a finite-point constellation diagram, significantly reducing the transmission information volume and communication overhead.

[0058] Complete signal demodulation to obtain discrete index values; retrieve the corresponding codewords from the locally stored codebook based on these discrete index values, and finally input the discrete feature representation into the edge inference module; the discrete feature representation refers to the embedding vector corresponding to the discrete codeword index mapped from the continuous latent vector through vector quantization (VQ) (such as the pre-trained vector in the codebook). It corresponds to the discrete symbol index after quantization by VQ-VAE at the transmitter and is the compressed expression of semantic information in the low-dimensional space. For example, in image tasks, the discrete feature representation corresponds to the quantization encoding of the typical semantic patterns (such as edges, textures) of image patches. The specific implementation process includes:

[0059] 1) Compensate for signal attenuation and phase offset through channel estimation; 2) Calculate the bit probabilities of the received symbols based on the maximum likelihood criterion or soft demodulation algorithms (such as log-likelihood ratio, LLR); 3) Recover the discrete index through hard decision or soft-input decoding (such as LDPC decoding); 4) Extract the corresponding compressed feature vector (such as the embedding vector of VQ-VAE) from the codebook according to the index.

[0060] The edge inference module usually adopts a lightweight neural network structure (such as a small convolutional neural network or a fully connected network) to adapt to the computing resource limitations of edge devices.

[0061] The discrete codebook of the transceiver is represented as , where represents the total number of vectors in the codebook. Each vector ek is a D-dimensional embedding vector.

[0062] Use digital modulation methods to map each index of the discrete codebook to constellation symbols and ensure coupling the discrete space to the modulation constellation, including: Convert each index of the discrete codebook into a transmitted complex signal , and perform digital modulation and point constellation design: First, convert the codebook index into a complex symbol through a predefined constellation mapping rule (such as QAM, PSK, etc.). For example, the 16-QAM constellation maps a 4-bit index to 16 discrete points on the two-dimensional complex plane; Secondly, design a non-uniform or geometrically shaped constellation diagram according to the channel characteristics, and optimize the constellation point distribution by maximizing the minimum Euclidean distance or minimizing the average symbol energy; Finally, perform pulse shaping (such as root-raised cosine filtering) and carrier modulation (such as OFDM subcarrier mapping) on the complex symbol to form a time-frequency domain transmission signal. The specific implementation includes three core links:

[0063] 1) Constellation mapping: Establish a bijective relationship between the index and the complex symbols, and use Gray code encoding to reduce the bit error rate of adjacent symbols; 2) Geometric shaping: Match the channel capacity boundary through probabilistic shaping or amplitude-phase optimization (APSK); 3) Adaptive modulation: Dynamically adjust the modulation order (such as QPSK → 64QAM) and code rate according to the channel state information (CSI) to achieve a dynamic balance between spectral efficiency and reliability. This process requires joint optimization of codebook training and channel coding constraints to ensure the transmission robustness of discrete symbols in AWGN or fading channels.

[0064] At the channel processing level, aiming at the problem that the dynamic changes commonly existing in wireless channels cause the system to frequently retrain the encoder (traditional methods have defects such as high data dependence, long time consumption, and high repetition cost), a two-level solution is proposed: First, use the Gaussian mixture model (GMM) to establish a channel feature expression framework, and the new channel distribution characteristics can be captured through a small number of samples; at the same time, deploy an adaptation layer at the decoding end, and perform the optimal inverse affine transformation based on the channel-related features of the new channel to effectively compensate for the feature distribution offset caused by channel fluctuations.

[0065] Aiming at the problem that the dynamic changes of wireless channels cause the system to frequently retrain the encoder (traditional methods consume high resources and have long time delays due to relying on a large amount of data and repeated training), the present invention proposes a two-level channel adaptation scheme: Channel feature modeling and rapid update: Use the Gaussian mixture model (GMM) to model the channel distribution, and learn the update rule of channel parameters from a small number of new channel samples through small sample learning; for example, when the channel switches from Rayleigh fading to Rice fading, the GMM can adaptively adjust the weights and covariances of each Gaussian component to capture the statistical characteristics of the new channel.

[0066] Inverse affine transformation compensation mechanism: Deploy an adaptation layer at the receiving end, and perform an inverse affine transformation on the semantic features output by the encoder based on the channel feature parameters extracted by the GMM. This operation aligns the feature distribution distorted by the channel to the original channel distribution assumed at the time of encoder design, so that the influence of channel changes can be eliminated without modifying the encoder. For example, if the new channel introduces an additional phase shift, the adaptation layer restores the consistency of the feature distribution through reverse phase correction.

[0067] Based on the channel distribution characteristics predicted by the Gaussian mixture model (GMM), the adaptation layer quickly optimizes the feature transformation parameters through a small number of new channel samples, so that the reconstructed semantic features are aligned with the task requirements under the current channel conditions; including: In wireless channel modeling, in order for the neural networks at the receiving and transmitting ends to learn through a method based on gradient descent (SGD-based optimization), a differentiable reverse path needs to be created from the decoder to the encoder. The present invention proposes a generative channel model in the form of a mixture probability density function, and uses the generative channel model to approximately simulate the probability density of the real channel conditions ; is the parameter of the generative channel model, which can accurately simulate random factors such as noise and phase offset in the channel. In this work, the conditional density of the channel is modeled by a set of Gaussian Mixture Models (GMMs) as follows:

[0068] ; where k is the number of components, is the mean vector, is the covariance matrix, and [0, 1] is the prior probability of component i, and the prior probability of the component is represented by the softmax function; refers to the Conditional Probability Density Function, which represents the probability density that the signal / feature actually observed at the receiving end is x under the condition that the discrete symbol / index sent by the transmitting end is z; is the parameter set of the generative channel model (described by GMM); x represents the observed signal or feature vector at the receiving end; this is usually the signal representation after being transmitted through the wireless channel and preliminarily processed by the receiver (such as demodulation, but possibly before the adaptive layer). z represents the discrete symbol or codebook index sent by the transmitting end; it is generated by the JSC encoder and represents the semantic information of the original input data xinput. : represents the mean vector of the i-th Gaussian component in the GMM, and this mean depends on (is conditioned on) the transmitted symbol z. It means that for different transmitted symbols z, the influence of the channel on the signal (manifested as the means of the GMM components) may be different.

[0069] Through this modeling method, the system can learn the characteristics of the channel in an end-to-end manner without prior knowledge of the exact physical model of the channel. In addition, the Gaussian mixture model also has strong approximation characteristics and analytical and computational feasibility, and these advantages are beneficial for adaptation between different fields and channel modeling.

[0070] Since any two Gaussian distributions can be transformed into each other through an affine transformation, the present invention utilizes this property and uses an affine transformation. Thus, without directly changing the architecture or all parameters of the Gaussian mixture distribution model, the existing model can be adjusted to adapt to the new data distribution. In fact, it is a process of a Gaussian mixture model adapting from the original channel characteristics to the new channel conditions.

[0071] When adapting a Gaussian mixture model to new features, the basic idea is to adjust the mean, covariance, and mixing weights of each Gaussian component in the original GMM model (which refers to the Gaussian mixture model (GMM) trained or learned under the source channel conditions when the system is initially trained or deployed. That is, before the channel changes (entering the new target channel), the GMM model used by the system to describe the statistical characteristics of the channel at that time. The goal of adaptation is to adjust this "original model" so that it can match the characteristics of the "new channel" (target domain)) through a set of specific transformation parameters to match the statistical characteristics of the target channel; these transformation parameters include:

[0072] Affine transformation parameters for the mean: For the mean vector of each Gaussian component, the affine transformation consists of a linear transformation matrix and an offset vector The transformed mean (z) is calculated by the following formula: (z) = (z) + ; where, refers to the linear transformation matrix (Linear Transformation Matrix) for the mean of the i-th Gaussian component; it is responsible for rotating, scaling, or shearing the original mean vector; refers to the offset vector (Offset Vector) or translation vector (Translation Vector) for the mean of the i-th Gaussian component; it is responsible for translating the result of the linear transformation. and are the affine transformation parameters that need to be learned through a small number of new channel samples;

[0073] Affine transformation parameters for the covariance: For the covariance matrix of each Gaussian component, the affine transformation is achieved by a transformation matrix which multiplies the covariance matrix on the left and right to achieve transformation and rescaling. The transformed covariance is calculated by the following formula:

[0074] = ; wherein, refers to the covariance matrix of the i-th component in the original GMM model under the given z; refers to the new covariance matrix of the i-th component in the GMM model adapted to the new channel under the given z after transformation; refers to the matrix for linear transformation of the covariance of the i-th Gaussian component (Linear Transformation Matrix); it describes how the feature space is linearly stretched or compressed.

[0075] Affine transformation parameters of the mixing weights: The mixing weight or prior probability of each Gaussian component is transformed by the scaling factor and the offset to obtain the adjusted weight; the transformed prior logic (z) is calculated by the following formula: (z) = (z) + ; (z) refers to the logit value corresponding to the mixing weight of the i-th component in the original GMM model (i.e., the input of the softmax function); refers to the scaling factor (Scaling Factor) for the logit value of the i-th component; refers to the offset (Offset) for the logit value of the i-th component.

[0076] The decoder front-end designs an optimal affine transformation layer to compensate for the distribution shift caused by channel changes, ensuring high decoding accuracy and transmission reliability.

[0077] Use feature transformation to adapt the decoder, especially in the scenario of few-shot domain adaptation. Calculate an efficient feature transformation , which can transform the input of the target domain (target domain features) into features more consistent with the source distribution (source domain features). This transformation is based on the optimal affine transformation parameters , and no changes need to be made to the trained encoder and decoder networks. The transformed Gaussian distribution can be obtained through the inverse transformation formula:

[0078] ; This inverse affine transformation consists of several components: is the covariance matrix 's inverse matrix, where i refers to the i-th component in the Gaussian mixture model.

[0079] is the affine transformation matrix.

[0080] is the bias vector of the affine transformation.

[0081] is the mean vector of the i-th Gaussian distribution conditional on the latent variable z.

[0082] The derivation of the inverse transformation formula is as follows: 1. Transformation definition: Define a transformation , which can map source domain features to the target domain for conversion between different domains. This transformation is to solve the inverse problem from to .

[0083] 2. Inverse transformation: To transform from to , we need the inverse of , that is .

[0084] 3. Affine transformation: A general affine transformation includes a linear transformation A and a translation transformation b. Here, A is replaced by , and is the bias vector of the affine transformation.

[0085] 4. Translation: The translation part in the transformation is represented by , which translates the transformation to the appropriate position.

[0086] 5. Repositioning: By adding , the reference point of the transformation is repositioned to the mean of the source Gaussian distribution.

[0087] This transformation requires knowing the transmitted symbol z and the mixture component i, which are not observable at the decoder. We solve this problem by considering the joint posterior distribution, which is based on the target Gaussian mixture, and the formula is:

[0088] ; Here is the conditional probability, representing the probability that the symbol z and the component i occur simultaneously given . P represents probability, and are the model parameters.

[0089] In the numerator of the formula: is the prior probability of the symbol z appearing. is the probability of selecting the i-th component given the symbol z. Given the symbol z and the component i, is the probability density function, where is the mean, is the covariance matrix.

[0090] The denominator of the formula is the sum of the numerator parts for all possible and j, ensuring that the sum of all conditional probabilities is 1. This is a normalization factor that guarantees an effective probability distribution.

[0091] According to Bayes' theorem, is expressed as multiplied by the prior probability and , and then divided by the marginal probability . For the mixture model, can be represented as the conditional Gaussian distribution . The marginal probability is calculated by summing over all and j possible combinations of .

[0092] The expected inverse feature transformation is defined at the decoder as: ; The above formula is derived as follows: expresses the expectation of the feature transformation from the target domain back to the source domain given the channel output . Here, the conditional probability is used to calculate the expected value of the inverse affine transformation .

[0093] This expectation is a weighted average of the inverse affine transformations for all possible symbols z and all possible components i associated with each symbol, with weights being the conditional probabilities .

[0094] This process can be decomposed into the following steps: Inverse affine transformation : For each symbol z and component i, calculate the inverse affine transformation , which maps the features in the target domain back to the feature space of the source domain according to the previously defined method.

[0095] Calculate the conditional probability : This probability reflects the probability that symbol z and component i occur simultaneously after observing the channel output .

[0096] Calculation of the expected value: The expected value is the weighted average of the inverse affine transformations of all possible z and i , with the weights being the corresponding conditional probabilities .

[0097] Mathematical expression of the weighted average: Multiply all by their respective conditional probabilities , and then add them together. This formula gives a feature transformation in the form of a weighted average, which takes into account the contributions of all possible z and i to calculate the features that are most likely to correspond to the source domain distribution at the decoder end.

[0098] This process is essentially a post - processing process of the channel output to better align the features of the source distribution, thereby improving the adaptability and recognition ability of the decoder to the incoming signal. This method can effectively compensate for the feature distribution shift caused by channel fluctuations.

[0099] In the adaptive layer, using the channel information learned from a small number of new samples, an inverse affine transformation is performed on the demodulated features to cancel the influence of channel changes. The transformed features are then fed into the subsequent vector quantization decoding module for final semantic feature reconstruction and task inference. Simply put, it is a key step for the adaptive layer to "correct" the received signal before decoding to cope with channel changes.

[0100] The transmitted signal experiences fading on the wireless channel. Therefore, at the receiving end, it is represented as: ; where represents the demodulated symbol, ; is the demodulator function; CN represents a complex Gaussian distribution, represents the covariance matrix of the distribution, is the variance, is the identity matrix.

[0101] The discrete feature representation reconstructed by the receiving - end network is input to the edge task inference module for edge inference tasks; refers to the embedding vector corresponding to the discrete codeword index mapped from the continuous latent vector through vector quantization (VQ) (such as the pre - trained vector in the codebook). It corresponds to the discrete symbol index after quantization by VQ - VAE at the transmitter end and is a compressed expression of semantic information in the low - dimensional space. As shown below:

[0102] ; Among them, is the edge task inference module with the parameter set , and is the result label output by the edge task inference module.

[0103] The entire task-oriented semantic communication is composed of a probabilistic graphical model, and the Markov chain of the probabilistic graphical model is as follows: (IV) Among them, , , , , , and are random variables, while , , , , , and are their instances respectively; in (IV), according to the corresponding codebook index is obtained, and the corresponding codebook index is obtained; after demodulation, according to the received index the discrete feature representation can be restored.

[0104] Embodiment 3 A dynamic environment adaptive semantic communication system according to Embodiment 2, wherein: This embodiment aims to verify the effectiveness and robustness of the proposed dynamic environment adaptive ToSC system (FA-NFM) based on few-shot learning in a dynamic channel environment.

[0105] Dataset and task: The public image dataset CIFAR-10 is used for simulation experiments, and the task is set as image classification.

[0106] Model architecture: Transmitter: The feature extractor uses a CNN convolutional layer; the JSC encoder is composed of several multi-layer perceptron (MLP) layers and a vector quantization (VQ) layer, the codebook size is set to N = 512, and the embedding vector dimension is D = 64.

[0107] Receiver: It includes a VQ decoding module corresponding to the transmitter, an adaptive layer based on GMM and few-shot learning, and an edge task inference module (classifier) composed of a small fully connected network.

[0108] Communication settings: Modulation method: The codebook index is transmitted using a 16-QAM digital modulation scheme.

[0109] Channel model: Simulated various channel conditions and their dynamic switching scenarios, including: Additive White Gaussian Noise (AWGN) channel, Uniform Rayleigh Fading channel, Rician Fading channel. The Signal-to-Noise Ratio (SNR) range considered in the simulation is from 2 dB to 10 dB.

[0110] Comparison method: The FA-NFM method of the present invention is compared with the following baseline methods: STWA, ATR, FTR, SNA (these methods represent different types of existing channel adaptation or transmission strategies).

[0111] Adaptive setting: After the channel condition switches, a small number of samples from the new channel (target domain) are used to update the parameters of the receiver's adaptive layer. As Figure 5 shown, even if only a small number of target domain samples (e.g., K = 40) are used for each class, the method of the present invention can effectively adapt.

[0112] Performance metrics: Mainly evaluate the quality of reconstructing semantic features under different channel conditions and SNR, with Peak Signal-to-Noise Ratio (PSNR) as the measurement standard. The higher the PSNR, the better the reconstruction quality, which is more beneficial for subsequent edge inference tasks. At the same time, the final image classification accuracy can also be further evaluated.

[0113] Result summary and analysis: Dynamic channel switching performance (refer to Figure 2 , Figure 3 , Figure 4 ): The experimental results are as shown in Figure 2 , Figure 3 and Figure 4 .

[0114] Figure 2 is the PSNR performance graph from the condition of uniform fading to Rician fading; Abscissa (X-axis): Signal-to-Noise Ratio (SNR), in decibels (dB), representing the ratio of signal strength to noise strength, usually ranging from a lower value (e.g., 2 dB) to a higher value (e.g., 10 dB). Ordinate (Y-axis): Peak Signal-to-Noise Ratio (PSNR), in decibels (dB), which is a common metric for measuring the reconstruction quality of an image (or signal), and a higher value indicates better reconstruction quality. Figure 2Shows the curves of PSNR vs. SNR for different methods when the channel condition dynamically changes from uniform fading (Rayleigh) to Rician fading. The expected result is that the FA-NFM method (the curve marked with stars) is higher than other methods at most or all SNR points, indicating its better adaptability and performance when the channel switches from one fading type to another.

[0115] Figure 3 PSNR performance graph for the transition from Rician fading to uniform fading conditions; Abscissa (X-axis): Signal-to-noise ratio (SNR) (dB). Ordinate (Y-axis): Peak signal-to-noise ratio (PSNR) (dB). Similar to Figure 2 but shows the performance when the channel condition switches back from Rician fading to uniform fading. Figure 3 Also used to verify the robustness and superiority of the FA-NFM method under dynamic channel switching. It is expected that the FA-NFM curve performs the best. The curves may cross in the low SNR region, but the FA-NFM should be dominant in the medium and high SNR regions.

[0116] Figure 4 PSNR performance graph for the transition from additive white Gaussian noise (AWGN) to uniform fading conditions; Abscissa (X-axis): Signal-to-noise ratio (SNR) (dB). Ordinate (Y-axis): Peak signal-to-noise ratio (PSNR) (dB). Figure 4 Shows the performance comparison when the channel switches from a relatively simple additive white Gaussian noise (AWGN) channel to a more complex uniform fading channel. This simulates the scenario from an ideal channel to an actual fading channel to examine the adaptability of the method. It is expected that the FA-NFM method can better adapt to this change and maintain a high PSNR.

[0117] In the simulated scenarios where the channel switches from uniform fading to Rician fading ( Figure 2 ), from Rician fading to uniform fading ( Figure 3 ), and from AWGN to uniform fading ( Figure 4 ), the FA-NFM method of the present invention (the curve marked with stars in the figure) has significantly better PSNR performance than all the compared baseline methods (STWA, ATR, FTR, SNA) in the entire SNR range from 2 dB to 10 dB. This indicates that the system of the present invention can effectively adapt to the dynamic changes of the channel and maintain high-quality feature transmission after the change.

[0118] Figure 5To change the PSNR performance graph under different adaptive sample counting conditions; Abscissa (X-axis): The number of target domain adaptation samples per class, e.g., from 20 to 100. Ordinate (Y-axis): Peak signal-to-noise ratio (PSNR) (dB), usually measured under a fixed SNR condition. Figure 5 It aims to explore the sensitivity of different methods to the number of samples in the new channel (target domain) for adaptation. It shows how the performance (PSNR) of various methods changes as the number of samples available for adaptation increases. The expected result is that the FA-NFM method (star-marked curve) can achieve a relatively high PSNR even with a small number of samples, and its performance may increase slightly or remain stable as the number of samples increases, outperforming other methods that require more samples to achieve the same performance. This highlights the advantage of "few-shot learning" of the present invention.

[0119] In summary, the simulation results of this embodiment verify that the FA-NFM system proposed by the present invention can achieve fast adaptation using a small number of samples in a dynamically changing wireless channel environment, effectively compensate for the impact of channel changes, maintain the robustness and high quality of semantic feature transmission, and has obvious advantages compared with the prior art.

[0120] Embodiment 4 A semantic communication method for adapting to a dynamic environment, comprising: Identifying and extracting features related to a specific task from the original input data; Mapping the continuous feature vectors output by the feature extractor to a discrete symbol representation suitable for wireless transmission; Mapping the continuous feature vectors output by the feature extractor to discrete codeword indices through a vector quantization mechanism; The wireless channel is used to: convert the discrete codeword indices into physically transmissible electromagnetic signals and complete the transmission process of the electromagnetic signals in a noisy and fading environment; Completing signal demodulation, semantic feature reconstruction, and task inference.

Claims

1. A dynamic environment adaptive semantic communication system for realizing the entire task-oriented semantic communication, characterized in that, Comprising: A transmitting - end network, a codebook space, a wireless channel, and a receiving - end network. The transmitting - end network includes a feature extractor and a joint source - channel encoder, i.e., a JSC encoder; The feature extractor is used to: identify and extract features related to a specific task from the original input data; The JSC encoder is used to: map the continuous feature vectors output by the feature extractor to discrete symbol representations suitable for wireless transmission; The codebook space is used to: map the continuous feature vectors output by the feature extractor to discrete codeword indices through a vector quantization mechanism; The wireless channel is used to: convert the discrete codeword indices into physically - transmittable electromagnetic signals and complete the transmission process of the electromagnetic signals in a noisy and fading environment; The receiving - end network is used to: complete signal demodulation, semantic feature reconstruction, and task inference.

2. The dynamic environment adaptive semantic communication system according to claim 1, wherein, The receiving - end network includes a vector quantization decoding module, an adaptive layer, and an edge task inference module; The vector quantization decoding module is used to: demodulate the received electromagnetic signals into discrete codeword indices and reconstruct the semantic feature vector zd(x) based on the mapping relationship of the codebook space; The adaptive layer is used to: dynamically compensate for the feature distribution offset caused by channel changes; Based on the channel distribution characteristics predicted by the Gaussian mixture model, the adaptive layer quickly optimizes the feature transformation parameters through a small number of new channel samples, so that the reconstructed semantic features are aligned with the task requirements under the current channel conditions; The edge task inference module is used to: perform terminal tasks on the reconstructed semantic features.

3. The dynamic environment adaptive semantic communication system according to claim 1, characterized in that Identify and extract features related to a specific task from the original input data; map the continuous feature vectors output by the feature extractor to discrete symbol representations suitable for wireless transmission; ; Among them, is a jointly designed feature extractor and JSC encoder, with the original input data as the input , and the parameters are ; is the set of trainable parameters including all the weights and bias terms of the feature extractor and JSC encoder; is a vector, each element of which corresponds to an encoded symbol, represents the discrete symbol representation obtained after passing through the feature extractor and JSC encoder, and each vector dimension is D.

4. A dynamic environment adaptive semantic communication system according to claim 1, characterized in that, Identify and extract features related to a specific task from the original input data; Comprising: The feature extractor is a convolutional neural network CNN; The convolutional neural network CNN includes a convolutional layer, a pooling layer, and a fully - connected layer; The conversion from the original input data to continuous semantic features is realized through the convolutional neural network CNN; Specifically, it includes: the convolutional layer extracts the spatial hierarchical features of the image through the local receptive field and weight - sharing mechanism, the pooling layer performs spatial down - sampling, and finally the fully - connected layer generates continuous semantic feature vectors of a fixed dimension.

5. The dynamic environment adaptive semantic communication system according to claim 1, characterized in that, Map the continuous feature vectors output by the feature extractor to discrete symbol representations suitable for wireless transmission; Comprising: First, perform a non - linear transformation on the continuous feature vectors through a multi - layer perceptron or a convolutional network, project the continuous feature vectors into a low - dimensional latent space to obtain continuous latent vectors; The specific implementation process is: input the continuous feature vectors into the fully - connected layer of the multi - layer perceptron or the convolutional layer of the convolutional network, and extract high - order feature correlations layer by layer through non - linear activation functions, and output low - dimensional continuous latent vectors; Secondly, adopt a codebook - based vector quantization mechanism, map the continuous latent vectors to discrete symbol indices through nearest - neighbor search, and this codebook is trained by jointly optimizing the source compression ratio and channel noise resistance; Finally, perform one - hot encoding on the discrete symbol indices to obtain discrete symbol representations.

6. The dynamic environment adaptive semantic communication system according to claim 1, wherein Adopt a codebook - based vector quantization mechanism, map the continuous latent vectors to discrete symbol indices through nearest - neighbor search; Comprising: Calculate the Euclidean distance between the current continuous latent vector and all the embedding vectors in the pre-trained codebook. The embedding vectors in the pre-trained codebook refer to a set of fixed vectors learned during the VQ-VAE training phase, and each embedding vector represents a discrete codeword. Determine the codeword index with the highest matching degree through nearest neighbor search, including: Select the index of the embedding vector with the smallest Euclidean distance, which is the obtained discrete symbol index.

7. A dynamic environment adaptive semantic communication system according to claim 1, characterized in that, Convert the discrete codeword index into an electromagnetic signal that can be physically transmitted, and complete the transmission process of the electromagnetic signal in a noisy and fading environment, including: Map the discrete codeword index to a specific constellation symbol through digital modulation. Each constellation point is represented in complex form as: s = I + jQ, where I and Q are the real and imaginary parts respectively, corresponding to the in-phase and quadrature components. Convert the discrete digital signal, i.e., the constellation point s, into an electromagnetic signal that can be physically transmitted. Adopt Gray coding to design the constellation mapping rule to ensure that adjacent constellation points have only one-bit difference. Complete signal demodulation to obtain the discrete index value, and retrieve the corresponding codeword from the locally stored codebook based on this discrete index value. Finally, input the discrete feature representation into the edge inference module. The specific implementation process includes: 1) Compensate for signal attenuation and phase offset through channel estimation. 2) Calculate the bit probability of the received symbol based on the maximum likelihood criterion or soft demodulation algorithm. 3) Recover the discrete index through hard decision or soft input decoding. 4) Extract the corresponding compressed feature vector from the codebook according to the index.

8. The dynamic environment adaptive semantic communication system according to claim 1, characterized in that Based on the channel distribution characteristics predicted by the Gaussian mixture model, the adaptive layer quickly optimizes the feature transformation parameters through a small number of new channel samples to align the reconstructed semantic features with the task requirements under the current channel conditions, including: A generative channel model is used to approximately simulate the probability density of the true channel conditions ; are the parameters of the generative channel model. The conditional density of the channel is modeled by a set of Gaussian mixture models as follows: ; where k is the number of components, is the mean vector, is the covariance matrix, and [0, 1] is the prior probability of component i, and the prior probabilities of the components are represented by the softmax function; refers to the conditional probability density function, which represents the probability density that the signal / feature actually observed at the receiving end is x under the condition that the discrete symbol / index sent by the transmitting end is z; is the parameter set of the generative channel model; x represents the observed signal or feature vector at the receiving end; z represents the discrete symbol or codebook index sent by the transmitting end; Adjust the mean, covariance, and mixing weights of each Gaussian component in the original GMM model through a set of specific transformation parameters to match the statistical characteristics of the target channel. These transformation parameters include: Affine transformation parameters for the mean For the mean vector of each Gaussian component, the affine transformation consists of a linear transformation matrix and an offset vector . The transformed mean (z) is calculated by the following formula: (z) = (z) + ; wherein, refers to a matrix for linearly transforming the mean value of the i-th Gaussian component; refers to an offset vector or translation vector for the mean value of the i-th Gaussian component; Affine transformation parameters for the covariance For the covariance matrix of each Gaussian component, the affine transformation is implemented by a transformation matrix and the transformed covariance is calculated by the following formula: = ; Among them, refers to the covariance matrix of the i-th component in the original GMM model given z; refers to the new covariance matrix of the i-th component in the GMM model adapted to the new channel given z after transformation; refers to the matrix for linearly transforming the covariance of the i-th Gaussian component; Affine transformation parameters for the mixing weights The mixing weight or prior probability of each Gaussian component is transformed by a scaling factor and an offset to obtain an adjusted weight; the transformed prior logic (z) is calculated by the following formula: (z) = (z) + ; $(z)$ refers to the logit value corresponding to the mixing weight of the $i$-th component in the original GMM model; refers to the scaling factor for the logit value of the $i$-th component; refers to the offset for the logit value of the $i$-th component.

9. A dynamic environment adaptive semantic communication system according to any one of claims 1-8, characterized in that, The transmitted signal experiences fading on the wireless channel. Therefore, at the receiving end, it is represented as: ; Among them, represents the demodulated symbol, ; is the demodulator function; CN represents a complex Gaussian distribution, represents the covariance matrix of the distribution, is the variance, is the identity matrix; Further preferably, the discrete feature representation reconstructed by the receiving-end network is input into the edge task inference module to perform edge inference tasks, as follows: ; Among them, is the edge task inference module with the parameter set , and is the result label output by the edge task inference module.

10. A dynamic environment adaptive semantic communication method, characterized in that, Including: Identify and extract features related to a specific task from the original input data. Map the continuous feature vector output by the feature extractor to a discrete symbol representation suitable for wireless transmission. Map the continuous feature vector output by the feature extractor to a discrete codeword index through a vector quantization mechanism. The wireless channel is used to convert the discrete codeword index into an electromagnetic signal that can be physically transmitted and complete the transmission process of the electromagnetic signal in a noisy and fading environment. Complete signal demodulation, semantic feature reconstruction, and task inference.

Citation Information

Cited By

  • Robust semantic communication method, device and equipment based on sparse vector coding

    CN120768503A

  • Image adaptive quantization communication method and system

    CN121442092A

  • Non-contact electromagnetic heart monitoring method based on signal semantic deconstruction

    CN121795867A