MIMO semantic communication method and system capable of resisting channel mismatch

By embedding a channel matrix adapter in the MIMO semantic communication system and adopting a two-stage training strategy and knowledge distillation mechanism, the channel mismatch problem was solved, the robustness and transmission quality of the system under non-ideal channel conditions were improved, and high-fidelity data reconstruction was achieved.

CN121585503APending Publication Date: 2026-02-27SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511899017.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing MIMO semantic communication systems suffer from severe channel mismatch problems in real-world scenarios where channel estimation errors are unavoidable, leading to a sharp decline in system performance and making it difficult to maintain robust transmission performance.

Method used

By embedding a channel matrix adapter in the channel codec, a three-layer channel processing architecture is constructed to dynamically sense and compensate for channel estimation errors. A two-stage training strategy and a knowledge distillation mechanism are used to optimize the parameters of the neural network modules, thereby achieving fine-tuning of semantic features.

Benefits of technology

It effectively improves the robustness and transmission quality of the system under non-ideal channel conditions, and ensures high-fidelity data reconstruction in large-scale antenna array scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121585503A_ABST
    Figure CN121585503A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless communication, and discloses an MIMO semantic communication method and system capable of resisting channel mismatch. The method comprises the following steps: acquiring to-be-transmitted data and an estimated channel matrix between a transmitting end and a receiving end, and performing semantic coding on the to-be-transmitted data to obtain a coding semantic feature; based on the estimation channel matrix, executing three-layer adaptive channel coding on the coding semantic feature to obtain a sending coding feature; pre-coding the sending coding characteristics, transmitting pre-coded signals to a receiving end to obtain receiving signals, and balancing the receiving signals based on the estimation channel matrix to obtain balanced characteristics; and based on the estimated channel matrix, performing three-layer adaptive channel decoding on the equilibrium feature to obtain a decoded semantic feature, and performing semantic decoding on the decoded semantic feature to reconstruct data. According to the invention, the dynamic compensation of the estimation channel matrix error is realized, and the robustness and transmission quality of the MIMO semantic communication system in an actual non-ideal channel environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to a MIMO semantic communication method and system that resists channel mismatch. Background Technology

[0002] With the rapid popularization of the Internet of Things and smart devices, MIMO semantic communication technology achieves efficient semantic feature transmission by jointly optimizing source and channel coding through deep neural networks. Existing MIMO semantic communication systems typically assume that the receiver can obtain ideal channel state information and design adaptive semantic codecs based on accurate channel matrices. They utilize singular value decomposition to decompose the MIMO channel into multiple equivalent sub-channels, thereby achieving high-fidelity image transmission.

[0003] However, in practical MIMO systems, especially in large-scale antenna array scenarios, obtaining an accurate and complete channel matrix is ​​costly and impractical due to limitations such as pilot overhead, hardware precision, and feedback delay. The actual channel matrix obtained often contains significant estimation errors. When a semantic communication system uses an ideal channel matrix during training but encounters an erroneous estimated channel matrix during inference, a severe channel mismatch problem arises, leading to a sharp decline in system performance. Experiments show that this mismatch can cause an average PSNR loss exceeding 1.8 dB and significantly degrade the quality of reconstructed images. Even simple fine-tuning of the model based on a non-ideal channel matrix is ​​insufficient to effectively compensate for the semantic information distortion caused by channel mismatch. Therefore, how to maintain robust transmission performance in MIMO semantic communication systems in real-world scenarios where channel estimation errors are unavoidable has become a pressing technical problem. Summary of the Invention

[0004] The main objective of this invention is to address the problem that existing MIMO semantic communication methods fail to effectively address channel estimation errors, leading to channel mismatch during the training and inference phases and a significant decrease in system transmission performance.

[0005] The first aspect of this invention provides a MIMO semantic communication method resistant to channel mismatch, the MIMO semantic communication method resistant to channel mismatch comprising: The data to be transmitted and the estimated channel matrix between the sender and receiver are obtained, and the data to be transmitted is semantically encoded to obtain the encoded semantic features. Based on the estimated channel matrix, the first channel sub-coding, the transmitter channel adaptation process, and the second channel sub-coding are sequentially performed on the encoded semantic features to obtain the transmission coding features. The transmitted coding features are precoded, and the precoded signal is transmitted to the receiving end through a MIMO channel to obtain the received signal. Based on the estimated channel matrix, the received signal is equalized to obtain equalization features. Based on the estimated channel matrix, the equalization features are sequentially subjected to first channel sub-decoding, receiver channel adaptation processing, and second channel sub-decoding to obtain decoded semantic features. Then, the decoded semantic features are semantically decoded to reconstruct the data.

[0006] A second aspect of the present invention provides a MIMO semantic communication system resistant to channel mismatch, the MIMO semantic communication system resistant to channel mismatch comprising: The semantic coding module is used to acquire the data to be transmitted and the estimated channel matrix between the sender and receiver, and to perform semantic coding on the data to be transmitted to obtain coded semantic features. The channel coding module is used to perform first channel sub-coding, transmitter channel adaptation processing and second channel sub-coding on the coding semantic features based on the estimated channel matrix to obtain the transmission coding features. The transmission processing module is used to precode the transmission coding features, transmit the precoded signal to the receiving end through the MIMO channel to obtain the received signal, and equalize the received signal based on the estimated channel matrix to obtain equalization features. The receiving decoding module is used to perform first channel sub-decoding, receiver channel adaptation processing and second channel sub-decoding on the equalization features in sequence based on the estimated channel matrix to obtain decoded semantic features, and to perform semantic decoding on the decoded semantic features to reconstruct the data.

[0007] The above-described MIMO semantic communication method and system for resisting channel mismatch is described above. In this embodiment of the invention, by acquiring the data to be transmitted and the estimated channel matrix between the transmitting end and the receiving end, the data to be transmitted is semantically encoded to obtain encoded semantic features; based on the estimated channel matrix, the encoded semantic features are sequentially subjected to first channel sub-coding, transmitting end channel adaptation processing, and second channel sub-coding to obtain transmitting encoded features; the transmitting encoded features are pre-coded, and the pre-coded signal is transmitted to the receiving end through the MIMO channel to obtain the received signal, and the received signal is equalized based on the estimated channel matrix to obtain equalization features; based on the estimated channel matrix, the equalization features are sequentially subjected to first channel sub-decoding, receiving end channel adaptation processing, and second channel sub-decoding to obtain decoded semantic features, and the decoded semantic features are semantically decoded to reconstruct the data. This application constructs a three-layer channel processing architecture by embedding a channel matrix adapter in the channel codec. It dynamically senses and compensates for channel estimation errors in the intermediate stage of feature compression, thereby achieving fine-grained adjustment of semantic features. This solves the problem of sharp performance degradation caused by channel mismatch in the actual deployment of existing MIMO semantic communication systems, effectively improves the robustness and transmission quality of the system under non-ideal channel state information conditions, and ensures high-fidelity data reconstruction in large-scale antenna array scenarios.

[0008] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0009] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the first embodiment of the MIMO semantic communication method for resisting channel mismatch in this invention. Figure 2 This is an example diagram of a MIMO semantic communication system model that resists channel mismatch in an embodiment of the present invention; Figure 3 This is a schematic diagram of an embodiment of the MIMO semantic communication system that resists channel mismatch in this invention. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0013] To facilitate understanding of this embodiment, the specific process of this embodiment is described below. Please refer to [link / reference]. Figure 1 The first embodiment of the MIMO semantic communication method for resisting channel mismatch in this invention includes: 101. Obtain the data to be transmitted and the estimated channel matrix between the sending end and the receiving end, and perform semantic encoding on the data to be transmitted to obtain encoded semantic features; In this embodiment, before acquiring the data to be transmitted and the estimated channel matrix between the transmitter and receiver, the corresponding parameters of the MIMO semantic communication system are adjusted through a two-stage training strategy: First, a training dataset of multiple training samples is acquired, wherein the training samples include the training data to be transmitted, the corresponding non-ideal training channel matrix with estimation error, and the corresponding ideal training channel matrix. The ideal training channel matrix is ​​then used to pre-train the semantic encoder, the first channel sub-encoder, the transmitter channel adapter, the second channel sub-encoder, the first channel sub-decoder, the receiver channel adapter, the second channel sub-decoder, and the semantic decoder to obtain pre-training parameters. In the first training stage, the pre-training parameters are used to initialize the semantic encoder, the first channel sub-encoder, the transmitter channel adapter, the second channel sub-encoder, the first channel sub-decoder, the receiver channel adapter, the second channel sub-decoder, and the semantic decoder. The parameters of the semantic encoder and the semantic decoder are frozen. The non-ideal channel matrix is ​​then used to pre-train the first channel sub-encoder, the transmitter channel adapter, the second channel sub-encoder, the first channel sub-decoder, the receiver channel adapter, and the second channel sub-decoder. The channel sub-decoder undergoes end-to-end training to obtain the first-stage training parameters. In the second training stage, the pre-trained parameters are used as teacher model parameters and frozen, while the first-stage training parameters are used as student model parameters. Knowledge distillation training is performed using the training dataset. For each training sample, teacher coding features and teacher decoding features are obtained using the teacher model parameters and an ideal channel matrix, respectively. Student coding features and student decoding features are obtained using the student model parameters and a non-ideal channel matrix. A comprehensive loss function, including data reconstruction loss and distillation loss, is calculated. The distillation loss includes the KL divergence between student coding features and teacher coding features, as well as the KL divergence between student decoding features and teacher decoding features. The student model parameters are updated according to the comprehensive loss function. After training, the student model parameters are used as the final parameters for the semantic encoder, the first channel sub-encoder, the transmitting channel adapter, the second channel sub-encoder, the first channel sub-decoder, the receiving channel adapter, the second channel sub-decoder, and the semantic decoder. The channel mismatch-resistant MIMO semantic communication method is applied to MIMO semantic communication system models (such as...). Figure 2As shown, the MIMO semantic communication system model includes a semantic encoder, a first channel sub-encoder, a transmitting channel adapter, a second channel sub-encoder, a first channel sub-decoder, a receiving channel adapter, a second channel sub-decoder, and a semantic decoder. Channel signal-to-noise ratio (SNR) parameters are obtained, and multiple scale adjustment factors are generated based on these parameters. Multi-scale feature extraction is performed on the data to be transmitted to obtain multi-level feature maps. Affine transformations are then performed on the multi-level feature maps according to the scale adjustment factors to obtain adjusted multi-level feature maps. Feature fusion and dimensionality transformation are then performed on the adjusted multi-level feature maps to obtain encoded semantic features.

[0014] In practical applications, such as Figure 2 As shown, the MIMO semantic communication system model includes an SNR-adaptive semantic encoder (SemEnc-SA), a channel encoder (ChEnc), a pre-encoder (Prec), a post-processor (PostP), a channel decoder (ChDec), and an SNR-adaptive semantic decoder (SemDec-SA). Further, the channel encoder internally comprises a three-layer structure: a first channel sub-encoder, a transmitting-end channel adapter, and a second channel sub-encoder; similarly, the channel decoder internally comprises a three-layer structure: a first channel sub-decoder, a receiving-end channel adapter, and a second channel sub-decoder. Before the system is put into practical use, the parameters of the above eight neural network modules need to be optimized and adjusted using a two-stage training strategy. The preparation process of the training dataset is as follows: an image dataset containing multiple training samples is selected, such as the CIFAR10 dataset containing 50,000 RGB color images with a resolution of 32×32 pixels, as the training data to be transmitted. Each training sample is equipped with two types of channel matrices: an ideal training channel matrix. Non-ideal training channel matrix Ideal training channel matrix This represents perfect channel state information where there is no estimation error between the transmitter and receiver, and its dimension is... Line × A complex matrix of columns, where For the number of transmitting antennas, This represents the number of receiving antennas. The non-ideal training channel matrix. The generation method is as follows: in the ideal training channel matrix A random noise value is superimposed at each element position of the matrix, i.e. Here, addition refers to the element-wise addition of matrices, estimating the error matrix. Each element is independently sampled from a Gaussian distribution in the complex field, which has a mean of 0 and a variance of . A Gaussian distribution in the complex domain refers to a random variable whose real and imaginary parts each follow an independent real Gaussian distribution, reflecting the error characteristics of simultaneous amplitude and phase in practical channel estimation. Variance This determines the severity of the estimation error. To simulate scenarios with different channel estimation accuracies, each training sample... From the minimum value min to maximum value The data is uniformly and randomly sampled between 0 and 10, allowing the system to learn and adapt to various estimation error levels, from minor to severe. This data preparation method realistically reproduces the channel state information estimation deviations caused by limited pilot sequence length, insufficient receiver hardware quantization accuracy, and channel feedback delays in actual communication systems, providing a data foundation consistent with real-world application scenarios for subsequent training.

[0015] The pre-training phase establishes initial parameter configurations for the semantic encoder, the first channel sub-encoder, the transmitter channel adapter, the second channel sub-encoder, the first channel sub-decoder, the receiver channel adapter, the second channel sub-decoder, and the semantic decoder. This phase uses an ideal training channel matrix. The training process is as follows: A batch of image samples is extracted from the training dataset. Each image x is input into a semantic encoder to extract multi-scale semantic features, and the encoded semantic features are output. These encoded semantic features are then compressed into intermediate features by the first channel sub-encoder. The transmitting channel adapter then uses the ideal channel matrix... The intermediate features are processed (at this point, since the channel matrix is ​​error-free, the adapter mainly learns the structured adjustment of the features), and the second channel sub-encoder further compresses them into transmit coded features. The transmit coded features are then subjected to singular value decomposition and pre-coding by the pre-encoder, simulating transmission through the MIMO channel. At the receiver, the post-processor performs channel equalization to obtain equalization features. The equalization features are then sequentially expanded in dimension by the first channel sub-decoder, and the receiver channel adapter, based on the ideal channel matrix, performs the final equalization. The second channel sub-decoder further expands and recovers the decoded semantic features; finally, the semantic decoder reconstructs the output image. During training, the original image x and the reconstructed image x are compared. The sum of absolute errors at the pixel level is used as the loss function. A smaller loss function value indicates higher reconstruction quality. The gradient descent optimization method (specifically the Adam optimizer, which adaptively adjusts the learning rate of each parameter) is used to calculate the gradient of the loss function with respect to the parameters of each module, and backpropagation is used to update the parameters of all eight modules. This process is repeated for multiple training rounds until the decrease in the loss function is less than a preset threshold for several consecutive rounds (e.g., a decrease of less than 0.001 for 10 consecutive rounds). At this point, training is considered converged, and the current parameters are saved as pre-training parameters. The pre-training phase ensures that the system has accurate semantic feature extraction, compressed transmission, and high-fidelity reconstruction capabilities under ideal channel conditions. The learned parameters provide a good initialization foundation for subsequent adaptation to non-ideal channel environments, avoiding the slow convergence and suboptimal performance problems caused by training with random parameters.

[0016] The first training phase is for adaptation training in a non-ideal channel environment. This phase begins by loading pre-trained parameters into the semantic encoder, first channel sub-encoder, transmitter channel adapter, second channel sub-encoder, first channel sub-decoder, receiver channel adapter, second channel sub-decoder, and semantic decoder as initial parameters. Subsequently, the parameters of the semantic encoder and semantic decoder are frozen. Specifically, in a deep learning training framework (such as PyTorch), a parameter freeze interface is called to disable gradient calculation for all weight matrices and bias vectors in these two modules. This ensures that these parameters are not updated during subsequent backpropagation and remain at their pre-training values. The reason for freezing the semantic encoder and decoder is that they are primarily responsible for extracting high-level semantic information (such as object contours and texture patterns) and reconstructing image details. These capabilities are independent of channel characteristics and have already been sufficiently learned during pre-training. The goal of the first training phase is to enable the three-layer structure of the channel encoder and decoder to adapt to the transmission mismatch caused by the non-ideal channel matrix. Training all modules simultaneously could easily lead to the semantic layer parameters being degraded by noise interference from the non-ideal channel. During training, each training sample uses the corresponding non-ideal training channel matrix. Forward propagation calculations are performed. Specifically, at the transmitting end channel adapter, the intermediate features output by the first channel sub-encoder and the estimated channel matrix are combined. The common input adapter module, internally uses a Transformer structure to... The flattened projection is used as channel state words, which are concatenated with intermediate features and then subjected to feature interaction through a multi-head self-attention mechanism to learn and recognize them. The system estimates the error pattern and compensates for intermediate features, outputting adapted features; the receiver channel adapter processes the equalization features in the same way. The loss function still uses the L1 norm to calculate the pixel difference between the reconstructed image and the original image, but during backpropagation, only the parameters of six modules—the first channel sub-encoder, the transmitting channel adapter, the second channel sub-encoder, the first channel sub-decoder, the receiving channel adapter, and the second channel sub-decoder—are updated. The first-stage training parameters are obtained after training convergence. This stage specifically optimizes the three-layer structure of the channel encoder / decoder for estimation errors in non-ideal channel matrices. The transmitting and receiving channel adapters learn to identify channel mismatch patterns and dynamically compensate for semantic features, thus enabling accurate transmission of semantic information even in channel environments with estimation errors, thereby improving the system's robustness.

[0017] The second training phase introduces a knowledge distillation mechanism to further optimize system performance. Knowledge distillation is a training strategy that improves the performance of weaker student models by having them learn the output feature distribution of stronger teacher models. In this phase, the pre-trained parameter-configured system is used as the teacher model, and the system with the first-phase training parameter configuration is used as the student model. The teacher and student models have identical network structures (both containing the eight modules mentioned above), but all parameters of the teacher model are frozen and no longer updated, and the ideal training channel matrix is ​​always used. The calculations are performed; all parameters of the student model can be updated, and a non-ideal training channel matrix is ​​used. The computation is performed. For each training sample in the training dataset, after the teacher model processes the sample, it extracts the teacher-encoded features at the output position of the second channel sub-encoder. (This feature contains the semantic representation after three layers of channel coding under ideal channel conditions), and the teacher decoding feature is extracted at the output position of the second channel sub-decoder. (This feature includes the semantic representation recovered after three-layer channel decoding under ideal channel conditions); After processing the same sample, the student model extracts the student-coded features at the same location. Decoding features of students The comprehensive loss function is defined as follows: ,in Reconstructing images for student models compared to the original images loss, The KL divergence is represented by β, where β is a weight hyperparameter (typically ranging from 0.1 to 1.0) used to balance the contributions of reconstruction loss and distillation loss. The KL divergence is a measure in information theory of the difference between two probability distributions; a smaller value indicates that the two distributions are closer. In this scheme, the KL divergence is calculated by: [The text abruptly ends here, so the translation stops.] and teacher coding features Treating them as high-dimensional feature vectors, we apply the softmax normalization function (the softmax function maps any real vector to a probability distribution that sums to 1) to each, obtaining two probability distributions. Then, we calculate the KL divergence value of the student distribution relative to the teacher distribution; the KL divergence of the decoded features is calculated in the same way. During training, gradient calculations and backpropagation updates are performed on the parameters of all eight modules of the student model according to the comprehensive loss function. After training converges, the student model parameters are used as the final parameters of the system. The technical effect of introducing knowledge distillation is to solve the following problem: In actual communication, an estimated channel matrix... The estimation error may correspond to multiple different real channel matrices (due to the randomness of the estimation error). This "one-to-many" correspondence makes it difficult for the neural network to establish a stable input-output mapping pattern during the first training stage, leading to oscillations and non-convergence during training. Through knowledge distillation, the teacher model is trained based on the ideal channel matrix, and its encoded and decoded features represent the optimal semantic representation pattern under the condition of no estimation error. The student model is guided by KL divergence constraints to imitate the feature distribution of the teacher model. This is equivalent to adapting to the non-ideal channel matrix while learning from the knowledge under the ideal channel condition, effectively stabilizing the training process and improving the system's generalization ability under various estimation error levels. Experimental data show that the system trained by the second stage of knowledge distillation can improve the image reconstruction quality by approximately 0.12 dB compared to the system trained only by the first stage, under the same channel estimation error variance.

[0018] Specifically, before semantic encoding is performed at the sending end, the data to be transmitted (such as a 32×32 pixel RGB color image x) and the estimated channel matrix are first obtained. Simultaneously, the channel signal-to-noise ratio (SNR) parameter *r* is obtained. The SNR represents the ratio of useful signal power to noise power, measured in decibels (dB). A higher SNR value indicates better channel quality. The SNR parameter *r* is obtained as follows: When a connection is established in the communication system, the receiver analyzes the received pilot signal (a known reference signal sequence) to measure the current channel quality and calculate the SNR value. This value is then fed back to the transmitter via the control channel. The transmitter reads the latest SNR feedback value as the parameter *r* before each encoded transmission. The semantic encoder integrates an adaptive SNR module, which uses a multi-layer perceptron (MLP) network to convert the SNR parameter *r* into multiple scale adjustment factors. The MLP network consists of several fully connected neural network layers. The input is a single numerical value *r*, which is transformed layer by layer by a nonlinear activation function (such as the ReLU function), and the output is a set of numerical vectors. The length of the output vector is equal to twice the number of feature layers generated by subsequent multi-scale feature extraction, because each feature map layer requires a pair of scale adjustment factors: a scaling factor γ and an offset factor β. For example, if three feature maps are subsequently extracted, the MLP outputs six values, which are: These scaling factors are dynamically adjusted based on channel quality: under low SNR conditions (e.g., r=0 dB), the network learns to generate larger scaling factors to enhance the amplitude of key features and improve noise immunity; under high SNR conditions (e.g., r=15 dB), it generates scaling factors close to 1 and offset factors close to 0, preserving more original feature details to fully utilize the good channel transmission quality. This adaptive mechanism enables the semantic encoder to optimize coding strategies for different channel environments.

[0019] The semantic encoder employs a Swin-Transformer backbone network for multi-scale feature extraction of the data to be transmitted. The Swin-Transformer is a hierarchical deep neural network that extracts feature representations at different spatial scales and semantic levels from the input image through layer-by-layer downsampling and feature abstraction. Specifically, the input image x (3×32×32, representing 3 color channels, 32 pixels high, and 32 pixels wide) first passes through image segmentation and embedding layers, dividing it into multiple small regions and mapping them to initial feature vectors. Then, it sequentially passes through multiple Swin-Transformer blocks for feature extraction. After passing through several blocks, spatial downsampling is performed (resolution halved, feature channel number doubled), outputting multi-level feature maps at different network depths. For example, the first layer outputs a shallow feature map. (With dimensions of 128×16×16, representing 128 feature channels and a spatial size of 16×16), this feature map contains low-level visual semantics such as image edges and textures; the second layer outputs a mid-level feature map. (Dimensions are 256×8×8), including intermediate semantics such as local structure of objects; the third layer outputs a deep feature map. (Dimensions are 512×4×4), including high-level semantics such as object category and overall layout. After obtaining multi-level feature maps, an affine transformation is performed on each feature map layer according to the scale adjustment factors generated in the previous step. The specific operation of the affine transformation is as follows: for the i-th layer feature map... For each feature channel, multiply the feature values ​​of all spatial locations within that channel element-wise by the corresponding scaling factor. Then add the offset factor to each element. The adjusted feature map is obtained. Adjust the formula to For example, for the 128×16×16 feature map of the first layer, a scaling factor is used. and offset factor Perform a channel-level affine transformation; use the second layer... and And so on. Different levels of feature maps require different adjustment factors because shallow features mainly contain detailed information and are sensitive to noise, while deep features mainly contain semantic information and are robust to noise. The protection and adjustment strategies for these features should differ under different channel qualities. The multi-level feature maps after affine transformation adaptively adjust the numerical distribution, enhancing robustness to channel noise while preserving task-related semantic information.

[0020] The adjusted multi-level feature maps are then fused and dimensionally transformed to obtain encoded semantic features. The purpose of feature fusion is to integrate semantic information from different levels of abstraction. The implementation method is as follows: First, the deep feature maps are upsampled (using bilinear interpolation) to restore them to the same spatial resolution as the shallow feature maps. Then, they are concatenated along the channel dimension. Specifically, the 512×4×4 feature map of the third layer is upsampled to 512×16×16, and the 256×8×8 feature map of the second layer is upsampled to 256×16×16. These are then concatenated with the 128×16×16 feature map of the first layer along the channel dimension to obtain a fused feature map with dimensions (128+256+512)×16×16=896×16×16. The fused feature map is then dimensionally transformed through several convolutional layers to adjust the number of feature channels and spatial size to a predetermined dimension c'×h'×w', outputting the encoded semantic features. Where x represents the input image and r represents the channel signal-to-noise ratio parameter. This represents a semantic encoding function. This represents the learnable parameters, such as weights and biases, of all neural network layers within the semantic encoder. It encodes semantic features. It is a three-dimensional tensor that contains the core semantic content of the input image. Its data volume is significantly compressed compared to the original image (e.g., from 3×32×32=3072 values ​​to c'×h'×w' values). Furthermore, it undergoes adaptive adjustment based on channel quality, focusing on preserving key semantic features at low SNR and retaining richer detailed features at high SNR. This provides a high-quality semantic representation that is adapted to the current channel environment for subsequent channel coding and cross-MIMO channel transmission, ensuring efficient and reliable semantic communication under different channel conditions.

[0021] 102. Based on the estimated channel matrix, the first channel sub-coding, the transmitter channel adaptation process, and the second channel sub-coding are sequentially performed on the encoded semantic features to obtain the transmission coding features; In this embodiment, the encoded semantic features are subjected to a first-dimensional compression transformation to obtain a first encoded feature; based on the estimated channel matrix, the first encoded feature is adapted to the transmitter channel to obtain a first adapted feature (wherein, adapting the first encoded feature to the transmitter channel based on the estimated channel matrix to obtain the first adapted feature includes: flattening the estimated channel matrix in terms of dimensions, and projecting the flattened estimated channel matrix into transmitter channel state terms; adding the transmitter channel state terms to the beginning of the sequence of the first encoded feature to obtain an extended feature sequence; performing multi-layer feature interaction on the extended feature sequence through a multi-head self-attention mechanism to obtain an enhanced feature sequence, removing channel state terms from the enhanced feature sequence, and extracting feature sequences corresponding to the positions of the first encoded features to generate the first adapted feature); the first adapted feature is subjected to a second-dimensional compression transformation to obtain a second encoded feature whose characterization feature dimension matches the number of transmitter antennas, and the second encoded feature is power normalized to obtain a transmitter encoded feature.

[0022] In practical applications, encoding semantic features After passing through the semantic encoder, the input 3D tensor enters the first channel sub-encoder for the first dimension compression transformation. The first channel sub-encoder uses a fully connected layer network, and the specific implementation process is as follows: the input 3D tensor... (Dimensions are c'×h'×w') First, the vector is flattened into a one-dimensional vector in the order of channel-height-width. For example, a 512×4×4 tensor is flattened into a vector of length 8192. Then, this vector is input into a fully connected neural network with 2 to 3 layers. Each layer performs matrix multiplication with the input vector using the weight matrix, adds the bias vector, and then passes it through the ReLU activation function to gradually reduce the feature dimension. The final output dimension is... The first encoded feature in two-dimensional matrix form ,in d' represents the number of transmit antennas (e.g., 16), and d' represents the dimension of the intermediate compressed feature (e.g., 128). First coded feature The matrix form is represented as follows: each position in the first dimension corresponds to a transmitting antenna, and the second dimension is the feature vector corresponding to that antenna. For example, The first row represents the 128-dimensional feature vector corresponding to the first antenna, the second row represents the feature vector of the second antenna, and so on. This can be expressed by the formula: ,in Represents the first channel sub-coding function. This represents the learnable parameters of the encoder (including the weight matrix and bias vector of each fully connected layer). The first-dimensional compression transformation compresses the amount of semantic feature data from 8192 values ​​to 2048 values ​​(based on the example above), significantly reducing transmission overhead. At the same time, it organizes the features into a matrix structure according to the number of antennas, so that each antenna corresponds to a clear feature vector, which facilitates independent processing and optimization of the transmission content of each antenna in the MIMO system.

[0023] First coding feature After obtaining it, based on the estimated channel matrix Perform channel adaptation processing at the transmitting end. This step first involves estimating the channel matrix. (dimension is) Line × The complex matrix (column-wise) is flattened by arranging its elements into a one-dimensional vector in row-wise order. For example, arranging the 16 elements in the first row and the 16 elements in the second row of a 16×16 matrix sequentially yields a complex vector of length 256. This flattened vector is then mapped through a fully connected layer to a transmitter channel state term (CSO) of dimension d' (the same as the feature dimension of the first coded feature, e.g., 128). The CSO is a feature vector containing a compressed representation of the estimated channel matrix after a linear transformation, carrying information about the current channel's amplitude fading, phase shift, and other characteristics. After obtaining the CSO, it is inserted as a new row vector into the first coded feature. The 0th row of the matrix (where the original 1st row becomes the 1st row, the original 2nd row becomes the 2nd row, and so on) forms an extended feature sequence with dimension . It contains one channel state term and There are 10 antenna feature vectors. The expanded feature sequence is input into a Transformer structure for feature interaction. The Transformer structure contains N Transformer blocks (4 layers are optimal), each block containing a multi-head self-attention layer and a feedforward network layer. The multi-head self-attention layer works by: processing each row of the expanded feature sequence (total...)... Each row is linearly transformed into a query vector, key vector, and value vector using three different weight matrices. Then, the dot product of each query vector and all key vectors is calculated (the sum of corresponding element-wise multiplications of two vectors). The dot product result is normalized using a softmax function to obtain the attention weight, which indicates which other rows in the sequence the row's features should focus on. Finally, the attention weights are used to perform a weighted average of all value vectors to obtain the updated features for that row. Multi-head processing refers to using eight sets (e.g.,) of different weight matrices in parallel for the above calculations, and then concatenating the eight sets of results to enhance the expressive power of the features. Through this mechanism, channel information in the channel state term can be propagated to the feature vector of each antenna, and the features of each antenna can also perceive the feature content of other antennas and channel characteristics, achieving global information interaction. The extended feature sequence is processed by four Transformer blocks to obtain the enhanced feature sequence, with the dimension still being [dimensionality missing]. Remove line 0 (channel state term) from the enhanced feature sequence, retaining the subsequent lines. The matrix formed by these rows is extracted as the first fitting feature. , dimension Expressed as a formula ,in This represents the channel matrix adapter function. This represents the learnable parameters of the adapter (including the channel term projection layer, the weight matrix in the Transformer block, etc.). The technical advantage of transmitter channel adaptation lies in: estimating the channel matrix. The channel estimation error included in the method can cause the actual channel characteristics to deviate from expectations during transmission. Traditional methods directly use... Precoding can lead to severe mismatches. This step uses the global information exchange mechanism of the Transformer to adaptively adjust the characteristics of each antenna according to the channel estimation. This is equivalent to pre-compensating for the possible channel mismatch at the feature level, which significantly improves the transmission robustness of the system under non-ideal channel conditions.

[0024] First adaptation feature The input then undergoes a second dimensionality compression transformation in the second channel sub-encoder. The second channel sub-encoder employs a fully connected layer network to process the input... The feature matrix is ​​further compressed by passing each row of the matrix (the feature vector of each antenna) through a fully connected layer containing one or two layers, compressing the feature dimension from d' to d, resulting in... Second coding feature of dimension , where d represents the number of complex symbols transmitted per antenna (e.g., 64). This can be expressed by the formula: ,in This represents the second channel sub-coding function. This represents the learnable parameters of the encoder. Second coding feature. Upon acquisition, the power is immediately normalized. Power normalization is a non-trainable mathematical operation, specifically implemented by calculating... The sum of the squares of the moduli of all elements (because) The total power is obtained from a complex matrix. Then Divide each element by ,in This is the total allowable transmit power constraint value for the system (determined by the rated power of the power amplifier components in the transmitter hardware). For example, if the calculated... Watts, and power constraints If the value is 1, then each element is divided by 1. This process ensures that the normalized total power is exactly 10 watts. The power-normalized features are used as transmission coding features, ready to enter the pre-encoder. The second dimensional compression and power normalization transform the channel-adapted optimized features into a signal format that meets the transmission requirements of the MIMO physical layer: the feature dimension d matches the number of symbols that each antenna can transmit in one transmission slot, satisfying spectrum resource allocation; power normalization ensures that the total power of the transmitted signal does not exceed the hardware limitations of the transmitter, avoiding signal distortion caused by the power amplifier entering the nonlinear saturation region, or insufficient signal-to-noise ratio at the receiver due to excessively low transmit power, thus ensuring stable and reliable signal transmission in the MIMO channel.

[0025] 103. The transmitted coding features are precoded, and the precoded signal is transmitted to the receiving end through the MIMO channel to obtain the received signal. Based on the estimated channel matrix, the received signal is equalized to obtain the equalization features. In this embodiment, singular value decomposition is performed on the estimated channel matrix to obtain a right singular matrix; matrix multiplication is performed between the right singular matrix and the transmitted coding features to obtain a precoded signal; the precoded signal is transmitted to the MIMO channel through a preset transmit antenna array, and the receiver obtains the received signal through a preset receive antenna array; singular value decomposition is performed on the estimated channel matrix to obtain a left singular matrix, and the conjugate transpose of the left singular matrix is ​​calculated; matrix multiplication is performed between the conjugate transpose and the received signal to obtain equalization features.

[0026] In practical applications, after obtaining the transmission coding features, precoding is performed. The precoding process is as follows: first, the estimated channel matrix is ​​processed... (dimension is) Line × A complex matrix of columns, where For the number of transmitting antennas, Singular Value Decomposition (SVD) is performed on the number of receiving antennas. SVD is a matrix factorization method that decomposes any complex matrix into the product of three matrices, mathematically represented as: Where U is the left singular matrix (dimension 1) ), where Σ is the singular value matrix (dimension). (Diagonal elements are non-negative real singular values, and the rest are zero), V is a right singular matrix (dimension...). The superscript H indicates the conjugate transpose operation (after matrix transpose, each element takes its complex conjugate, i.e., the real part remains unchanged while the imaginary part changes sign). SVD decomposition is implemented by calling a dedicated algorithm from a standard numerical computation library (such as the LAPACK library), which takes the channel matrix as input. Output three decomposition matrices U, Σ, and V. After obtaining the right singular matrix V, it is compared with the transmitted encoded features. Perform matrix multiplication to obtain the precoded signal. ,in The dimension is ×d (where d is the number of symbols per antenna), the specific operation of matrix multiplication is that each row of V is multiplied by... The corresponding columns are multiplied element-wise and then summed. The dimension of the result z is... ×d. The technical advantage of precoding is that during transmission, the signals from multiple transmitting antennas in a MIMO channel are mixed and superimposed, making it difficult for the receiver to separate the signals from each antenna. By using SVD decomposition to find the characteristic structure of the channel matrix, the right singular matrix V contains the optimal signal transmission direction. After linearly transforming the transmission coding features with V, it is equivalent to converting the originally complex MIMO channel into multiple independent and parallel equivalent sub-channels. The gain of each sub-channel corresponds to a singular value, eliminating mutual interference between antenna signals and greatly improving transmission reliability.

[0027] The precoded signal z is transmitted to the MIMO channel through a pre-configured transmit antenna array at the transmitter, and the received signal is obtained at the receiver through a pre-configured receive antenna array. The transmit antenna array consists of... The system consists of physical antennas (e.g., 16 antennas), each equipped with an RF transmission module. The RF transmission module operates as follows: the k-th row of data (k complex symbols) of the precoded signal z is first input to a digital-to-analog converter (DAC) to convert the digital signal into a continuous analog voltage signal; the analog signal enters an up-converter (Mixer) and is multiplied by a high-frequency carrier signal (e.g., 5 GHz) generated by a local oscillator to shift the baseband signal to the RF band; the RF signal is amplified to sufficient power (e.g., 100 milliwatts) by a power amplifier (PA) and finally radiated into space through the k-th antenna. The antennas transmit simultaneously, and the electromagnetic waves propagate through space, reaching the receiver after fading and noise interference from the wireless channel. The receiver is equipped with... The system uses 16 receiving antennas, each capturing electromagnetic signals in space. These signals then enter the RF receiving module. The RF receiving module works as follows: the weak RF signals received by the antennas first enter a low-noise amplifier (LNA) for low-noise amplification to prevent additional noise from being introduced into subsequent processing; the amplified signal then enters a downconverter, where it is multiplied by the local oscillator signal to downconvert the RF signal to the baseband frequency band; the baseband analog signal is then input to an analog-to-digital converter (ADC), sampled and quantized into a digital signal to obtain the received signal. , dimension ×d, where the j-th row represents the d complex symbols received by the j-th receiving antenna. The mathematical representation of the received signal is: =Hz + N, where H is the actual channel matrix (determined by the electromagnetic wave propagation path from the transmitting antenna to the receiving antenna, including characteristics such as signal attenuation and phase shift), and N is the additive noise (including receiver thermal noise and environmental electromagnetic interference, each element independently follows a complex Gaussian distribution).

[0028] Received signal After acquisition, equalization processing is performed to recover the transmitted coding features. The equalization process is as follows: The estimated channel matrix... Perform singular value decomposition (which was already completed in the precoding step; the decomposition result is used directly here) to obtain the left singular matrix U, and then calculate the conjugate transpose of U. (After transposing U, the real part of each element remains unchanged, and the sign of the imaginary part changes). The conjugate transpose matrix... (dimension) × ) and received signal (dimension) ×d) Perform matrix multiplication to obtain the equilibrium characteristics. , dimension ×d. The equilibrium characteristic can be expressed as It utilizes The SVD decomposition form. Since the left singular matrix U produced by the SVD decomposition has the unitary matrix property (each column vector of U is orthogonal and has a length of 1), it satisfies... ( The identity matrix has diagonal elements of 1 and all other elements of 0. Therefore, the equilibrium characteristic simplifies to: The technical effect of equalization is as follows: the precoding step transforms the signal using the right singular matrix V to adapt it to the characteristics of the transmitting end of the channel; the equalization step performs an inverse transformation on the received signal using the conjugate transpose of the left singular matrix U, eliminating signal mixing and interference introduced by the channel matrix H, and separating and restoring the aliased signals received by multiple receiving antennas into signal components corresponding to each transmitting antenna. Since the estimated channel matrix is ​​actually used... Features after equalization, rather than the true channel matrix H The data still contains residual mismatch interference and noise caused by channel estimation errors. The subsequent receiver channel adapter will compensate for these errors and restore accurate semantic features.

[0029] 104. Based on the estimated channel matrix, perform first channel sub-decoding, receiver channel adaptation processing and second channel sub-decoding on the equalization features in sequence to obtain decoded semantic features, and perform semantic decoding on the decoded semantic features to reconstruct the data.

[0030] In this embodiment, the equalization feature undergoes a first dimensional expansion transformation to obtain a first decoding feature; based on the estimated channel matrix, the first decoding feature is adapted to the receiver channel to obtain a second adapted feature (wherein, adapting the first decoding feature to the receiver channel based on the estimated channel matrix to obtain the second adapted feature includes: flattening the estimated channel matrix in terms of dimensions, and projecting the flattened estimated channel matrix into receiver channel state words; concatenating the receiver channel state words with the first decoding feature to obtain an extended decoding sequence; performing multi-layer feature interaction on the extended decoding sequence through a multi-head self-attention mechanism to obtain an enhanced decoding sequence, and extracting feature sequences corresponding to the positions of the first decoding features from the enhanced decoding sequence to generate the second adapted feature); the second adapted feature undergoes a second dimensional expansion transformation to obtain a decoding semantic feature that represents the same feature dimension as the encoded semantic feature. The channel signal-to-noise ratio (SNR) parameter is obtained, and multiple decoding scale adjustment factors are generated based on the channel SNR parameter. Multi-scale feature reconstruction is performed on the decoded semantic features to obtain a multi-level reconstructed feature map. Based on each of the decoding scale adjustment factors, an affine transformation of channel-by-channel scaling and offset is performed on the multi-level reconstructed feature map to obtain an adjusted multi-level reconstructed feature map. Spatial resolution restoration and detail enhancement are performed on the adjusted multi-level reconstructed feature map to obtain reconstructed data.

[0031] In practical applications, equilibrium characteristics After acquisition, the data enters the first channel sub-decoder for the first dimensionality expansion transformation. The first channel sub-decoder uses a fully connected layer network, and the implementation process is as follows: the input equalization features are processed... (dimension is) ,in (where d is the number of transmitting antennas and d is the number of symbols per antenna). The input consists of a fully connected neural network with 1 to 2 layers. Each layer performs matrix multiplication with the input using a weight matrix, adds a bias vector, and then passes through a ReLU activation function to progressively increase the feature dimension. The final output dimension is... First decoding feature , where d' is the intermediate decoding feature dimension (e.g., 128). Dimension expansion transformation is the inverse process of dimension compression transformation, expanding the feature vector corresponding to each antenna from d dimensions to d' dimensions, increasing the capacity of the feature representation to accommodate more information. This can be expressed by the formula: ,in This represents the first channel sub-decoding function. This represents the learnable parameters of the decoder (including the weight matrix and bias vector of the fully connected layer). The first dimensionality expansion transformation restores the compressed features after MIMO channel transmission and equalization to the intermediate dimension, providing sufficient feature space for error compensation and information recovery for the subsequent receiver channel adapter, while maintaining the matrix structure so that the feature vector corresponding to each receiving antenna can be processed independently.

[0032] First decoding feature After obtaining it, based on the estimated channel matrix Perform receiver-side channel adaptation processing. This step is symmetrical to the transmitter-side channel adaptation structure. First, estimate the channel matrix H{est} (dimension 1). Line × The matrix (columns) is flattened, and its elements are arranged into a one-dimensional vector in row order. The flattened vector is then mapped through a fully connected layer to a receiver channel state term of dimension d', which contains the compressed and encoded representation of the estimated channel matrix. After obtaining the receiver channel state term, it is inserted as a new row vector into the first decoded feature. The 0th row of the matrix (where the original 1st row becomes the 1st row, the original 2nd row becomes the 2nd row, and so on) forms the extended decoding sequence, with dimension [missing information]. It contains one channel state term and Each antenna feature vector is used. The extended decoded sequence is input into a Transformer structure for feature interaction. This structure contains four Transformer blocks, each using a multi-head self-attention mechanism to propagate channel state lexical information from the sequence to the feature vector of each antenna. Simultaneously, the antenna features can also perceive and interact with each other. After processing through four Transformer blocks, the enhanced decoded sequence is obtained, still retaining its dimensionality of [missing information]. Remove line 0 (channel state term) from the enhanced decoded sequence, retaining the subsequent lines. The rows are used to extract the matrix formed by these rows as the second adaptation feature. , dimension Expressed as a formula ,in This represents the receiver channel matrix adapter function. These represent the learnable parameters of the adapter. The technical effect of receiver channel adaptation lies in: equalization characteristics. Residual mismatch interference and transmission noise caused by channel estimation errors can distort features, and direct semantic decoding can lead to a decrease in the quality of reconstructed images. By introducing receiver channel state terms and utilizing the global information interaction mechanism of Transformer, the decoding features corresponding to each antenna can be adaptively corrected according to the characteristics of the estimated channel matrix, compensating for the distortion caused by channel mismatch and significantly improving the signal recovery accuracy under non-ideal channel conditions.

[0033] Second adaptation feature The input then undergoes a second dimensionality expansion transformation in the second channel sub-decoder. The second channel sub-decoder also employs a fully connected layer network to transform the input... The dimensional features are further expanded, specifically by: The two-dimensional matrix is ​​flattened into a one-dimensional vector in row order, and then the dimension is mapped through a fully connected layer containing two to three layers, reducing the vector length from... Expanding to c'×h'×w' (e.g., from 16×128=2048 to 512×4×4=8192), and then rearranging the one-dimensional vector according to the dimensions of c', h', and w' into a three-dimensional tensor, we obtain the decoded semantic features. The dimension is c'×h'×w', which is exactly the same as the dimension of the encoded semantic features. This can be expressed by the formula: ,in This represents the second channel sub-decoding function. This represents the learnable parameters of the decoder. For example, if the second adaptation feature dimension is 16×128 and the target decoding semantic feature dimension is 512×4×4, the fully connected layer first maps 16×128=2048 values ​​to 8192 values, and then reorganizes these 8192 values ​​into a three-dimensional tensor with 512 channels and a spatial size of 4×4 for each channel. The second dimension expansion transformation completes the final transformation from the intermediate features optimized by channel adaptation to the semantic feature space, restoring the same feature dimension and data structure as the encoded semantic features, enabling the subsequent semantic decoder to reconstruct these features into the original image. The three-layer channel decoding structure is symmetrically designed with the three-layer channel coding structure of the transmitter. The first layer expands the feature dimension, the second layer performs channel adaptation compensation, and the third layer restores the semantic feature dimension, progressively eliminating various distortions caused by MIMO channel transmission and channel estimation errors, ensuring accurate recovery of semantic information.

[0034] Decoding semantic features After acquisition, the image is reconstructed using an SNR-adaptive semantic decoder. The semantic decoder first obtains the channel signal-to-noise ratio parameter *r* (this parameter is obtained at the encoding end via the feedback channel; the decoding end directly reads the same *r* value to maintain encoding-decoding consistency). Based on *r*, it generates multiple decoding scale adjustment factors through a multilayer perceptron network. The input to the multilayer perceptron network is a single numerical value *r*, and the output is a set of numerical vectors. The vector length is twice the number of feature layers generated by subsequent multi-scale feature reconstruction. Each feature map corresponds to a pair of decoding scale adjustment factors: a scaling factor *γ* and an offset factor *β*. The decoding scale adjustment factors are generated in the same way as the encoding scale adjustment factors, but the parameters are independent, allowing for adaptive adjustment based on the characteristics of the decoding process. The semantic decoder uses a Swin-Transformer as the backbone network to decode semantic features. Multi-scale feature reconstruction is performed. The Swin-Transformer's decoding mode is the opposite of its encoding mode; it reconstructs image features at different spatial scales from compressed semantic features through layer-by-layer upsampling and feature abstraction. The specific implementation process is as follows: Decoding semantic features. (Dimensions 512×4×4) First, features are processed through several Swin-Transformer blocks, then spatial upsampling is performed through transposed convolutional layers (resolution doubled, feature channel count halved), outputting multi-level reconstructed feature maps at different network depths. For example, the first layer outputs a deep reconstructed feature map. (Dimensions are 512×4×4), the second layer outputs the reconstructed feature map of the middle layer. (Dimensions are 256×8×8), the third layer outputs a shallow reconstructed feature map. (Dimensions are 128×16×16). After obtaining the multi-level reconstructed feature maps, each layer of feature map undergoes a channel-wise scaling and offset affine transformation according to the corresponding decoding scale adjustment factor. Channel-wise means: for the i-th layer feature map... For each feature channel (e.g., the first layer has 512 channels), the feature values ​​of all spatial locations within that channel are multiplied element-wise by the corresponding scaling factor. Then add the offset factor to each element. The adjusted feature map is obtained. Affine transformation enables the numerical distribution of the decoded feature map to be adaptively adjusted according to the channel quality, suppressing noise amplification under low SNR conditions and finely recovering image details under high SNR conditions, forming a closed-loop optimization with the adaptive adjustment at the encoding end.

[0035] The adjusted multi-level reconstructed feature maps are then subjected to spatial resolution restoration and detail enhancement to obtain reconstructed data. Spatial resolution restoration is achieved by using transposed convolutional layers (also known as deconvolutional layers, which achieve upsampling by sliding the convolutional kernel across the input feature map and inserting padding values ​​at each position) to progressively enlarge the spatial size of the deep feature maps. For example, the 512×4×4 feature map of the first layer is upsampled to 512×8×8 through a transposed convolutional layer, then upsampled to 512×16×16, ultimately restoring it to the same spatial resolution as the input image (e.g., 32×32). Detail enhancement is achieved by using skip connections between reconstructed feature maps of different levels. Skip connections refer to the operation of directly passing shallow feature maps to deep layers and fusing them with deep feature maps. This preserves the detailed information (such as edges and textures) in shallow features while utilizing the semantic information (such as objects and layouts) in deep features. Specifically, the shallow reconstructed feature maps... The 128×16×16 deep reconstructed feature map is concatenated with the upsampled deep reconstructed feature map (upsampled to the same 16×16 spatial size) along the channel dimension (e.g., 128 channels + 256 channels = 384 channels) to obtain a fused feature map (384×16×16). The fused feature map is then refined through several convolutional layers. Each convolutional layer uses a kernel to slide across the feature map to extract local patterns, gradually adjusting the number of feature channels to match the input image (e.g., 3 color channels), and outputting the reconstructed data. (Dimensions are 3×32×32). Expressed as a formula: ,in This represents the reconstructed output image. This represents a semantic decoding function. This represents the learnable parameters of the semantic decoder. Spatial resolution restoration and detail enhancement ensure that the reconstructed image approximates the original input image as closely as possible in terms of spatial size, pixel detail, and color reproduction, completing end-to-end reconstruction from compressed semantic features to a complete image. The SNR adaptive semantic decoding mechanism enables the decoder to dynamically adjust the reconstruction strategy according to channel quality. When channel conditions are poor, it prioritizes ensuring the image's recognizability and the correctness of key semantic content, while when channel conditions are good, it fully restores the image's fine details and color fidelity, achieving high-quality semantic communication in various wireless channel environments.

[0036] In this embodiment of the invention, by acquiring the data to be transmitted and the estimated channel matrix between the transmitting and receiving ends, semantic encoding is performed on the data to be transmitted to obtain encoded semantic features. Based on the estimated channel matrix, the encoded semantic features are sequentially processed by first channel sub-coding, transmitting end channel adaptation processing, and second channel sub-coding to obtain transmitting encoded features. The transmitting encoded features are pre-coded, and the pre-coded signal is transmitted to the receiving end through a MIMO channel to obtain the received signal. The received signal is then equalized based on the estimated channel matrix to obtain equalization features. Based on the estimated channel matrix, the equalization features are sequentially processed by first channel sub-decoding, receiving end channel adaptation processing, and second channel sub-decoding to obtain decoded semantic features. The decoded semantic features are then semantically decoded to reconstruct the data. This application constructs a three-layer channel processing architecture by embedding a channel matrix adapter in the channel codec. It dynamically senses and compensates for channel estimation errors in the intermediate stage of feature compression, achieving fine-grained adjustment of semantic features. This solves the problem of sharp performance degradation caused by channel mismatch in the actual deployment of existing MIMO semantic communication systems, effectively improving the robustness and transmission quality of the system under non-ideal channel state information conditions, and ensuring high-fidelity data reconstruction in large-scale antenna array scenarios.

[0037] The above describes the MIMO semantic communication method against channel mismatch in the embodiments of the present invention. The following describes the MIMO semantic communication system against channel mismatch in the embodiments of the present invention. Please refer to [link / reference]. Figure 3 One embodiment of the MIMO semantic communication system resistant to channel mismatch in this invention includes: The semantic coding module 201 is used to acquire the data to be transmitted and the estimated channel matrix between the sending end and the receiving end, and to perform semantic coding on the data to be transmitted to obtain coded semantic features. Channel coding module 202 is used to perform first channel sub-coding, transmitter channel adaptation processing and second channel sub-coding on the coding semantic features in sequence based on the estimated channel matrix to obtain transmission coding features; The transmission processing module 203 is used to precode the transmission coding features, transmit the precoded signal to the receiving end through the MIMO channel to obtain the received signal, and equalize the received signal based on the estimated channel matrix to obtain equalization features. The receiving decoding module 204 is used to perform first channel sub-decoding, receiver channel adaptation processing and second channel sub-decoding on the equalization features in sequence based on the estimated channel matrix to obtain decoded semantic features, and to perform semantic decoding on the decoded semantic features to reconstruct the data.

[0038] In this embodiment of the invention, by acquiring the data to be transmitted and the estimated channel matrix between the transmitting and receiving ends, semantic encoding is performed on the data to be transmitted to obtain encoded semantic features. Based on the estimated channel matrix, the encoded semantic features are sequentially processed by first channel sub-coding, transmitting end channel adaptation processing, and second channel sub-coding to obtain transmitting encoded features. The transmitting encoded features are pre-coded, and the pre-coded signal is transmitted to the receiving end through a MIMO channel to obtain the received signal. The received signal is then equalized based on the estimated channel matrix to obtain equalization features. Based on the estimated channel matrix, the equalization features are sequentially processed by first channel sub-decoding, receiving end channel adaptation processing, and second channel sub-decoding to obtain decoded semantic features. The decoded semantic features are then semantically decoded to reconstruct the data. This application constructs a three-layer channel processing architecture by embedding a channel matrix adapter in the channel codec. It dynamically senses and compensates for channel estimation errors in the intermediate stage of feature compression, achieving fine-grained adjustment of semantic features. This solves the problem of sharp performance degradation caused by channel mismatch in the actual deployment of existing MIMO semantic communication systems, effectively improving the robustness and transmission quality of the system under non-ideal channel state information conditions, and ensuring high-fidelity data reconstruction in large-scale antenna array scenarios.

[0039] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0040] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A MIMO semantic communication method resistant to channel mismatch, characterized in that, The MIMO semantic communication method that resists channel mismatch includes: The data to be transmitted and the estimated channel matrix between the sender and receiver are obtained, and the data to be transmitted is semantically encoded to obtain the encoded semantic features. Based on the estimated channel matrix, the first channel sub-coding, the transmitter channel adaptation process, and the second channel sub-coding are sequentially performed on the encoded semantic features to obtain the transmission coding features. The transmitted coding features are precoded, and the precoded signal is transmitted to the receiving end through a MIMO channel to obtain the received signal. Based on the estimated channel matrix, the received signal is equalized to obtain equalization features. Based on the estimated channel matrix, the equalization features are sequentially subjected to first channel sub-decoding, receiver channel adaptation processing, and second channel sub-decoding to obtain decoded semantic features. Then, the decoded semantic features are semantically decoded to reconstruct the data.

2. The MIMO semantic communication method for resisting channel mismatch according to claim 1, characterized in that, The semantic encoding of the data to be transmitted to obtain the encoded semantic features includes: Obtain the channel signal-to-noise ratio (SNR) parameters, and generate multiple scale adjustment factors based on the channel SNR parameters; Multi-scale feature extraction is performed on the data to be transmitted to obtain multi-level feature maps, and affine transformation is performed on the multi-level feature maps according to each of the scale adjustment factors to obtain adjusted multi-level feature maps. The adjusted multi-level feature maps are subjected to feature fusion and dimensional transformation to obtain encoded semantic features.

3. The MIMO semantic communication method for resisting channel mismatch according to claim 1, characterized in that, Based on the estimated channel matrix, the first channel sub-coding, the transmitter channel adaptation process, and the second channel sub-coding are sequentially performed on the encoded semantic features to obtain the transmitted encoded features, including: The encoded semantic features are subjected to a first-dimensional compression transformation to obtain the first encoded features; Based on the estimated channel matrix, the first coding feature is adapted to the transmitting channel to obtain the first adapted feature; The first adaptation feature is subjected to a second dimensional compression transformation to obtain a second coding feature whose dimension matches the number of transmit antennas. The second coding feature is then power normalized to obtain the transmit coding feature.

4. The MIMO semantic communication method for resisting channel mismatch according to claim 3, characterized in that, The step of performing transmitter channel adaptation on the first coding feature based on the estimated channel matrix to obtain the first adapted feature includes: The estimated channel matrix is ​​dimensionally flattened, and the flattened estimated channel matrix is ​​projected as a transmitter channel state term; The transmitting end channel state lexical is added to the beginning of the sequence of the first encoded feature to obtain the extended feature sequence; The extended feature sequence is subjected to multi-level feature interaction through a multi-head self-attention mechanism to obtain an enhanced feature sequence. Channel state terms are removed from the enhanced feature sequence, and feature sequences corresponding to the first coding feature positions are extracted to generate the first adaptation feature.

5. The MIMO semantic communication method for resisting channel mismatch according to claim 1, characterized in that, The step of precoding the transmitted coding features and transmitting the precoded signal to the receiving end through a MIMO channel to obtain the received signal includes: Singular value decomposition is performed on the estimated channel matrix to obtain the right singular matrix; Perform matrix multiplication on the right singular matrix and the transmitted coding feature to obtain the precoded signal; The precoded signal is transmitted to the MIMO channel through a preset transmit antenna array, and the receiver obtains the received signal through a preset receive antenna array.

6. The MIMO semantic communication method for resisting channel mismatch according to claim 1, characterized in that, The process of equalizing the received signal based on the estimated channel matrix to obtain equalization features includes: Singular value decomposition is performed on the estimated channel matrix to obtain the left singular matrix, and the conjugate transpose of the left singular matrix is ​​calculated. The equalization feature is obtained by performing matrix multiplication between the conjugate transpose matrix and the received signal.

7. The MIMO semantic communication method for resisting channel mismatch according to claim 1, characterized in that, The equalization features are sequentially subjected to first channel sub-decoding, receiver channel adaptation processing, and second channel sub-decoding to obtain decoded semantic features including: The first dimension expansion transformation is performed on the equalization feature to obtain the first decoding feature; Based on the estimated channel matrix, the first decoding feature is adapted to the receiver channel to obtain the second adaptation feature; The second adaptation feature is subjected to a second dimensional expansion transformation to obtain a decoded semantic feature that represents the same feature dimension as the encoded semantic feature.

8. The MIMO semantic communication method for resisting channel mismatch according to claim 7, characterized in that, The step of performing receiver channel adaptation on the first decoding feature based on the estimated channel matrix to obtain the second adaptation feature includes: The estimated channel matrix is ​​dimensionally flattened, and the flattened estimated channel matrix is ​​projected as receiver channel state terms; The receiving end channel state vocabulary is concatenated with the first decoding feature to obtain an extended decoding sequence; The extended decoding sequence is subjected to multi-layer feature interaction through a multi-head self-attention mechanism to obtain an enhanced decoding sequence. Then, the feature sequence corresponding to the first decoding feature position is extracted from the enhanced decoding sequence to generate the second adaptation feature.

9. The MIMO semantic communication method for resisting channel mismatch according to claim 1, characterized in that, The step of semantically decoding the decoded semantic features and reconstructing the data includes: The channel signal-to-noise ratio (SNR) parameter is obtained, and multiple decoding scale adjustment factors are generated based on the channel SNR parameter. Multi-scale feature reconstruction is performed on the decoded semantic features to obtain a multi-level reconstructed feature map. Then, according to the decoding scale adjustment factors, the multi-level reconstructed feature map is subjected to a channel-by-channel scaling and offset affine transformation to obtain an adjusted multi-level reconstructed feature map. Spatial resolution restoration and detail enhancement are performed on the adjusted multi-level reconstructed feature maps to obtain reconstructed data.

10. A MIMO semantic communication system resistant to channel mismatch, characterized in that, The MIMO semantic communication system that is resistant to channel mismatch includes: The semantic coding module is used to acquire the data to be transmitted and the estimated channel matrix between the sender and receiver, and to perform semantic coding on the data to be transmitted to obtain coded semantic features. The channel coding module is used to perform first channel sub-coding, transmitter channel adaptation processing and second channel sub-coding on the coding semantic features based on the estimated channel matrix to obtain the transmission coding features. The transmission processing module is used to precode the transmission coding features, transmit the precoded signal to the receiving end through the MIMO channel to obtain the received signal, and equalize the received signal based on the estimated channel matrix to obtain equalization features. The receiving decoding module is used to perform first channel sub-decoding, receiver channel adaptation processing and second channel sub-decoding on the equalization features in sequence based on the estimated channel matrix to obtain decoded semantic features, and to perform semantic decoding on the decoded semantic features to reconstruct the data.