Semantic communication retransmission method and system based on long short-term memory network, terminal and storage medium
By adopting the semantic communication retransmission method based on LSTM in semantic communication, the existing decoder structure is solved, and the complete transmission of semantic content and the improvement of model generalization capabilities is achieved.
Patent Information
- Application Number
- CN202510527724.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing semantic retransmission decoder has complex structure, which increases the complexity and implementation cost of the system. It is difficult to train and weak generalization ability, so it is unable to effectively handle different retransmissions and changing channel conditions.
The semantic communication retransmission method based on long and short-term memory network (LSTM) is adopted. By acquiring image data sets, semantic features are extracted, and transmitting or retransmitting in a noisy wireless channel, the LSTM classifier is used to process and classify the noisy semantic features, and output image classification results.
Simplify the decoder structure, effectively capture and retain semantic features at different time steps during the transmission process, ensure the complete transmission of semantic content, improve the generalization ability of the model, and reduce the complexity of the model and training complexity.
Smart Images

Figure CN120050004A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information and communication technologies, and particularly to a semantic communication retransmission method, system, terminal, and storage medium based on a long short-term memory network. Background Art
[0002] With the rapid development of modern communication technologies, traditional communication methods are facing increasingly severe challenges. Traditional communication systems are based on Shannon's theory, which simplifies the communication process to the bit-level transmission and recovery of information. Their core objective is to ensure the symbol-level accuracy of channel transmission, while ignoring the semantic value of the information itself. This paradigm has demonstrated remarkable efficacy in traditional application scenarios. However, in the face of the explosive growth of emerging intelligent applications such as the Internet of Things, autonomous driving, virtual reality, and augmented reality, its transmission efficiency bottleneck has become increasingly prominent. Against this backdrop, the vision of intelligent interconnection of all things proposed by the Sixth Generation Wireless Communication System (6G) poses a twofold innovation requirement for communication technologies: on the one hand, there is an urgent need to achieve a deep integration of communication systems and artificial intelligence (AI) technologies; on the other hand, it is required to construct a more efficient intelligent encoding and decoding mechanism. As a core technical direction of the 6G system, semantic communication (SC), empowered by a deep neural network (DNN), has opened up an innovative path to break through the traditional efficacy bottleneck. Compared with simply pursuing the fidelity of symbol transmission, a semantic communication system focuses on the intelligent extraction and encoding of task-related semantic information, achieving an improvement in transmission efficiency by eliminating redundant information. With the help of DNN, a semantic communication system can better understand environmental information and given tasks. A large number of studies have shown that this new communication paradigm exhibits significant advantages in data-intensive and latency-sensitive intelligent scenarios such as image classification, speech recognition, and video transmission.
[0003] To ensure efficient semantic transmission in a noisy wireless channel, systematic optimization is required from two dimensions: the coding and decoding mechanism and transmission reliability. First, by designing Joint Source and Channel Coding (JSCC), the training of the end-to-end DNN model is realized, and then the system-level optimization of the coding process is achieved. To simplify the analysis and training process, the semantic feature vector is usually represented as continuous real values and transmitted through the channel, while the receiving end receives a noisy version of the continuous feature vector. When modern digital communication systems transmit feature vectors, the feature vectors need to be quantized, encoded, and modulated into QAM (Quadrature Amplitude Modulation) signals for transmission. At the same time, the receiving end needs to demodulate the signal through digital-to-analog conversion and reconstruct the feature vector. The distortion introduced by quantization and modulation may seriously damage the semantic structure of the feature vector, resulting in a decline in the performance of intelligent tasks. Second, Automatic Repeat reQuest (ARQ) is widely used in digital communication systems to achieve reliable data transmission over a noisy wireless channel. By retransmitting the extracted feature vectors, the reliability of communication is guaranteed, thus better completing intelligent tasks.
[0004] If data is transmitted in an extremely poor channel condition or other harsh communication environments, no retransmission or a single retransmission may not be able to obtain the desired inference result. In this case, multiple retransmissions may be required to ensure transmission reliability. Currently, there are mainly two types of decoder structures for semantic retransmission in communication systems: Multi-decoders and Unified Single Decoder. In the multi-decoder structure, multiple decoders need to be trained and maintained, which increases the complexity and computational resources of the system. Second, multiple decoders may learn duplicate parameters, especially when the task features are similar, thus reducing the efficiency of the model. In the unified single decoder structure, since all transmitted information is processed by the same decoder, the performance of the multi-decoder scheme may not be achieved in some specific cases (such as when the number of retransmissions is specified in advance); the single decoder needs to be able to process coding information of different lengths, which increases the difficulty of training.
[0005] In summary, the existing decoders for semantic retransmission have the following technical defects: 1) Whether maintaining a complex single decoder or multiple specialized decoders will increase the complexity and implementation cost of the system; 2) The training difficulty is high. The existing decoder structures need to be trained to adapt to various transmission situations, especially when dealing with different numbers of retransmissions, which may involve complex training processes and tuning; 3) Weak generalization ability, insufficient generalization ability when dealing with new transmission situations, especially when facing changing channel conditions and transmission requirements.
[0006] Therefore, the existing technology still needs to be improved. Summary of the Invention
[0007] The technical problem to be solved by the present invention is that, aiming at the defects of the existing technology, the present invention provides a semantic communication retransmission method, system, terminal and storage medium based on a long short-term memory network to solve the problems of complex system and low generalization ability existing in the existing decoding schemes for semantic retransmission.
[0008] The technical solution adopted by the present invention to solve the technical problem is as follows: In the first aspect, the present invention provides a semantic communication retransmission method based on a long short-term memory network, including: Obtain an image data set to be classified and processed; Extract semantic features from the image data set based on the feature extractor of the classification network; Transmit or retransmit the extracted semantic features, and process and classify the noisy semantic features after transmission or retransmission based on the LSTM classifier of the classification network; Output the image classification result corresponding to the image data set.
[0009] In one implementation, the extracting semantic features from the image data set based on the feature extractor of the classification network includes: Use the feature extractor deployed on the edge device and hyperparameters to extract semantic features from the images in the image data set; Quantize the extracted semantic features through an analog-to-digital conversion module, and modulate them through a digital modulator to obtain the modulated semantic features.
[0010] In one implementation, the transmitting or retransmitting the extracted semantic features includes: Transmit the modulated semantic features to the LSTM classifier deployed on the edge server through a noisy and fading wireless channel; Or based on the noisy and fading wireless channel, use a retransmission mechanism for retransmitting modulated features to retransmit the modulated semantic features, add new Gaussian white noise, and transmit the retransmitted signal to the LSTM classifier.
[0011] In one implementation, the processing and classifying the noisy semantic features after transmission or retransmission based on the LSTM classifier of the classification network includes: Scale the signal corresponding to the semantic feature with noise; Perform demodulation and DAC digital-to-analog conversion on the scaled signal to obtain the estimated semantic feature; Input the estimated semantic feature into the LSTM classifier with hyperparameters to obtain the image classification result. In one implementation, the LSTM classifier based on the classification network processes and classifies the semantic feature with noise after transmission or retransmission, and further includes:
[0012] Perform demodulation and DAC digital-to-analog conversion on the retransmitted semantic feature to obtain the retransmitted noisy semantic feature; Input the retransmitted noisy semantic feature into the LSTM classifier and jointly decode it with the feature before retransmission to obtain the image classification result with retransmission. In one implementation, the method further includes:
[0013] According to the preset uniform quantization parameter, divide the range of the input feature extracted from the image into equal intervals; Train the classification network based on the function and the cross-entropy loss function. In one implementation, training the classification network based on the
[0014] function and the cross-entropy loss function includes: Randomly select training batches from the training dataset ; Extract the corresponding semantic features for the selected dataset to obtain the quantization result ; ; When transmitting and retransmitting the same batch of pictures, calculate the signal with noise , where the channel signal-to-noise ratio is the same; when transmitting and retransmitting different pictures, generate random signal-to-noise ratios with the same mean and noise , and calculate the signal with noise ; Obtain the estimated feature ; Obtain the classification result ; Calculate the training loss based on the cross-entropy loss function , and update the parameters through backpropagation, , until the parameter , converges.
[0015] In a second aspect, the present invention provides a semantic communication retransmission system based on a long short-term memory network, including: An acquisition module, configured to acquire an image data set to be classified and processed; A feature extraction module, configured to extract semantic features from the image data set based on a feature extractor of a classification network; A classification module, configured to transmit or retransmit the extracted semantic features, and process and classify the noisy semantic features after transmission or retransmission based on an LSTM classifier of the classification network; An output module, configured to output an image classification result corresponding to the image data set.
[0016] In a third aspect, the present invention provides a terminal, including: a processor and a memory, where the memory stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on the long short-term memory network is executed by the processor, it is used to implement the operations of the semantic communication retransmission method based on the long short-term memory network as described in the first aspect.
[0017] In a fourth aspect, the present invention further provides a medium, where the medium is a computer-readable storage medium, and the medium stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on the long short-term memory network is executed by a processor, it is used to implement the operations of the semantic communication retransmission method based on the long short-term memory network as described in the first aspect.
[0018] The present invention adopts the above technical solutions and has the following effects: The present invention uses an LSTM classifier to replace the traditional multi-decoder structure and unified single-decoder structure. While simplifying the decoder structure, it can effectively capture and retain semantic features at different time steps during the transmission process, ensuring the complete transmission of semantic content; at the same time, it supports dynamic-length input, and sequentially inputs the semantic features of each retransmission into each time step of the LSTM network, which is suitable for the semantic retransmission scenario; the present invention improves the generalization ability, avoids the problem that the number of tests of the multi-decoder structure must be equal to the number of trainings, and reduces the complexity of the model and the complexity of training. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the structures shown in these drawings.
[0020] Figure 1 is a flowchart of the semantic communication retransmission method based on long short-term memory network in the present invention.
[0021] Figure 2 is a schematic diagram of the LSTM unit structure in the present invention.
[0022] Figure 3 is a schematic diagram of the classification network system model in the present invention.
[0023] Figure 4 is a schematic diagram of the convergence trend of the CLFNet model in the present invention.
[0024] Figure 5 is a functional schematic diagram of the terminal in an implementation manner of the present invention.
[0025] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the drawings. Specific Embodiments
[0026] To make the purpose, technical solutions, and advantages of the present invention clearer and more definite, the following will further describe the present invention in detail with reference to the embodiments and the drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0027] Exemplary Method In current communication systems, there are mainly two types of decoder structures for semantic retransmission: multi-decoders and unified single decoder. In the multi-decoder structure, N different decoders are respectively responsible for different transmission and retransmission tasks. Each time information is received, the corresponding decoder concatenates it with the previously transmitted information and decodes the corresponding task from this concatenated information. Here, the information can be the received bitstream or the feature vector. The advantage of this decoder structure is that with each retransmission, the decoder can use the previously transmitted information (incremental knowledge) to jointly perform task inference, thus gradually improving the transmission performance. However, there are also some problems with this structure. Firstly, it requires training and maintaining multiple decoders, which increases the complexity of the system and computational resources. Secondly, multiple decoders may learn duplicate parameters, especially when the task features are similar, thus reducing the efficiency of the model. For the second structure, the unified single decoder, regardless of how many times the information is transmitted, only uses a single unified decoder to process all transmitted information. This decoder is trained to be able to process encoded information of different sizes, and if the number of transmissions does not reach the preset maximum value, those extra dimensions will be automatically filled with zero vectors. During the training process, even at randomly selected numbers of transmissions, the system will learn how to extract effective information from the retransmitted semantic knowledge. For the unified single decoder structure, it only requires one decoder, which simplifies the system design. Compared with multiple decoders, the single decoder scheme saves more storage and computational resources. Secondly, it has higher flexibility. By filling with zero vectors to process encoded information of different lengths, the system can flexibly handle different transmission situations. In the unified single decoder structure, since all transmitted information is processed by the same decoder, it may not be able to achieve the performance of the multi-decoder scheme in some specific cases (such as when the number of retransmissions is specified in advance). Second, the single decoder needs to be able to process encoded information of different lengths, which increases the difficulty of training.
[0028] In summary, the existing decoders for semantic retransmission have the following technical defects: 1) Whether maintaining a complex single decoder or multiple specialized decoders will increase the complexity and implementation cost of the system; 2) The training difficulty is high. The existing decoder structures need to be trained to adapt to various transmission situations, especially when dealing with different numbers of retransmissions, which may involve complex training processes and tuning; 3) The generalization ability is weak, and the generalization ability is insufficient when dealing with new transmission situations, especially when facing changing channel conditions and transmission requirements.
[0029] In view of the above technical problems, an embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network. The method mainly obtains an image data set to be classified and processed; then, extracts semantic features from the image data set based on a feature extractor of a classification network; finally, transmits or retransmits the extracted semantic features, and processes and classifies the noisy semantic features after transmission or retransmission based on an LSTM classifier of the classification network; and outputs an image classification result corresponding to the image data set. In the embodiment of the present invention, the LSTM classifier is used to replace the traditional multi-decoder structure and the unified single-decoder structure, which simplifies the decoder structure while effectively capturing and retaining semantic features at different time steps during the transmission process, ensures the complete transmission of semantic content, and improves the generalization ability of the model.
[0030] As Figure 1 shown, an embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network, including the following steps: Step S100, obtaining an image data set to be classified and processed; Step S200, extracting semantic features from the image data set based on a feature extractor of a classification network.
[0031] In this embodiment, a decoding structure of a semantic retransmission system based on a long short-term memory network (LSTM) is designed. As a special recurrent neural network (RNN), LSTM is specifically designed to learn and remember long-term dependencies in sequential data. The LSTM-based semantic decoder uses its memory unit and three gates (forget gate, input gate, and output gate), and through the control of these gates, it decides which information needs to be remembered, updated, or forgotten. The cooperation of these three gates enables the LSTM to flexibly adjust the memory content according to the input data and tasks, retain important information, and obtain a high classification accuracy.
[0032] In this embodiment, the LSTM-based decoder structure has several important parameters: input_size, hidden_size, num_layers, num_classes, seq_length, which respectively represent the dimension of the LSTM input features, the dimension of the hidden layer, the number of stacked layers, the number of output classes, and the length of the input sequence. Among them, seq_length does not need to be specified, that is, the designed LSTM decoder structure in this embodiment can process variable-length sequences, which greatly reduces the training cost of the system and improves the flexibility of retransmission and subsequent optimization control.
[0033] In the scenario of the semantic retransmission system designed in this embodiment, the features of the first transmission obtained at the receiving end and the features of subsequent retransmissions are respectively used as the inputs for each time step of the LSTM network. Their dimensions are the same, that is, the size of input_size is the same. Through the action of the memory unit and three gates inside the LSTM, the classification result is finally obtained by a fully connected layer. The LSTM-based decoder structure designed in this embodiment can be used to process variable-length sequences, solving the drawback that traditional decoders must be trained for a specific number of transmissions, greatly reducing the training complexity, improving the flexibility in processing retransmitted data, and the generalization ability of classification.
[0034] LSTM is a special type of RNN designed to address the problem of long-term dependencies that traditional RNNs cannot handle. Although traditional RNNs can capture correlations in sequences, they are prone to the problems of vanishing gradients or exploding gradients when dealing with longer sequences, making it difficult for the model to remember earlier inputs. LSTM controls the flow of information through a special "gating" mechanism, enabling it to effectively handle long-term dependencies.
[0035] Table 1 Variable Annotation Table
[0036] An LSTM cell mainly consists of five parts: the cell state, the hidden state, the forget gate, the input gate, and the output gate. These gates are controlled by trainable neural network layers, through which the LSTM cell can selectively remember or discard information to achieve the effect of long-term memory.
[0037] (1) Cell state: The cell state is a main line running through the LSTM cell, transmitting information between previous and subsequent time steps. LSTM adjusts the cell state through the forget gate and the input gate, enabling the model to "remember" or "forget" certain information: (1); (2) Hidden state: The hidden state is the output of the LSTM and also serves as the input for the next time step. The hidden state not only affects the output of the current time step but also provides context information for the next time step: (2); (3) Forget gate: The forget gate determines which information needs to be discarded through a sigmoid function (an activation function). Based on the current input and the hidden state of the previous time step, it generates a vector with values between 0 and 1, where 0 represents "completely forget" and 1 represents "completely retain": (3); (4) Input Gate: The input gate controls how much new information is to be written into the cell state. The input gate consists of a sigmoid function layer and a tanh function layer (a type of activation function layer). The sigmoid function layer determines the proportion to be written, and the tanh function layer provides the candidate information: (4); (5) Output Gate: The output gate determines which information will be output from the current cell state. It filters the information in the cell state through sigmoid and tanh operations and outputs it to the next time step.
[0038] (5); The structural diagram of the LSTM cell is as Figure 2 shown.
[0039] Specifically, in one implementation manner of this embodiment, step S200 includes the following steps: Step S201, using the feature extractor deployed on the edge device and hyperparameters to extract semantic features from the images in the image dataset; Step S202, quantifying the extracted semantic features through an analog-to-digital conversion module and modulating them through a digital modulator to obtain the modulated semantic features.
[0040] In this embodiment, a semantic retransmission system for device-edge collaborative inference is considered, as Figure 3 shown. This system consists of an edge device (Edge Device, ED) and an edge server (Edge Server, ES), which cooperate to execute a shared AI task. Here, a ten-class classification task for pictures is considered (for example, CIFAR-10 dataset, MNIST dataset, Nature 10 dataset, etc. Here, the MNIST dataset is used, and each picture is composed of 28×28 pixel handwritten digital pictures from 0 to 9 for the ten-class classification task of pictures). This system uses a classification network based on DNN, and the classification network is divided into two parts: a feature extractor deployed on the ED and an LSTM classifier deployed on the ES. The ED transmits the extracted semantic features to the ES through a digital communication link, and the LSTM classifier located on the ES processes and classifies the received semantic features with noise. Below, the inference process of this system will be introduced.
[0041] In this embodiment, the image processed by the ED is represented as , where is the image dataset. The ED uses a DNN-based feature extractor and hyperparameters From the input image Extracting semantic features from , expressed as: (6); in, Quantization and digital modulator after analog-to-digital conversion (ADC) After modulation, the semantic features obtained are: (7); Transmitted from ED to ES through a noisy wireless channel.
[0042] In the above quantization and digital modulator process, the specific process is: first, Normalize to the range [0,1], the normalization formula is: (x - x_min) / (x_max - x_min), the normalized value is multiplied by , To quantize the number of bits, map it to an integer range. If in training mode, use the sin_round function (a SQL base function) for quantization, and if in non-training mode, use the torch.round function (a function that rounds each element in the input tensor to the nearest integer) for rounding quantization.
[0043] The quantized output is then converted into binary numbers for transmission, followed by QPSK (Quadrature PhaseShift Keying, a digital modulation technology) modulation. First, the input binary data stream is grouped into pairs of two bits, and then the symbol index is calculated and mapped to the constellation diagram, and finally the complex symbol is output.
[0044] like Figure 1 As shown, an embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network, comprising the following steps: Step S300, transmitting or retransmitting the extracted semantic features, and processing and classifying the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network; Step S400: outputting the image classification result corresponding to the image data set.
[0045] In this embodiment, an LSTM classifier different from the traditional LSTM structure is used to process and classify the semantic features with noise after transmission or retransmission; the structure of the LSTM network in this embodiment is similar to the traditional LSTM structure, and the network parameters are adjusted to make it more suitable for classification tasks. Finally, a fully connected layer is added to the last layer inside the LSTM to obtain the classification result.
[0046] After the data in this embodiment is input in the format given by the LSTM network, it is processed by the internal network structure and passes through the last fully connected layer to output the unnormalized probability (logits, referring to the original prediction value output by the last layer of the model), so as to obtain the image classification result corresponding to the image data set.
[0047] Specifically, in one implementation manner of this embodiment, step S300 includes the following steps: Step S301, transmit the modulated semantic features through a wireless channel with noise and fading to the LSTM classifier deployed on the edge server; Step S302, or based on the wireless channel with noise and fading, use a retransmission mechanism for retransmitting the modulated features to retransmit the modulated semantic features, add new Gaussian white noise, and transmit the retransmitted signal to the LSTM classifier.
[0048] Based on the above step S301, in one implementation manner of this embodiment, the step S300 includes the following steps: Step S3011, perform scaling processing on the signal corresponding to the semantic features with noise; Step S3012, perform demodulation and DAC digital-to-analog conversion on the scaled signal to obtain the estimated semantic features; Step S3013, input the estimated semantic features into the LSTM classifier with hyperparameters to obtain the image classification result.
[0049] In this embodiment, the signal received at the ES can be expressed as: (8); represents the channel fading coefficient, is the additive Gaussian white noise. Considering a slow fading scenario, where and the signal-to-noise ratio remain constant in each classification task, but vary independently between tasks. Assume that the channel information of the ES is perfect. After receiving the signal, the ES performs signal scaling: (9); Through demodulation and DAC digital-to-analog conversion , the estimated semantic features are obtained: (10); In this embodiment Demodulation is the inverse process of modulation. Specifically, first calculate the distance between the received signal and all possible symbols in the constellation diagram, find the nearest constellation point, convert it into binary bits, and finally restore it to a binary bit stream. Then restore the binary signal to the signal in the original quantized form and perform digital-to-analog conversion.
[0050] Digital-to-analog conversion, which is the inverse process of quantization. The specific process is as follows: restore the demodulated signal to an analog signal, including dividing by the quantization accuracy and de-normalization. The specific formula is x_analog = (x_max - x_min) / (2^q - 1) * x + x_min.
[0051] After obtaining the estimated semantic features, input them into the LSTM-based classifier with hyperparameters to obtain the classification result: (11); To combat channel fading and improve task performance, in the present invention, a retransmission mechanism that allows the ED to retransmit modulation features multiple times is considered. The principle of this mechanism is derived from Automatic Repeat-reQuest (ARQ). ARQ is an error correction protocol in the data link layer and transport layer of the OSI model. It performs reliable information transmission based on an unreliable service. If the sender does not receive an acknowledgment message within a certain period after transmission, it usually retransmits. Here, the ARQ protocol is used in the communication system to ensure the reliability of the features.
[0052] Based on the above step S302, in one implementation manner of this embodiment, the step S300 includes the following steps: Step S3021, perform demodulation and DAC digital-to-analog conversion on the retransmitted semantic features to obtain the noisy semantic features after retransmission; Step S3022, input the noisy semantic features after retransmission into the LSTM-based classifier and jointly decode them with the features before retransmission to obtain the image classification result with retransmission.
[0053] In this embodiment, the retransmission mechanism of the ED retransmitting modulation features , and the retransmitted signal is expressed as: (12); where additional white Gaussian noise . After receiving , the ES applies the same demodulation and DAC processes as in Equation (10) to obtain an estimate of the retransmission semantic feature , that is , where . Then, the ES will input into the LSTM-based classifier to obtain the prediction with retransmission: (13); In this embodiment, the overall framework of the classification network is as Figure 3 shown. The process of the classification network for processing data is to input the data in the shape of [batch_size, seq_length, input_size] into the LSTM. The LSTM will perform transformation processing through three gates and memory units, etc. at each time step according to Equations (1)-(5), etc., and finally pass through the fully connected layer to map to the class space to obtain the class probability.
[0054] In this embodiment, a semantic feature selection and weighting mechanism is adopted. Due to its memory unit and gating mechanism, the LSTM network can long-term remember important information and filter out irrelevant information; moreover, it improves the generalization ability for variable-length inputs. The LSTM network can flexibly handle different numbers of retransmission situations without redundant targeted training.
[0055] The embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network. Before steps S200~S300, the following steps are further included: Step S001, according to the preset uniform quantization parameter, divide the range of the input features extracted from the image into equal intervals.
[0056] The embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network. During the implementation of the entire method, the following is further included: Step S002, based on function and cross-entropy loss function to train the classification network.
[0057] Specifically, in an implementation manner of this embodiment, step S002 includes the following steps: Step S002a, randomly select training batches from the training data set ; Step S002b, extract the corresponding semantic features for the selected data set , obtain the quantization result ; Step S002c, when transmitting and retransmitting the same batch of pictures, calculate the signal with noise , where the channel signal-to-noise ratio is the same; when transmitting and retransmitting different pictures, generate random signal-to-noise ratios with the same mean and noise , calculate the signal with noise ; Step S002d, obtain the estimated features ; Step S002e, obtain the classification result ; Step S002f, calculate the training loss based on the cross-entropy loss function , update the parameters through backpropagation , , until the parameters , converge.
[0058] In this embodiment, the digital semantic communication training process (i.e., the training process of the classification network) is as follows: For the above classification network, in this embodiment, end-to-end training is performed on the classification network to jointly optimize the hyperparameters and . Since the feature extractor and classifier in the classification network are connected by a digital communication link, the extracted information is quantized and modulated before transmission. Considering uniform quantization, where the input range is divided into equal intervals. Among them, is the resolution of the ADC, that is, the quantization bit. In this embodiment, the default quantization bit is 4. Through uniform quantization, the real-valued input is rounded to the nearest integer value, and then represented as bit binary code, such as natural binary code. To mitigate the gradient explosion effect, a surrogate function is introduced in this embodiment: (14); Among them, is the input to be rounded, is the positive integer design parameter. provides a smooth and differentiable approximation of the rounding function, which makes the gradient propagate stably during the training process, thereby obtaining better convergence and better model performance. In this embodiment, is only used in the model training state to approximate quantization, while the torch.round() function is used for actual quantization operations during model inference.
[0059] Specifically, in this embodiment, is represented as The rounded output, i.e., . Refers to the features extracted by the feature extractor The quantization result during the training phase, i.e., after passing through the quantization function The output, i.e., Figure 3 in . During the model training process of the classification network, bypass the modulation process and transmit to ES. Therefore, the signal received at ES is . It should be noted that during the actual inference process, the discrete feature is represented as bit binary code, and then modulated into a complex signal before transmission .
[0060] Let be the training image set. In this embodiment, the cross-entropy loss function is used to train the classification network as: (15); where, is the number of image categories, is whether the image is the label of the th category, is the th retransmission classifier 's predicted output. In this embodiment, the algorithm flow of the classification network is summarized as follows:
[0061] Table 2 Network structure of CLFNet
[0062] In this embodiment, numerical simulation is used to evaluate the system performance. Consider a digit classification task and use MNIST as the dataset. The MNIST dataset consists of 70,000 images with 28×28 grayscale pixels. The dataset is divided into three parts: 48,000 training samples, 12,000 validation samples, and 10,000 test samples. The network structure of the classification network is shown in Table 2. In particular, in the classification network, in this embodiment, a tanh activation function is introduced in the last fully connected layer to limit the output value within the range of [−1, 1]. Between ES and ED, this embodiment considers a Rayleigh fading channel, where the SNR received by ES follows an exponential distribution with a default average value of 。This embodiment believes that the received SNR varies from sample to sample. To capture the impact of different channel conditions, this embodiment considers 100 random SNRs for each image during the training, validation, and testing phases. Additionally, this embodiment uses J = 100 random to calculate the expected cost of retransmission. This embodiment considers uniform quantization during the quantization process and quadrature phase shift keying (QPSK) during modulation.
[0063] This embodiment experimentally tests the system by training LSTM classifiers with different numbers of retransmissions during training, with a maximum of 5 retransmissions during training, and setting the maximum number of retransmissions to 7 during testing. The classification accuracy is shown in Table 3.
[0064] According to the results, it can be found that the accuracy increases by 5 - 6 percentage points when retransmitting once compared to not retransmitting, and the improvement effect is the most obvious. Finally, the classification accuracy of the model saturates at around 96%. For the training of the model, the more times of retransmission during training, the greater the cost. The model trained with two retransmissions can achieve the expected results. The convergence trend of the model trained with two retransmissions is as Figure 4 shown, and the cross - entropy loss finally converges to around 1.4. It should be noted that the number of retransmissions during testing can be greater than that during training, and the accuracy improves slightly, which fully demonstrates the generalization ability of the LSTM network as a decoder.
[0065] Table 3 Overall performance of the model
[0066] The decoding structure of a semantic retransmission system based on LSTM designed in this embodiment can effectively adapt to various semantic retransmission situations and exhibit stable and good picture classification performance. This embodiment uses the publicly available picture classification dataset MNIST to complete the ten - classification task of digits 0 - 9. The system only trains the case of two retransmissions. During testing, the classification accuracies without retransmission and with 1 - 7 retransmissions are 85.94%, 91.28%, 93.32%, 94.43%, 95.34%, 95.75%, 96.01%, 96.12% respectively.
[0067] This embodiment achieves the following technical effects through the above technical solutions: In this embodiment, an LSTM classifier is used to replace the traditional multi-decoder structure and unified single-decoder structure. While simplifying the decoder structure, it can effectively capture and retain the semantic features at different time steps during the transmission process, ensuring the complete transmission of semantic content. At the same time, it supports variable-length inputs, and the semantic features of each retransmission are sequentially input into each time step of the LSTM network, which is suitable for the semantic retransmission scenario. This embodiment improves the generalization ability, avoids the problem that the number of tests of the multi-decoder structure must be equal to the number of training times, and reduces the complexity of the model and the training complexity.
[0068] Exemplary device Based on the above embodiments, the present invention further provides a semantic communication retransmission system based on a long short-term memory network, including: An acquisition module, configured to acquire an image data set to be classified and processed; A feature extraction module, configured to extract semantic features from the image data set based on a feature extractor of a classification network; A classification module, configured to transmit or retransmit the extracted semantic features, and process and classify the noisy semantic features after transmission or retransmission based on the LSTM classifier of the classification network; An output module, configured to output an image classification result corresponding to the image data set.
[0069] This embodiment achieves the following technical effects through the above technical solutions: In this embodiment, an LSTM classifier is used to replace the traditional multi-decoder structure and unified single-decoder structure. While simplifying the decoder structure, it can effectively capture and retain the semantic features at different time steps during the transmission process, ensuring the complete transmission of semantic content. At the same time, it supports variable-length inputs, and the semantic features of each retransmission are sequentially input into each time step of the LSTM network, which is suitable for the semantic retransmission scenario. This embodiment improves the generalization ability, avoids the problem that the number of tests of the multi-decoder structure must be equal to the number of training times, and reduces the complexity of the model and the training complexity.
[0070] Based on the above embodiments, the present invention further provides a terminal, and its principle block diagram can be as Figure 5 shown.
[0071] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected through a system bus; wherein, the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect external devices; the display screen is used to display corresponding information; the communication module is used to communicate with a cloud server or other devices.
[0072] When the computer program is executed by a processor, it is used to implement the operations of the semantic communication retransmission method based on a long short-term memory network.
[0073] Those skilled in the art can understand that Figure 5 the principle block diagram shown in [the figure] is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0074] In one embodiment, a terminal is provided, which includes: a processor and a memory. The memory stores a semantic communication retransmission program based on a long short-term memory network. When the semantic communication retransmission program based on a long short-term memory network is executed by the processor, it is used to implement the operations of the semantic communication retransmission method based on a long short-term memory network as described above.
[0075] In one embodiment, a storage medium is provided, which stores a semantic communication retransmission program based on a long short-term memory network. When the semantic communication retransmission program based on a long short-term memory network is executed by the processor, it is used to implement the operations of the semantic communication retransmission method based on a long short-term memory network as described above.
[0076] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention may include non-volatile and volatile memories.
[0077] In summary, the present invention provides a semantic communication retransmission method, system, terminal, and storage medium based on a long short-term memory network, including: obtaining an image data set to be classified and processed; extracting semantic features from the image data set based on a feature extractor of a classification network; transmitting or retransmitting the extracted semantic features, and processing and classifying the noisy semantic features after transmission or retransmission based on an LSTM classifier of the classification network; outputting an image classification result corresponding to the image data set. The present invention uses an LSTM classifier to replace the traditional multi-decoder structure and unified single-decoder structure, which simplifies the decoder structure while effectively capturing and retaining the semantic features at different time steps during the transmission process, ensuring the complete transmission of semantic content, and improving the generalization ability of the model.
[0078] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or changes can be made according to the above description, and all such improvements and changes shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A semantic communication retransmission method based on long short-term memory network, characterized in that: include: Obtain an image data set to be classified; A feature extractor based on a classification network extracts semantic features from the image dataset; The extracted semantic features are transmitted or retransmitted, and the semantic features with noise after transmission or retransmission are processed and classified based on the LSTM classifier of the classification network; Output the image classification result corresponding to the image dataset.
2. The semantic communication retransmission method based on long short-term memory network according to claim 1 is characterized in that: The classification network-based feature extractor extracts semantic features from the image dataset, including: Leveraging feature extractors deployed on edge devices and hyperparameters extracting semantic features from images in the image dataset; The extracted semantic features are quantized by the analog-to-digital conversion module and Modulation is performed to obtain modulated semantic features.
3. The semantic communication retransmission method based on long short-term memory network according to claim 2 is characterized in that: The transmitting or retransmitting the extracted semantic features includes: Transmitting the modulated semantic features to an LSTM classifier deployed on an edge server through a noisy and fading wireless channel; Or based on the noisy and fading wireless channel, the modulated semantic features are retransmitted using a retransmission mechanism of the retransmission modulation feature, new Gaussian white noise is added, and the retransmitted signal is transmitted to the LSTM classifier.
4. The semantic communication retransmission method based on long short-term memory network according to claim 1 is characterized in that: The LSTM classifier based on the classification network processes and classifies the semantic features with noise after transmission or retransmission, including: Performing scaling processing on the signal corresponding to the semantic feature with noise; After scaling, the signal Demodulation and DAC digital-to-analog conversion to obtain estimated semantic features; The estimated semantic features are input into the hyperparameters as LSTM-based classifier The image classification results are obtained.
5. The semantic communication retransmission method based on long short-term memory network according to claim 1 is characterized in that: The LSTM classifier based on the classification network processes and classifies the semantic features with noise after transmission or retransmission, and also includes: The semantic features after retransmission are Demodulation and DAC digital-to-analog conversion are performed to obtain semantic features of the noisy data after retransmission. The noisy semantic features after retransmission are input into the LSTM classifier and decoded together with the features before retransmission to obtain the image classification result with retransmission.
6. The semantic communication retransmission method based on long short-term memory network according to claim 1 is characterized in that: The method further comprises: According to the preset uniform quantization parameter, the range of features extracted from the input image is divided into equal intervals; based on The classification network is trained using the function and the cross entropy loss function.
7. The semantic communication retransmission method based on long short-term memory network according to claim 6 is characterized in that: The basis Function and cross entropy loss function are used to train the classification network, including: Randomly from the training data set Select training batches; Extract corresponding semantic features for the selected dataset , and obtain the quantitative results ; Calculate the noisy signal when transmitting and retransmitting the same batch of images , where the channel signal-to-noise ratio is the same; for different pictures, when transmitting and retransmitting, a random signal-to-noise ratio with the same mean is generated and noise , calculate the signal with noise ; Get estimated features ; Get classification results ; Calculate the training loss based on the cross entropy loss function , update the parameters through back propagation , , until the parameter , convergence.
8. A semantic communication retransmission system based on long short-term memory network, characterized in that: include: An acquisition module is used to acquire an image data set to be classified and processed; A feature extraction module, configured to extract semantic features from the image data set based on a feature extractor of a classification network; A classification module, used for transmitting or retransmitting the extracted semantic features, and processing and classifying the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network; The output module is used to output the image classification result corresponding to the image data set.
9. A terminal, characterized in that: include: A processor and a memory, wherein the memory stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on a long short-term memory network is executed by the processor, it is used to implement the operation of the semantic communication retransmission method based on a long short-term memory network as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on a long short-term memory network is executed by a processor, it is used to implement the operation of the semantic communication retransmission method based on a long short-term memory network as described in any one of claims 1-7.
Citation Information
Patent Citations
Two-stage encrypted traffic classification method for programmable switch
CN119011493A
Archive key element identification and classification method based on improved LSTM semantic modeling
CN119026601A
Self-adaptive semantic communication method for machine vision task
CN119544143A
Machine learning data representations, architectures, and systems that intrinsically encode and represent benefit, harm, and emotion to optimize learning
US20200104726A1
Information transmitting method and apparatus based on intent-driven network, electronic device, and medium
WO2023138238A1