A semantic communication retransmission method, system, terminal and storage medium based on long short-term memory network

By using an LSTM classifier to replace the traditional decoder structure, the semantic communication retransmission process is simplified and optimized, which solves the problems of high decoder complexity and weak generalization ability in the existing technology, and realizes efficient semantic content transmission and flexible retransmission processing.

CN120050004BActive Publication Date: 2025-09-05SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510527724.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-05
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing semantic retransmission decoders have complex structures, high costs, are difficult to train, and have weak generalization capabilities, especially when faced with constantly changing channel conditions and transmission requirements.

Method used

A semantic communication retransmission method based on the long short-term memory network (LSTM) is adopted. The traditional multi-decoder and unified single decoder structure are replaced by the LSTM classifier. The memory unit and gating mechanism of LSTM are used to process noisy semantic features, support dynamic length input and perform classification.

Benefits of technology

The decoder structure is simplified, the complete transmission capability of semantic content is improved, the generalization capability of the model is enhanced, the system complexity and training complexity are reduced, and it can adapt to different numbers of retransmissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050004B_ABST
    Figure CN120050004B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of information and communication technologies and discloses a semantic communication retransmission method, system, terminal, and storage medium based on a long short-term memory network. The method comprises: obtaining an image dataset to be classified; extracting semantic features from the image dataset using a feature extractor based on a classification network; transmitting or retransmitting the extracted semantic features, and processing and classifying the noisy semantic features after transmission or retransmission using an LSTM classifier based on the classification network; and outputting an image classification result corresponding to the image dataset. The present invention uses an LSTM classifier to replace traditional multi-decoder structures and unified single-decoder structures, simplifying the decoder structure while effectively capturing and retaining semantic features at different time steps during the transmission process, ensuring the complete transmission of semantic content and improving the model's generalization capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information communication technology, and in particular to a semantic communication retransmission method, system, terminal and storage medium based on a long short-term memory network. Background Art

[0002] With the rapid development of modern communication technologies, traditional communication methods are facing increasingly severe challenges. Based on Shannon's theory, traditional communication systems simplify the communication process to bit-level transmission and recovery of information. Their core goal is to ensure symbol-level accuracy in channel transmission, while neglecting the semantic value of the information itself. This paradigm has demonstrated significant effectiveness in traditional application scenarios. However, with the explosive growth of emerging intelligent applications such as the Internet of Things, autonomous driving, virtual reality, and augmented reality, its transmission efficiency bottleneck has become increasingly prominent. Against this backdrop, the vision of the intelligent interconnection of all things, proposed by the sixth-generation wireless communication system (6G), places dual demands on communication technology for innovation: on the one hand, the urgent need for deep integration of communication systems and artificial intelligence (AI) technologies; on the other, the need for more efficient intelligent encoding and decoding mechanisms. As a core technology direction of the 6G system, semantic communication (SC), empowered by deep neural networks (DNNs), has pioneered an innovative path to break through traditional efficiency bottlenecks. Compared to simply pursuing the fidelity of symbolic transmission, semantic communication systems focus on the intelligent extraction and encoding of task-relevant semantic information, improving transmission efficiency by eliminating redundant information. With the help of DNNs, semantic communication systems can better understand environmental information and given tasks. Numerous studies have demonstrated that this new communication paradigm exhibits significant advantages in data-intensive, latency-sensitive intelligent scenarios such as image classification, speech recognition, and video transmission.

[0003] To ensure efficient semantic transmission over noisy wireless channels, systematic optimization is required across two dimensions: encoding and decoding mechanisms and transmission reliability. First, joint source and channel coding (JSCC) is designed to enable end-to-end DNN model training, thereby achieving system-level optimization of the encoding process. To simplify analysis and training, semantic feature vectors are typically represented as continuous real values ​​and transmitted over the channel, while the receiver receives a noisy version of the continuous feature vector. Modern digital communication systems quantize, encode, and modulate feature vectors into QAM (Quadrature Amplitude Modulation) signals for transmission. Simultaneously, the receiver must perform digital-to-analog conversion to demodulate the signal and reconstruct the feature vectors. The distortion introduced by quantization and modulation can severely damage the semantic structure of feature vectors, resulting in a degradation in intelligent task performance. Second, automatic repeat request (ARQ) is widely used in digital communication systems to achieve reliable data transmission over noisy wireless channels. ARQ ensures communication reliability by retransmitting extracted feature vectors, thereby enhancing the performance of intelligent tasks.

[0004] When transmitting data under extremely poor channel conditions or other harsh communication environments, the desired inference results may not be achieved without retransmission or with a single retransmission. In this case, multiple retransmissions may be necessary to ensure transmission reliability. Currently, there are two main decoder architectures for semantic retransmission in communication systems: multi-decoders and unified single decoders. In the multi-decoder architecture, multiple decoders need to be trained and maintained, increasing system complexity and computational resources. Furthermore, multiple decoders may learn duplicate parameters, especially when task characteristics are similar, thus reducing model efficiency. In the unified single decoder architecture, since all transmitted information is processed by the same decoder, the performance of a multi-decoder solution may not be achieved in certain situations (such as when the number of retransmissions is predefined). Furthermore, the single decoder needs to be able to process encoded information of varying lengths, which increases the difficulty of training.

[0005] In summary, existing decoders for semantic retransmission have the following technical drawbacks:

[0006] 1) Maintaining a complex single decoder or multiple specialized decoders will increase system complexity and implementation cost;

[0007] 2) Training is difficult. Existing decoder structures need to be trained to adapt to various transmission conditions, especially when handling different numbers of retransmissions. This may involve complex training and tuning.

[0008] 3) Weak generalization ability, which is insufficient in handling new transmission situations, especially when facing changing channel conditions and transmission requirements.

[0009] Therefore, the existing technology needs to be improved. Summary of the Invention

[0010] The technical problem to be solved by the present invention is that, in response to the defects of the existing technology, the present invention provides a semantic communication retransmission method, system, terminal and storage medium based on long short-term memory network to solve the problems of system complexity and low generalization ability of existing decoding schemes for semantic retransmission.

[0011] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0012] In a first aspect, the present invention provides a semantic communication retransmission method based on a long short-term memory network, comprising:

[0013] Obtain the image dataset to be classified;

[0014] A feature extractor based on a classification network extracts semantic features from the image dataset;

[0015] Transmitting or retransmitting the extracted semantic features, and processing and classifying the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network;

[0016] Output the image classification result corresponding to the image dataset.

[0017] In one implementation, the classification network-based feature extractor extracts semantic features from the image dataset, including:

[0018] Leveraging feature extractors deployed on edge devices and hyperparameters extracting semantic features from images in the image dataset;

[0019] The extracted semantic features are quantized by the analog-to-digital conversion module and Modulation is performed to obtain modulated semantic features.

[0020] In one implementation, transmitting or retransmitting the extracted semantic features includes:

[0021] Transmitting the modulated semantic features to an LSTM classifier deployed on an edge server through a noisy and fading wireless channel;

[0022] Or based on the noisy and fading wireless channel, the modulated semantic features are retransmitted using a retransmission mechanism of the retransmission modulation feature, new Gaussian white noise is added, and the retransmitted signal is transmitted to the LSTM classifier.

[0023] In one implementation, the LSTM classifier based on the classification network processes and classifies the semantic features with noise after transmission or retransmission, including:

[0024] Performing scaling processing on the signal corresponding to the semantic feature with noise;

[0025] The scaled signal is processed Demodulation and DAC digital-to-analog conversion to obtain estimated semantic features;

[0026] The estimated semantic features are input into the hyperparameters LSTM-based classifier , and obtain the image classification results.

[0027] In one implementation, the LSTM classifier based on the classification network processes and classifies the semantic features with noise after transmission or retransmission, further comprising:

[0028] The semantic features after retransmission are Demodulation and DAC digital-to-analog conversion are performed to obtain semantic features of the retransmitted noisy data.

[0029] The noisy semantic features after retransmission are input into the LSTM classifier and decoded together with the features before retransmission to obtain the image classification result with retransmission.

[0030] In one implementation, the method further includes:

[0031] According to the preset uniform quantization parameter, the range of features extracted from the input image is divided into equal intervals;

[0032] based on The classification network is trained using the function and the cross entropy loss function.

[0033] In one implementation, the The classification network is trained using the function and the cross entropy loss function, including:

[0034] Randomly from the training data set Select training batches;

[0035] Extract corresponding semantic features for the selected dataset , and obtain the quantitative results ;

[0036] Calculate the noisy signal when transmitting and retransmitting the same batch of images , where the channel signal-to-noise ratio is the same; for different pictures, when they are transmitted and retransmitted, a random signal-to-noise ratio with the same mean is generated and noise , calculate the signal with noise ;

[0037] Get estimated features ;

[0038] Get classification results ;

[0039] Calculate the training loss based on the cross entropy loss function , update the parameters through back propagation , , until the parameter , convergence.

[0040] In a second aspect, the present invention provides a semantic communication retransmission system based on a long short-term memory network, comprising:

[0041] An acquisition module is used to obtain an image dataset to be classified;

[0042] A feature extraction module, configured to extract semantic features from the image dataset using a feature extractor based on a classification network;

[0043] A classification module, configured to transmit or retransmit the extracted semantic features, and process and classify the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network;

[0044] The output module is used to output the image classification results corresponding to the image dataset.

[0045] In a third aspect, the present invention provides a terminal comprising: a processor and a memory, wherein the memory stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on a long short-term memory network is executed by the processor, it is used to implement the operation of the semantic communication retransmission method based on a long short-term memory network as described in the first aspect.

[0046] In a fourth aspect, the present invention also provides a medium, which is a computer-readable storage medium, and the medium stores a semantic communication retransmission program based on a long short-term memory network. When the semantic communication retransmission program based on a long short-term memory network is executed by a processor, it is used to implement the operation of the semantic communication retransmission method based on a long short-term memory network as described in the first aspect.

[0047] The present invention adopts the above technical solution to achieve the following effects:

[0048] The present invention uses an LSTM classifier to replace the traditional multi-decoder structure and unified single decoder structure, simplifying the decoder structure while effectively capturing and retaining the semantic features of different time steps during the transmission process, ensuring the complete transmission of semantic content. At the same time, it supports dynamic length input and sequentially inputs the semantic features of each retransmission into each time step of the LSTM network, which is suitable for semantic retransmission scenarios. The present invention improves generalization capability, avoids the problem that the number of tests for multi-decoder structures must be equal to the number of training times, and reduces the complexity of the model and the complexity of training. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0050] Figure 1 It is a flow chart of the semantic communication retransmission method based on long short-term memory network in the present invention.

[0051] Figure 2 It is a schematic diagram of the LSTM unit structure in the present invention.

[0052] Figure 3 It is a schematic diagram of the classification network system model in the present invention.

[0053] Figure 4 It is a schematic diagram of the convergence trend of the CLFNet model in the present invention.

[0054] Figure 5 It is a functional principle diagram of a terminal in one implementation of the present invention.

[0055] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0057] Exemplary Methods

[0058] Currently, there are two main decoder architectures for semantic retransmission in communication systems: multi-decoders and unified single decoders. In a multi-decoder architecture, N different decoders are responsible for different transmission and retransmission tasks. After each received message, the corresponding decoder concatenates it with the previously transmitted message and decodes the concatenated message to complete the corresponding task. The information here can be a received bitstream or a feature vector. The advantage of this decoder architecture is that with each retransmission, the decoder can leverage previously transmitted information (incremental knowledge) to jointly perform task reasoning, thereby gradually improving transmission performance. However, this architecture also has some problems. First, it requires training and maintaining multiple decoders, which increases system complexity and computational resources. Second, multiple decoders may learn duplicate parameters, especially when the task characteristics are similar, which reduces model efficiency. In the second architecture, a unified single decoder, a single decoder is used to process all transmitted messages, regardless of how many times the message is transmitted. This decoder is trained to handle encoded messages of varying sizes, and if the number of transmissions does not reach a preset maximum, the extra dimensions are automatically padded with zero vectors. During training, the system learns to extract valid information from the retransmitted semantic knowledge, even with randomly selected transmission times. The unified single decoder architecture requires only a single decoder, simplifying system design. Compared to multiple decoders, a single decoder solution saves storage and computational resources. Furthermore, it offers greater flexibility, handling encoded messages of varying lengths through zero-padding, enabling the system to adapt to diverse transmission scenarios. Because all transmitted messages are processed by the same decoder, the unified single decoder architecture may not achieve the performance of a multiple-decoder solution in certain situations (e.g., when the number of retransmissions is predefined). Furthermore, the single decoder needs to be able to handle encoded messages of varying lengths, which increases the difficulty of training.

[0059] In summary, existing decoders for semantic retransmission have the following technical drawbacks:

[0060] 1) Maintaining a complex single decoder or multiple specialized decoders will increase system complexity and implementation cost;

[0061] 2) Training is difficult. Existing decoder structures need to be trained to adapt to various transmission conditions, especially when handling different numbers of retransmissions. This may involve complex training and tuning.

[0062] 3) Weak generalization ability, which is insufficient in handling new transmission situations, especially when facing changing channel conditions and transmission requirements.

[0063] In response to the above technical problems, an embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network. The method mainly obtains an image dataset to be classified and processed; then, a feature extractor based on a classification network extracts semantic features from the image dataset; finally, the extracted semantic features are transmitted or retransmitted, and an LSTM classifier based on the classification network processes and classifies the noisy semantic features after transmission or retransmission; and the image classification results corresponding to the image dataset are output. The embodiment of the present invention uses an LSTM classifier to replace the traditional multi-decoder structure and unified single decoder structure, simplifying the decoder structure while effectively capturing and retaining the semantic features at different time steps during the transmission process, ensuring the complete transmission of the semantic content and improving the generalization ability of the model.

[0064] like Figure 1 As shown, an embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network, comprising the following steps:

[0065] Step S100, obtaining an image data set to be classified;

[0066] Step S200: A feature extractor based on a classification network extracts semantic features from the image dataset.

[0067] In this embodiment, a decoding structure of a semantic retransmission system based on a long short-term memory (LSTM) network is designed. LSTM, as a special recurrent neural network (RNN), is specifically designed to learn and memorize long-term dependencies in sequential data. The LSTM-based semantic decoder uses its memory unit and three gates (forget gate, input gate, and output gate) to control which information needs to be remembered, updated, or forgotten. The interaction of these three gates enables LSTM to flexibly adjust memory content according to input data and tasks, retain important information, and achieve higher classification accuracy.

[0068] In this embodiment, the LSTM-based decoder structure has several important parameters: input_size, hidden_size, num_layers, num_classes, and seq_length, which respectively represent the dimension of the LSTM input feature, the dimension of the hidden layer, the number of stacked layers, the number of output categories, and the length of the input sequence. Among them, seq_length does not need to be specified, that is, the LSTM decoder structure designed in this embodiment can process variable-length sequences, which greatly reduces the training cost of the system and improves the flexibility of retransmission and subsequent optimization control.

[0069] In the context of the semantic retransmission system designed in this embodiment, the features of the first transmission and subsequent retransmissions obtained at the receiving end are used as inputs to the LSTM network at each time step. Their dimensions, i.e., the same input_size, are identical. After passing through the memory unit and three gates within the LSTM, a fully connected layer finally generates the classification result. The LSTM-based decoder structure designed in this embodiment can be used to process variable-length sequences, resolving the drawback of traditional decoders that require training for a specific number of transmissions. This significantly reduces training complexity, improves flexibility in processing retransmitted data, and improves the generalization of classification capabilities.

[0070] LSTM is a special type of RNN designed to handle long-term dependencies that traditional RNNs struggle to handle. While traditional RNNs can capture correlations within sequences, they are prone to vanishing or exploding gradients when processing longer sequences, making it difficult for the model to remember earlier inputs. LSTM, on the other hand, uses a special "gating" mechanism to control the flow of information, effectively handling long-term dependencies.

[0071] Table 1 Variable annotation table

[0072]

[0073] An LSTM unit consists of five main components: a cell state, a hidden state, a forget gate, an input gate, and an output gate. These gates are controlled by trainable neural network layers, allowing the LSTM unit to selectively remember or discard information, achieving long-term memory.

[0074] (1) Cell state: The cell state is the main thread running through the LSTM unit, transmitting information from previous and next time steps. LSTM adjusts the cell state through the forget gate and input gate, allowing the model to "remember" or "forget" certain information:

[0075] (1);

[0076] (2) Hidden state: The hidden state is the output of LSTM and also serves as the input of the next time step. The hidden state not only affects the output of the current time step, but also provides contextual information for the next time step:

[0077] (2);

[0078] (3) Forget gate: The forget gate uses a sigmoid function (an activation function) to decide which information to discard. Based on the current input and the hidden state of the previous time step, it generates a vector with a value between 0 and 1, where 0 means "completely forgotten" and 1 means "completely retained":

[0079] (3);

[0080] (4) Input gate: The input gate controls how much new information is written into the cell state. The input gate consists of a sigmoid function layer and a tanh function layer (an activation function layer). The sigmoid function layer determines the proportion of writing, and the tanh function layer provides candidate information:

[0081] (4);

[0082] (5) Output gate: The output gate determines which information will be output from the current cell state. It filters the information in the cell state through sigmoid and tanh operations and outputs it to the next time step.

[0083] (5);

[0084] The structure diagram of LSTM unit is as follows Figure 2 shown.

[0085] Specifically, in one implementation of this embodiment, step S200 includes the following steps:

[0086] Step S201: Using the feature extractor deployed on the edge device and hyperparameters extracting semantic features from images in the image dataset;

[0087] Step S202: The extracted semantic features are quantized by the analog-to-digital conversion module and converted by the digital modulator. Modulation is performed to obtain modulated semantic features.

[0088] In this embodiment, a semantic retransmission system for device edge collaborative reasoning is considered, such as Figure 3As shown in Figure 1, the system consists of edge devices (EDs) and edge servers (ESs), which collaborate to perform shared AI tasks. Here, we consider the task of classifying images into ten categories (e.g., datasets like CIFAR-10, MNIST, and Nature 10). The MNIST dataset is used here, where each image consists of 28×28 pixels of handwritten digits from 0 to 9, for this task). This system uses a DNN-based classification network consisting of two parts: a feature extractor deployed on the ED and an LSTM-based classifier deployed on the ES. The ED transmits the extracted semantic features to the ES via a digital communication link. The LSTM-based classifier on the ES processes and classifies the received noisy semantic features. The following describes the system's inference process.

[0089] In this embodiment, the ED processed image is represented as ,in, is an image dataset. ED uses a feature extractor based on DNN and hyperparameters From the input image Extract semantic features from , expressed as:

[0090] (6);

[0091] in, Quantization and digital modulator after analog-to-digital conversion (ADC) After modulation, the semantic features obtained are:

[0092] (7);

[0093] Transmitted from ED to ES through a noisy wireless channel.

[0094] In the above quantization and digital modulator process, the specific process is: first, Normalize to the range of [0,1], the normalization formula is: (x - x_min) / (x_max - x_min), the normalized value is multiplied by , To quantize the number of bits, map it to an integer range. If in training mode, use the sin_round function (a SQL base function) for quantization. If in non-training mode, use the torch.round function (a function that rounds each element in the input tensor to the nearest integer) for rounding quantization.

[0095] The quantized output is then converted into binary numbers for transmission, followed by QPSK (Quadrature Phase Shift Keying, a digital modulation technique). The input binary data stream is first grouped into pairs of two bits, the symbol index is calculated, mapped to a constellation diagram, and finally the complex symbols are output.

[0096] like Figure 1 As shown, an embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network, comprising the following steps:

[0097] Step S300, transmitting or retransmitting the extracted semantic features, and processing and classifying the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network;

[0098] Step S400: outputting the image classification result corresponding to the image dataset.

[0099] In this embodiment, an LSTM classifier different from the traditional LSTM structure is used to process and classify semantic features with noise after transmission or retransmission; the structure of the LSTM network in this embodiment is similar to the traditional LSTM structure, and the network parameters are adjusted to make it more suitable for classification tasks. Finally, a fully connected layer is added to the last layer inside the LSTM to obtain the classification results.

[0100] In this embodiment, data is input according to the format given by the LSTM network, processed by the internal network structure, and passed through the last fully connected layer to output unnormalized probabilities (logits, which refer to the original predicted values ​​output by the last layer of the model), thereby obtaining the image classification results corresponding to the image dataset.

[0101] Specifically, in one implementation of this embodiment, step S300 includes the following steps:

[0102] Step S301: transmitting the modulated semantic features to an LSTM classifier deployed on an edge server through a noisy and fading wireless channel;

[0103] Step S302, or based on the noisy and fading wireless channel, retransmit the modulated semantic features using the retransmission mechanism of the retransmission modulation feature, add new Gaussian white noise, and transmit the retransmitted signal to the LSTM classifier.

[0104] Based on the above step S301, in one implementation of this embodiment, step S300 includes the following steps:

[0105] Step S3011, performing scaling processing on the signal corresponding to the semantic feature with noise;

[0106] Step S3012: perform the scaling on the signal. Demodulation and DAC digital-to-analog conversion to obtain estimated semantic features;

[0107] Step S3013: input the estimated semantic features into the hyperparameter LSTM-based classifier , and obtain the image classification results.

[0108] In this embodiment, the signal received at the ES can be expressed as:

[0109] (8);

[0110] represents the channel fading coefficient, is additive white Gaussian noise. Consider a slow fading scenario where and signal-to-noise ratio It remains constant in each classification task, but varies independently between tasks. Assume that the channel information of ES is perfect. After receiving the signal, ES performs signal scaling:

[0111] (9);

[0112] pass Demodulation and DAC digital-to-analog conversion , get the estimated semantic features:

[0113] (10);

[0114] In this embodiment Demodulation is the inverse of modulation. Specifically, the received signal is first calculated from all possible symbols in the constellation, the nearest constellation point is found, and the signal is converted to binary bits, which are then restored to a binary bit stream. The binary signal is then restored to its original quantized form for digital-to-analog conversion.

[0115] Digital-to-analog conversion is the inverse of quantization. The specific process is to restore the demodulated signal to the analog signal, including dividing by the quantization accuracy and denormalizing. The specific formula is x_analog = (x_max - x_min) / (2^q - 1) x + x_min.

[0116] After obtaining the estimated semantic features, input them into the hyperparameters LSTM-based classifier To obtain the classification results:

[0117] (11);

[0118] In order to combat channel fading and improve mission performance, the present invention considers a modulation feature that allows multiple ED retransmissions. The retransmission mechanism is derived from Automatic Repeat-reQuest (ARQ), an error-correction protocol used in the data link and transport layers of the OSI model. It provides reliable information transmission over unreliable services. If the sender doesn't receive an acknowledgment within a certain period of time after transmission, it typically retransmits the message. The ARQ protocol is used in communication systems to ensure feature reliability.

[0119] Based on the above step S302, in one implementation of this embodiment, step S300 includes the following steps:

[0120] Step S3021, the semantic features after retransmission are processed. Demodulation and DAC digital-to-analog conversion are performed to obtain semantic features of the retransmitted noisy data.

[0121] Step S3022: input the noisy semantic features after retransmission into the LSTM classifier. and decoded together with the features before retransmission to obtain the image classification result with retransmission.

[0122] In this embodiment, the ED retransmission modulation characteristic The retransmission mechanism of , the retransmission signal is expressed as:

[0123] (12);

[0124] Adding new Gaussian white noise Upon receiving After that, ES is applied as in formula (10) The same demodulation and DAC process is used to obtain the estimation of the retransmitted semantic features. ,Right now ,in Then, ES will Input to LSTM-based classifier , we get the prediction with retransmission:

[0125] (13);

[0126] In this embodiment, the overall framework of the classification network is as follows: Figure 3As shown in the figure, the classification network processes data by inputting the data into the LSTM in the shape of [batch_size, seq_length, input_size]. The LSTM uses three gates and memory cells to perform transformation processing at each time step through formulas (1)-(5). Finally, it passes through the fully connected layer and is mapped to the category space to obtain the category probability.

[0127] This embodiment adopts a semantic feature selection and weighting mechanism. Due to its memory unit and gating mechanism, the LSTM network can memorize important information for a long time and filter out irrelevant information. Moreover, the generalization ability for variable-length input is improved. The LSTM network can flexibly handle different numbers of retransmissions without redundant targeted training.

[0128] An embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network, which further includes the following steps before steps S200 to S300:

[0129] Step S001: Divide the range of features extracted from the input image into equal intervals.

[0130] An embodiment of the present invention provides a semantic communication retransmission method based on a long short-term memory network. During the implementation of the entire method, the method further includes:

[0131] Step S002, based on The classification network is trained using the function and the cross entropy loss function.

[0132] Specifically, in one implementation of this embodiment, step S002 includes the following steps:

[0133] Step S002a, randomly select Select training batches;

[0134] Step S002b: extract the corresponding semantic features for the selected data set , and obtain the quantitative results ;

[0135] Step S002c, for the same batch of pictures, when transmitting and retransmitting, calculate the signal with noise , where the channel signal-to-noise ratio is the same; for different pictures, when they are transmitted and retransmitted, a random signal-to-noise ratio with the same mean is generated and noise , calculate the signal with noise ;

[0136] Step S002d, obtain estimated features ;

[0137] Step S002e, obtain classification results ;

[0138] Step S002f, calculating the training loss based on the cross entropy loss function , update the parameters through back propagation , , until the parameter , convergence.

[0139] In this embodiment, the digital semantic communication training process (i.e., the training process of the classification network) is:

[0140] For the above classification network, in this embodiment, the classification network is trained end-to-end to jointly optimize the hyperparameters and Since the feature extractor and classifier in the classification network are connected through a digital communication link, the extracted information is quantized and modulated before transmission. Consider uniform quantization, where the input range is divided into equal intervals. is the resolution of the ADC, i.e., the quantization bit. In this embodiment, the default quantization bit is 4. By uniform quantization, the real value input is rounded to the nearest integer value and then expressed as In order to alleviate the gradient explosion effect, a proxy function is introduced in this embodiment:

[0141] (14);

[0142] in, is the input to be rounded, Design parameters are positive integers. Provides a smooth differentiable approximation of the rounding function, which makes the gradient propagation stable during training, resulting in better convergence and better model performance. This example is only used in the model training state to approximate quantization, and the torch.round() function is used for actual quantization during model inference.

[0143] Specifically, in this embodiment, Expressed as The rounded output of . Refers to the features extracted by the feature extractor The result of quantization in the training phase, that is, after the quantization function The output after that is Figure 3 in In the model training process of the classification network, the modulation process is bypassed and is transmitted to ES. Therefore, the signal received at ES is It is worth noting that in the actual reasoning process, discrete features is represented as bit binary code, which is then modulated into a complex signal before transmission .

[0144] set up is the training image set. In this embodiment, the cross entropy loss function used to train the classification network is:

[0145] (15);

[0146] in, is the number of image categories, is an image Is it the first The labels of the categories, It is Secondary retransmission classifier The prediction output of . In this embodiment, the algorithm flow of the classification network is summarized as follows:

[0147]

[0148] Table 2 Network structure of CLFNet

[0149]

[0150] This embodiment uses numerical simulation to evaluate the system performance, considering a digital classification task and using MNIST as the dataset. The MNIST dataset consists of 70,000 images with 28×28 grayscale pixels. The dataset is divided into three parts: 48,000 training samples, 12,000 validation samples, and 10,000 test samples. The network structure of the classification network is shown in Table 2. In particular, in the classification network, this embodiment A tanh activation function is introduced in the last fully connected layer of to limit the output value to the range of [−1, 1]. Between ES and ED, this embodiment considers a Rayleigh fading channel, where the SNR received by ES follows an exponential distribution with a default average value of This embodiment considers that the received SNR varies from sample to sample. In order to capture the impact of different channel conditions, this embodiment considers 100 random SNRs for each image in the training, validation, and testing phases. In addition, this embodiment uses J = 100 random This embodiment considers uniform quantization in the quantization process and Quadrature Phase Shift Keying (QPSK) in the modulation process.

[0151] This embodiment experimentally tests the system to train LSTM classifiers with different retransmission times during training. The maximum number of retransmissions during training is 5 times, and the maximum number of retransmissions during testing is set to 7 times. The classification accuracy is shown in Table 3.

[0152] The results show that retransmission once increases the accuracy by 5-6 percentage points compared to no retransmission, which is the most significant improvement. Finally, the classification accuracy of the model tends to saturate at around 96%. For model training, the more retransmissions the training is, the greater the overhead is. The model with two retransmissions can achieve the expected results. The convergence trend of the model with two retransmissions is as follows: Figure 4 As shown in Figure 2, the cross entropy loss eventually converges to around 1.4. Note that the number of retransmissions during testing can be greater than the number during training, and the accuracy is slightly improved, which fully demonstrates the generalization ability of the LSTM network as a decoder.

[0153] Table 3 Overall performance of the model

[0154]

[0155] The decoding structure of the semantic retransmission system designed in this embodiment, based on an LSTM, can effectively adapt to various semantic retransmission scenarios and demonstrate stable and excellent image classification performance. This embodiment uses the public image classification dataset MNIST to complete the task of classifying the digits 0-9 into ten categories. The system is trained only with two retransmissions. During testing, the classification accuracy with no retransmissions and with one to seven retransmissions was 85.94%, 91.28%, 93.32%, 94.43%, 95.34%, 95.75%, 96.01%, and 96.12%, respectively.

[0156] This embodiment achieves the following technical effects through the above technical solution:

[0157] This embodiment uses an LSTM classifier to replace the traditional multi-decoder structure and the unified single decoder structure, which simplifies the decoder structure and can effectively capture and retain the semantic features of different time steps during the transmission process, ensuring the complete transmission of the semantic content. At the same time, it supports dynamic length input and inputs the semantic features of each retransmission into each time step of the LSTM network in sequence, which is suitable for semantic retransmission scenarios. This embodiment improves the generalization ability, avoids the problem that the number of tests of the multi-decoder structure must be equal to the number of training times, and reduces the complexity of the model and the complexity of training.

[0158] Exemplary devices

[0159] Based on the above embodiments, the present invention further provides a semantic communication retransmission system based on a long short-term memory network, comprising:

[0160] An acquisition module is used to obtain an image dataset to be classified;

[0161] A feature extraction module, configured to extract semantic features from the image dataset using a feature extractor based on a classification network;

[0162] A classification module, configured to transmit or retransmit the extracted semantic features, and process and classify the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network;

[0163] The output module is used to output the image classification results corresponding to the image dataset.

[0164] This embodiment achieves the following technical effects through the above technical solution:

[0165] This embodiment uses an LSTM classifier to replace the traditional multi-decoder structure and the unified single decoder structure, which simplifies the decoder structure and can effectively capture and retain the semantic features of different time steps during the transmission process, ensuring the complete transmission of the semantic content. At the same time, it supports dynamic length input and inputs the semantic features of each retransmission into each time step of the LSTM network in sequence, which is suitable for semantic retransmission scenarios. This embodiment improves the generalization ability, avoids the problem that the number of tests of the multi-decoder structure must be equal to the number of training times, and reduces the complexity of the model and the complexity of training.

[0166] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be shown as follows: Figure 5 shown.

[0167] The terminal includes: a processor, memory, interface, display screen and communication module connected via a system bus; wherein the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and computer programs; the internal memory provides an environment for the operation of the operating system and computer programs in the storage medium; the interface is used to connect to external devices; the display screen is used to display corresponding information; and the communication module is used to communicate with a cloud server or other devices.

[0168] When the computer program is executed by a processor, it is used to implement the operation of the semantic communication retransmission method based on the long short-term memory network.

[0169] It will be understood by those skilled in the art that Figure 5The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0170] In one embodiment, a terminal is provided, which includes: a processor and a memory, wherein the memory stores a semantic communication retransmission program based on a long short-term memory network, and the semantic communication retransmission program based on a long short-term memory network is used to implement the operation of the above-mentioned semantic communication retransmission method based on a long short-term memory network when executed by the processor.

[0171] In one embodiment, a storage medium is provided, wherein the storage medium stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on a long short-term memory network is executed by a processor, it is used to implement the operation of the above-mentioned semantic communication retransmission method based on a long short-term memory network.

[0172] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When executed, the computer program can include the processes in the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include both non-volatile and volatile memory.

[0173] In summary, the present invention provides a semantic communication retransmission method, system, terminal, and storage medium based on a long short-term memory network, comprising: obtaining an image dataset to be classified; extracting semantic features from the image dataset using a feature extractor based on a classification network; transmitting or retransmitting the extracted semantic features, and processing and classifying the noisy semantic features after transmission or retransmission using an LSTM classifier based on the classification network; and outputting an image classification result corresponding to the image dataset. The present invention uses an LSTM classifier to replace the traditional multi-decoder structure and unified single decoder structure, simplifying the decoder structure while effectively capturing and retaining semantic features at different time steps during the transmission process, ensuring the complete transmission of semantic content and improving the generalization ability of the model.

[0174] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A semantic communication retransmission method based on long short-term memory network, characterized in that: include: Obtain the image dataset to be classified; A feature extractor based on a classification network extracts semantic features from the image dataset; Transmitting or retransmitting the extracted semantic features, and processing and classifying the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network; Outputting the image classification result corresponding to the image dataset; The LSTM classifier based on the classification network processes and classifies the semantic features with noise after transmission or retransmission, including: Performing scaling processing on the signal corresponding to the semantic feature with noise; The scaled signal is demodulated and subjected to DAC digital-to-analog conversion to obtain estimated semantic features; the specific process of the DAC digital-to-analog conversion is: restoring the demodulated signal to an analog signal, including operations of dividing by quantization accuracy and denormalization; The estimated semantic features are input into the hyperparameters Based on the LSTM classifier, the image classification results are obtained; The LSTM classifier based on the classification network processes and classifies the semantic features with noise after transmission or retransmission, and further includes: Demodulate and convert the retransmitted semantic features into digital-to-analog (DAC) to obtain the noisy semantic features. The noisy semantic features after retransmission are input into the LSTM classifier and decoded together with the features before retransmission to obtain the image classification result with retransmission.

2. The semantic communication retransmission method based on long short-term memory network according to claim 1 is characterized in that The classification network-based feature extractor extracts semantic features from the image dataset, including: The hyperparameters deployed on the edge device are A feature extractor extracts semantic features from images of the image dataset; The extracted semantic features are quantized by an analog-to-digital conversion module and modulated by a digital modulator to obtain modulated semantic features.

3. The semantic communication retransmission method based on long short-term memory network according to claim 2 is characterized in that The transmitting or retransmitting the extracted semantic features includes: Transmitting the modulated semantic features to an LSTM classifier deployed on an edge server through a noisy and fading wireless channel; Or based on the noisy and fading wireless channel, the modulated semantic features are retransmitted using a retransmission mechanism of the retransmission modulation feature, new Gaussian white noise is added, and the retransmitted signal is transmitted to the LSTM classifier.

4. The semantic communication retransmission method based on long short-term memory network according to claim 1 is characterized in that The method further comprises: According to the preset uniform quantization parameter, the range of features extracted from the input image is divided into equal intervals; based on The classification network is trained using the function and the cross entropy loss function.

5. A semantic communication retransmission system based on a long short-term memory network, used to implement the semantic communication retransmission method based on a long short-term memory network according to any one of claims 1 to 4, characterized in that: include: An acquisition module is used to obtain an image dataset to be classified; A feature extraction module, configured to extract semantic features from the image dataset using a feature extractor based on a classification network; A classification module, configured to transmit or retransmit the extracted semantic features, and process and classify the semantic features with noise after transmission or retransmission based on the LSTM classifier of the classification network; The output module is used to output the image classification results corresponding to the image dataset.

6. A terminal, characterized in that: include: A processor and a memory, wherein the memory stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on a long short-term memory network is executed by the processor, it is used to implement the operation of the semantic communication retransmission method based on a long short-term memory network as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a semantic communication retransmission program based on a long short-term memory network, and when the semantic communication retransmission program based on a long short-term memory network is executed by a processor, it is used to implement the operation of the semantic communication retransmission method based on a long short-term memory network as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Archive key element identification and classification method based on improved LSTM semantic modeling

    CN119026601A

  • Self-adaptive semantic communication method for machine vision task

    CN119544143A