A multi-task-oriented semantic communication method, device, equipment and storage medium
By constructing a semantic communication framework based on maximizing mutual information, the transmission of semantic information is optimized, solving the performance and efficiency problems of multi-task communication, and realizing efficient semantic information transmission under different channel and task conditions. It is applicable to a variety of wireless communication systems and tasks.
Patent Information
- Application Number
- CN202411567889.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-05
AI Technical Summary
Existing technologies struggle to improve the transmission efficiency of semantic information and reduce communication and computational overhead while ensuring multitasking performance.
We construct a semantic communication framework based on maximizing mutual information, including a semantic communication sender and receiver. Through an adaptive variable-length semantic representation coding model and a quantization model, we optimize the end-to-end semantic communication system, enabling semantic information to simultaneously possess generative and discriminative properties, and adapt to different channel states and task requirements.
It improves the transmission efficiency of semantic information under multi-tasking conditions, reduces communication and computing overhead, and is applicable to various scenarios such as wireless networks, satellite networks, and Wi-Fi networks. It supports the transmission of information such as text, voice, images, and video, and completes tasks such as classification, recognition, reconstruction, clustering, and retrieval.
Smart Images

Figure CN119728012B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication technology, and specifically relates to a semantic communication method, apparatus, device, and storage medium for multi-task applications. Background Technology
[0002] Task-oriented semantic communication is an intelligent communication method that accomplishes task requirements by transmitting semantic information. Traditional syntax communication requires that the decoded information at the receiving end be strictly consistent with the encoded information at the sending end, i.e., to achieve bit-level error-free transmission. However, task-oriented semantic communication does not require strict matching of the encoding and decoding sequences at the sending and receiving ends; it only requires that the information processed at the receiving end can complete the corresponding task as needed.
[0003] Current research on task-oriented semantic communication is gradually becoming a hot topic in academia, and data transmission based on semantic information will be a highly competitive key technology. Designing task-oriented semantic communication systems allows the transmitted signals to simultaneously carry semantic information required by multiple tasks, thereby satisfying the needs of multiple tasks at the receiving end. Improving the transmission efficiency of semantic information and reducing communication and computational overhead while ensuring multi-tasking performance has practical application significance. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to propose a semantic communication method, apparatus, device and storage medium for multiple tasks, so as to optimize the end-to-end semantic communication system by maximizing the mutual information between source information and semantic representation, thereby transmitting semantic representations that are both generative and discriminative in the network, and providing the required semantic information for multiple target tasks of the receiving end.
[0005] In view of the above objectives, in a first aspect, the present invention provides a semantic communication method for multi-task communication, comprising:
[0006] A semantic communication framework based on maximizing mutual information is constructed. The semantic communication framework based on maximizing mutual information includes a semantic communication sender and a semantic communication receiver. The semantic communication sender performs semantic information representation, information encoding, and information transformation processing on the source information. The semantic communication receiver performs semantic information recovery, information decoding, and task reasoning processing on the received semantic information.
[0007] A semantic representation quantization model is constructed. The semantic representation quantization model performs preset precision quantization on the semantic information representation at the semantic communication sending end. The semantic symbols obtained after quantization are represented by the number of bits, and the semantic symbols obtained after quantization are mapped to constellation points.
[0008] An adaptive variable-length semantic representation coding model is constructed. Based on the current channel state and the target task requirements, the adaptive variable-length semantic representation coding model selectively activates specified dimensions of the semantic information representation sequentially and continuously. Under the premise of ensuring the performance of the target task, the semantic representation transmission overhead is adaptively reduced. The target task includes at least one of classification, recognition, reconstruction, clustering and retrieval.
[0009] As a preferred solution for a multi-task-oriented semantic communication method, the semantic communication sending end uses convolutional neural networks and deep neural networks to represent the source information semantically.
[0010] The semantic communication sending end performs information encoding on the source information, including semantic encoding, source encoding, and channel encoding.
[0011] As a preferred solution for multi-task semantic communication methods, the semantic representation quantification model utilizes artificial intelligence models and maximizes global and local mutual information to enable the semantic information representation to possess both generative and discriminative properties.
[0012] As a preferred solution for a multi-task-oriented semantic communication method, the semantic communication receiver uses an artificial intelligence model to reconstruct the received semantic information and restore the source information during the semantic information recovery process.
[0013] The semantic communication receiver performs information decoding on the received semantic information, including semantic decoding, source decoding, and channel decoding.
[0014] As a preferred solution for semantic communication methods oriented towards multiple tasks, an asymmetric quantization algorithm is used in the process of mapping the quantized semantic symbols to constellation points to make the quantized semantic symbols map to constellation points with amplitude and phase intervals, so as to adapt to the radio frequency hardware design in actual digital communication systems.
[0015] As a preferred solution for multi-task-oriented semantic communication methods, the adaptive variable-length semantic representation coding model is based on the current channel state, including the signal-to-noise ratio and noise variance at the current moment.
[0016] The adaptive variable-length semantic representation coding model is based on target task requirements, including semantic representation pruning threshold and target task performance threshold.
[0017] Secondly, the present invention provides a semantic communication device for multi-task applications, comprising:
[0018] The semantic communication basic processing module is used to construct a semantic communication basic framework based on maximizing mutual information. The semantic communication basic framework based on maximizing mutual information includes a semantic communication sender and a semantic communication receiver. The semantic communication sender performs semantic information representation, information encoding, and information transformation processing on the source information. The semantic communication receiver performs semantic information recovery, information decoding, and task reasoning processing on the received semantic information.
[0019] The semantic representation quantization module is used to construct a semantic representation quantization model. The semantic representation quantization model performs preset precision quantization on the semantic information representation at the semantic communication sending end. The semantic symbols obtained after quantization are represented by the number of bits, and the semantic symbols obtained after quantization are mapped to constellation points.
[0020] An adaptive variable-length semantic representation coding module is used to construct an adaptive variable-length semantic representation coding model. The adaptive variable-length semantic representation coding model selectively activates specified dimensions of the semantic information representation sequentially and continuously based on the current channel state and target task requirements. Under the premise of ensuring the performance of the target task, it adaptively reduces the semantic representation transmission overhead. The target task includes at least one of classification, recognition, reconstruction, clustering and retrieval.
[0021] As a preferred embodiment of a multi-task-oriented semantic communication device, the semantic communication basic processing module includes:
[0022] The semantic communication sending end uses convolutional neural networks and deep neural networks to represent the semantic information of the source information.
[0023] The semantic communication sending end performs information encoding on the source information, including semantic encoding, source encoding, and channel encoding;
[0024] Using an artificial intelligence model, the received semantic information is reconstructed to restore the source information;
[0025] The semantic communication receiver performs information decoding on the received semantic information, including semantic decoding, source decoding, and channel decoding.
[0026] As a preferred embodiment of a multi-task-oriented semantic communication device, the semantic representation quantization module includes:
[0027] The semantic representation quantification model utilizes artificial intelligence models and maximizes global and local mutual information to enable the semantic information representation to possess both generative and discriminative properties.
[0028] An asymmetric quantization algorithm is used to map the quantized semantic symbols to constellation points with amplitude and phase intervals, in order to adapt to the RF hardware design in practical digital communication systems.
[0029] As a preferred embodiment of a multi-task-oriented semantic communication device, the adaptive variable-length semantic representation encoding module includes:
[0030] The adaptive variable-length semantic representation coding model is based on the current channel state, including the signal-to-noise ratio and noise variance at the current moment;
[0031] The adaptive variable-length semantic representation coding model is based on target task requirements, including semantic representation pruning threshold and target task performance threshold.
[0032] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements a multitasking semantic communication method of the first aspect or any possible implementation thereof.
[0033] Fourthly, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a multi-task-oriented semantic communication method of the first aspect or any possible implementation thereof.
[0034] As can be seen from the above, the technical solution provided by this invention constructs a semantic communication framework based on maximizing mutual information. This framework includes a semantic communication transmitter and a semantic communication receiver. The transmitter performs semantic information representation, encoding, and transformation processing on the source information. The receiver performs semantic information recovery, decoding, and task inference processing on the received semantic information. A semantic representation quantization model is constructed, which quantizes the semantic information representation at a preset precision at the transmitter. The quantized semantic symbols are represented using a number of bits and mapped to constellation points. An adaptive variable-length semantic representation encoding model is constructed, which selectively and continuously activates specified dimensions of the semantic information representation based on the current channel state and target task requirements. This adaptively reduces the semantic representation transmission overhead while ensuring target task performance. The target task includes at least one of classification, recognition, reconstruction, clustering, and retrieval. This invention can complete multiple tasks simultaneously through a single semantic information extraction, thereby improving the transmission efficiency of semantic information and reducing communication and computational overhead while ensuring multi-task performance. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in this invention or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A schematic diagram of the semantic communication method for multi-task purposes provided in an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of a DNN-based simultaneous multi-task joint source-channel encoding and decoding framework provided in an embodiment of the present invention.
[0038] Figure 3 The training process for the joint source-channel coding and decoding framework based on maximizing mutual information provided in this embodiment of the invention;
[0039] Figure 4 This is a schematic diagram of the adaptive variable-length semantic representation coding model structure provided in an embodiment of the present invention;
[0040] Figure 5 This is a diagram illustrating the architecture of a multi-task-oriented semantic communication device provided in an embodiment of the present invention.
[0041] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0043] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this invention should have the ordinary meaning understood by those skilled in the art to which this invention pertains. The terms "comprising" or "including," or similar words used in the embodiments of this invention, mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.
[0044] In related technologies, task-oriented semantic communication (TOC) is an intelligent communication method that fulfills task requirements by transmitting semantic information. Traditional syntax-based communication requires strict consistency between the decoded information at the receiving end and the encoded information at the sending end, achieving bit-level error-free transmission. However, task-oriented semantic communication does not require strict matching of the encoding and decoding sequences at the sending and receiving ends; it only requires that the information processed by the receiving end can complete the corresponding task as needed. Because task-oriented semantic communication can further compress the transmitted content, it redefines the error requirements for information transmission in the system, increasing the fault tolerance of information transmission. This breaks through the transmission bottleneck of classic syntax-based communication systems, further improving transmission efficiency and providing a new solution for the evolution of next-generation wireless communication systems.
[0045] Current research on task-oriented semantic communication is gradually becoming a hot topic, and data transmission based on semantic information will be a highly competitive key technology. Further design of task-oriented semantic communication systems can enable the signals transmitted at the sending end to simultaneously carry semantic information required by multiple tasks, thereby simultaneously meeting the needs of multiple tasks at the receiving end, giving the semantic communication system richer functionality and higher transmission efficiency.
[0046] Among related technologies, Mutual Information Neural Estimation (MINE) is a newly proposed method that uses neural networks to assist in semi-supervised or unsupervised representation learning. Computing the mutual information of high-dimensional features using traditional methods is a very difficult task due to the complex probability distributions and high computational costs involved. However, MINE overcomes this challenge by leveraging the powerful representational capabilities of neural networks to effectively estimate the mutual information between high-dimensional features. Specifically, MINE captures potentially valuable information representations in the data by maximizing the lower bound of the mutual information estimated by the neural network. This method not only improves the accuracy of mutual information estimation but also enables the estimation of mutual information in high-dimensional data, making semi-supervised and unsupervised learning on complex datasets more feasible and efficient. In summary, MINE provides an effective tool for estimating the mutual information between high-dimensional features, greatly promoting research and applications in representation learning and related fields. By integrating mutual information estimation into the design of a multi-task semantic communication framework, the sender can actively extract useful information from the information source, rather than passively generating intermediate features through label matching or pixel reconstruction. The semantic representations extracted by this method can simultaneously improve the performance of both generation and discrimination tasks, thereby more effectively handling various complex task requirements.
[0047] In view of this, embodiments of the present invention propose a semantic communication method, apparatus, device, and storage medium for multi-task applications. By maximizing global and local mutual information, the semantic representation extracted by the sending end possesses both generative and discriminative properties, enabling it to be directly used to complete both generation and discrimination tasks. Furthermore, by minimizing the mean square error for end-to-end training of the sending and receiving ends, the receiving end can reconstruct the original content based on the received semantic representation. This allows multiple tasks to be completed simultaneously through a single semantic information extraction, thereby improving the transmission efficiency of semantic information and reducing communication and computational overhead while ensuring multi-task performance. The specific details of the embodiments of the present invention are as follows.
[0048] See Figure 1 , Figure 2 , Figure 3 and Figure 4 This invention provides a semantic communication method for multi-task applications, comprising the following steps:
[0049] S1. Construct a semantic communication framework based on maximizing mutual information. The semantic communication framework based on maximizing mutual information includes a semantic communication sender and a semantic communication receiver. The semantic communication sender performs semantic information representation, information encoding, and information transformation processing on the source information. The semantic communication receiver performs semantic information recovery, information decoding, and task reasoning processing on the received semantic information.
[0050] S2. Construct a semantic representation quantization model. The semantic representation quantization model performs preset precision quantization on the semantic information representation at the semantic communication sending end. The semantic symbols obtained after quantization are represented by the number of bits, and the semantic symbols obtained after quantization are mapped to constellation points.
[0051] S3. Construct an adaptive variable-length semantic representation coding model. The adaptive variable-length semantic representation coding model selectively activates specified dimensions of the semantic information representation sequentially and continuously based on the current channel state and the target task requirements. Under the premise of ensuring the performance of the target task, it adaptively reduces the semantic representation transmission overhead. The target task includes at least one of classification, recognition, reconstruction, clustering and retrieval.
[0052] In this embodiment, by maximizing the mutual information between source information and semantic representation, the end-to-end semantic communication system is optimized, thereby transmitting semantic representations that are both generative and discriminative in the network, providing the necessary semantic information for multiple target tasks at the receiving end. The network includes wireless networks, satellite networks, Wi-Fi networks, etc. Source information includes text, voice, images, video, etc. Target tasks include classification, recognition, reconstruction, clustering, retrieval, etc. Semantic representations include intermediate parameters output by various artificial intelligence models, which may include convolutional neural networks, deep neural networks, etc.
[0053] In this embodiment, a semantic communication framework for a joint source channel is constructed based on maximizing mutual information, enabling the extracted semantic representation to possess both generative and discriminative properties. An asymmetric quantization method is used to map the quantized semantic symbols to fewer constellation points with clearer amplitude and phase intervals. An adaptive semantic representation encoding method based on a multilayer perceptron allows the semantic representation to adapt to different channel states and sequentially and continuously activate different dimensions. This invention is applicable to various scenarios, including single-user and multi-user transmissions. The semantic communication framework based on maximizing mutual information can be implemented using various infrastructures such as deep neural networks (DNNs), convolutional neural networks (CNNs), and attention mechanisms.
[0054] For details, see Figure 2 This embodiment uses a DNN to illustrate the solution. The semantic communication framework based on maximizing mutual information can be applied to various tasks such as classification, retrieval, and reconstruction. This embodiment uses image classification and image reconstruction tasks to illustrate the solution.
[0055] At the transmitting end, the joint source-channel coding network E a (·) Source information The mapping to a 2n-dimensional semantic representation vector z is as follows:
[0056]
[0057] Then, the semantic representation vector z is transformed into n complex symbols, which serve as the modulation signal. The power is then normalized as follows:
[0058]
[0059] Unlike traditional information compression, semantic representation not only includes the semantic information of the source information but also includes compensation information to combat the randomness of the wireless channel. The transformed semantic symbols have continuously varying amplitudes and phases, which can be demodulated by a joint source-channel decoding network based on deep learning, or the semantic representation can be quantized before modulation.
[0060] Since the receiver has orthogonal time-frequency resource blocks, the impact of inter-signal interference on transmission can be ignored, and the received signal y is obtained as follows:
[0061] y = hx + n
[0062] Where h is the channel parameter, and n represents the independent and identically distributed channel noise vector, which has a mean of 0 and a variance of δ. 2 Symmetric complex Gaussian distribution Using channel estimation and channel equalization techniques, the receiver can obtain the channel parameters h, and then recover the received signal as follows:
[0063]
[0064] Therefore, it can be extended to Gaussian, Rayleigh, and Ricean channels, among others. The receiver uses the same transformation strategy as the transmitter to transform the noisy received signal. Transform back to noisy semantic representation As input to the receiver's joint source-channel decoding network, it simultaneously completes image reconstruction and image classification tasks.
[0065] For image reconstruction tasks, the receiver uses a joint source-channel decoding network D. β (·) Utilizing noisy semantic representations Obtain the reconstructed image as follows:
[0066]
[0067] Among them, the joint source-channel coding network E at the transmitting end α (·) and the joint source-channel decoding network D at the receiving end β (·) is jointly trained end-to-end. The performance of the image reconstruction task is measured by the peak signal-to-noise ratio (PSNR), which represents the visual difference between two images:
[0068]
[0069] Where MAX represents the maximum value of all pixels in the image, and MSE(,) represents the mean square error calculation.
[0070] For image classification tasks, noisy semantic representations are input into a pre-designed classifier ψ(·). As a functional, it can be designed according to the specific task at the receiving end and can be implemented in different forms, including DNN models, K-nearest neighbor classifiers, and support vector machines. Since the received semantic representations are discriminative and can be directly used for classification tasks, the classifier ψ(·) does not need to be integrated with the joint source-channel coding network E at the transmitting end. α (·) Perform end-to-end joint training. The performance of image classification tasks is measured by accuracy (ACC), which represents the proportion of correctly classified samples out of the total number of test samples.
[0071]
[0072] In traditional single-task semantic communication frameworks, the two tasks mentioned above cannot be completed simultaneously. This is because task-specific joint source-channel coding-decoding models are trained end-to-end, employing task-specific loss functions designed to adjust model parameters to make the output closer to the task's expected value. Different task objectives typically lead to different optimization directions, making it difficult to achieve multiple objectives simultaneously. Furthermore, in traditional end-to-end training paradigms, the output is actively optimized only based on the task objective, while the semantic representations output by intermediate layers are passively learned in an unsupervised manner. Therefore, the transmitted semantic representations are merely intermediate results for a specific objective task, lacking discriminative or interpretable qualities.
[0073] Therefore, the present invention aims to optimize semantic representation and task results simultaneously based on a joint source-channel coding and decoding framework that maximizes mutual information, thereby achieving the extraction of discriminative semantic information and improving the generation capabilities of semantic representation and joint source-channel decoder.
[0074] In this embodiment, at the transmitting end, global mutual information is first maximized by maximizing the Jensen-Shannon (JS) divergence between the source image s and the semantic representation z. The JS divergence is defined as follows:
[0075]
[0076] Among them, T θ (·) denotes the discriminator, which consists of a neural network with parameters θ, and s′ represents a pseudo-image sample. The JS divergence is maximized adversarially, and the discriminator aims to maximize the discrimination score T of the real samples. θ While minimizing the discrimination score T of the pseudo-samples (s, z), we can simultaneously achieve the desired result. θ (s′,z) Thus, the trained discriminator can distinguish whether the semantic representation z comes from the real source image, thereby maximizing the correlation between the semantic representation and the original sample (through the joint distribution term) and improving the independence between the semantic representation and irrelevant samples (through the marginal distribution term).
[0077] like Figure 3 As shown, by focusing on different parts of the image (global or local), the trained semantic representation can contain information specific to different tasks. Specifically, global mutual information, using the entire image as its receptive field, maximizes the JS divergence between the source image s and the semantic representation z by optimizing the joint source-channel coding network and the global discriminator.
[0078]
[0079] Where, θ gThis represents the neural network parameters of the global discriminator. Since global mutual information considers all parts of the image simultaneously, maximizing global mutual information can optimize the performance of image reconstruction tasks.
[0080] For image classification tasks, we should focus on local image patches rather than the entire image, because the edges detected in local image patches contain structural knowledge of the image, which can improve the performance of the classification task. For example... Figure 3 As shown, the joint source-channel coding network consists of two parts, satisfying... Part E ω The function of (·) is to extract feature maps from the source image. The N×N feature vectors of depth C′ in the feature map correspond to the structural information of N×N local image patches in the source image, as shown below:
[0081] f = E ω (s)
[0082] Part Two E φ The function of (·) is to summarize all feature vectors into a semantic representation.
[0083] Therefore, to enhance the discriminative performance of semantic representation in classification tasks, each local image patch of the image is used as the receptive field. The average JS divergence between the N×N local image patches of the source image and the semantic representation z is maximized by optimizing the joint source-channel coding network and the local discriminator as follows:
[0084]
[0085] Where i represents the local image patch index, θ l This represents the neural network parameters of the local discriminator. Unlike traditional image classification training methods based on label-supervised learning, the semantic representation obtained by joint source-channel encoding and decoding based on maximizing mutual information through local mutual information maximization includes knowledge of the edge structures detected in each local image patch. This knowledge can be used to distinguish image samples of different categories. The classification method is based on the visual feature similarities and differences between image samples, rather than relying on data matching between each sample and its label. Therefore, it belongs to the unsupervised learning method, and the training process of semantic representation does not require labels.
[0086] By jointly training a joint source-channel encoder-decoder network, a global discriminator, and a local discriminator based on maximizing mutual information, both global and local mutual information can be optimized, thereby learning a semantic representation that is both generative and discriminative. Furthermore, the performance of the image reconstruction task is affected not only by the quality of the semantic representation but also by the joint source-channel decoder network. Therefore, the joint source-channel encoder-decoder based on maximizing mutual information simultaneously minimizes the end-to-end mean square error to reduce pixel-level differences between the source and reconstructed images. The final loss function is defined as follows:
[0087]
[0088] Here, λ is used to adjust the balance between image reconstruction performance and semantic representation quality, while μ1 and μ2 control the degree of attention paid to global and local mutual information during training. By adjusting the above hyperparameters, the joint source-channel communication framework based on maximizing mutual information can focus only on the image classification task or the image reconstruction task, or it can balance the performance of the two tasks to simultaneously meet the needs of the receiver for both tasks.
[0089] In this embodiment, the semantic representation output by the joint source-channel coding network consists of 32-bit floating-point numbers. Directly mapping this representation yields semantic symbols with continuously varying and randomly distributed amplitude and phase values over a wide range. This near-continuous constellation mapping requires the system to identify high-resolution amplitude and phase shifts, which existing RF hardware designs cannot meet. To enable joint source-channel coding and decoding based on maximizing mutual information to be applied in practical wireless communication systems, a semantic representation quantization method based on asymmetric quantization is used. This method quantizes the semantic representation output by the joint source-channel coding network into fewer bits, thereby mapping it to a constellation diagram with fewer, more regularly distributed, and easier-to-distinguish amplitude and phase shifts.
[0090] Specifically, firstly, each floating-point number z in the semantic representation z j Quantized into q-bit integers
[0091]
[0092] Where, ρ s This represents a scaling factor, which maps 32-bit floating-point numbers from their original range to a smaller range [0, 2]. q -1]. The scaling factor is calculated as follows:
[0093]
[0094] The `round` function performs rounding, and the `clamp` function is used to remove values exceeding [0, 2]. q The quantization outlier of [-1] is defined as:
[0095]
[0096] Therefore, the quantized semantic representation can be mapped to discretely distributed semantic symbols, consisting only of 2 q -1 constellation point mapping, with easily identifiable amplitude and phase differences between constellation points.
[0097] Since quantization causes abrupt changes in the input parameters of the subsequent joint source-channel decoding network, leading to a decrease in learning performance, the quantized values cannot be used directly. Therefore, dequantization is required to ensure that the quantized semantic representation retains its discrete characteristics while having a numerical distribution similar to that before quantization. The dequantization operation is represented as follows:
[0098]
[0099] The dequantized semantic representation vector is represented as It can still be mapped to 2 q -1 constellation point, but numerically approximates the semantic representation vector z before quantization. A larger quantization bit number q results in a semantic representation closer to the original semantic representation with less information loss and better target task performance at the receiver, but also higher communication overhead. Adjusting the quantization bit number q can balance semantic transmission overhead and target task performance. Fine-tuning with a few rounds within a pre-trained joint source-channel coding / decoding framework based on maximizing mutual information allows the framework to adapt to different quantization bit numbers.
[0100] In this embodiment, since the channel environment in real-world applications is dynamically changing, the length of the semantic representation should be adjusted according to the changing channel conditions. Specifically, more dimensions should be transmitted under poor channel conditions, while fewer dimensions should be activated under good channel conditions to save communication resources. Therefore, the simultaneous multi-task semantic communication method introduces adaptive variable-length semantic representation coding, which can sequentially and continuously activate a corresponding number of semantic representation dimensions according to the current channel state, such as... Figure 4 As shown. Therefore, apart from transmitting the shortened semantic representation vector after the mask, the sender only needs to transmit an index indicating the end position to inform the receiver of the activation status of the semantic representation, thereby ensuring the normal progress of subsequent processes.
[0101] First, the transmitting end receives information about the channel state, including signal-to-noise ratio and channel fading, through the feedback channel. This channel state information can be used as auxiliary knowledge input into the joint source-channel coding network to support adaptive semantic representation variable-length coding. For example, the noise variance δ 2The input is fed into a lightweight multilayer perceptron g(·) deployed at the transmitter, generating an activation mask m(·) with the same dimensions as the original semantic representation. To enable activation of more dimensions even in poor channel conditions, the output g of each dimension of the multilayer perceptron... i (δ 2 This should be about the noise variance δ. 2 The activation mask is a non-negative monotonically increasing function. Furthermore, to achieve continuous activation characteristics of the semantic representation, the activation mask should be multiplied by the upper triangular identity matrix, thus ensuring that the value of each dimension of the mask is less than the value of the previous dimension, i.e., the value monotonically decreases starting from the first dimension. The calculation method for any dimension k of the activation mask is as follows:
[0102]
[0103]
[0104]
[0105] Therefore, the larger the noise variance, the larger the value of each dimension of the activation mask, decreasing sequentially from the first dimension. A predefined pruning threshold η can be defined based on the trade-off between task performance and communication overhead. Mask dimensions with values greater than η are assigned a value of 1, indicating that the semantic representation of that dimension is temporarily activated; conversely, those with values less than η are assigned a value of 0, indicating that the semantic representation of that dimension is temporarily pruned.
[0106]
[0107] Therefore, the larger the noise variance, the more dimensions of the activated semantic representation are, and the distribution is continuous. Finally, at the sending end, the assigned activation mask and the original semantic representation are multiplied by Hadamard to obtain the variable-length encoded semantic representation.
[0108]
[0109] By integrating a multilayer perceptron into the joint source-channel coding network for joint training, the resulting activation mask can continuously activate important dimensions in the semantic representation. The importance of each dimension can be automatically learned during the joint training process. Fine-tuning is then performed in a few rounds within the trained framework. In each mini-batch training process, a noise variance value is randomly selected from the possible range of noise variance values as the training input, enabling the framework to achieve adaptive variable-length semantic representation coding under different channel conditions.
[0110] In summary, this invention constructs a semantic communication framework based on maximizing mutual information. This framework includes a semantic communication transmitter and a semantic communication receiver. The transmitter performs semantic information representation, encoding, and transformation processing on the source information. The receiver performs semantic information recovery, decoding, and task inference processing on the received semantic information. A semantic representation quantization model is constructed, which quantizes the semantic information representation at a preset precision at the transmitter. The quantized semantic symbols are represented using a number of bits and mapped to constellation points. An adaptive variable-length semantic representation encoding model is constructed, which selectively and continuously activates specified dimensions of the semantic information representation based on the current channel state and target task requirements. This adaptively reduces the semantic representation transmission overhead while ensuring target task performance. The target task includes at least one of classification, recognition, reconstruction, clustering, and retrieval. This invention can complete multiple tasks simultaneously through a single semantic information extraction, thereby improving the transmission efficiency of semantic information and reducing communication and computational overhead while ensuring multi-task performance.
[0111] It should be noted that the method of this embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this embodiment, and the multiple devices will interact with each other to complete the method described.
[0112] It should be noted that the above description describes some embodiments of the present invention. In some cases, the described actions or steps can be performed in a different order than that shown in the above embodiments and the desired result can still be achieved. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0113] See Figure 5 Based on the same inventive concept, corresponding to any of the methods in the above embodiments, this embodiment of the invention also provides a semantic communication device for multi-task applications, comprising:
[0114] The semantic communication basic processing module 100 is used to construct a semantic communication basic framework based on maximizing mutual information. The semantic communication basic framework based on maximizing mutual information includes a semantic communication sender and a semantic communication receiver. The semantic communication sender performs semantic information representation, information encoding, and information transformation processing on the source information. The semantic communication receiver performs semantic information recovery, information decoding, and task reasoning processing on the received semantic information.
[0115] The semantic representation quantization module 200 is used to construct a semantic representation quantization model. The semantic representation quantization model performs preset precision quantization on the semantic information representation at the semantic communication sending end. The semantic symbols obtained after quantization are represented by the number of bits, and the semantic symbols obtained after quantization are mapped to constellation points.
[0116] An adaptive variable-length semantic representation coding module 300 is used to construct an adaptive variable-length semantic representation coding model. The adaptive variable-length semantic representation coding model selectively activates specified dimensions of the semantic information representation sequentially and continuously based on the current channel state and target task requirements. Under the premise of ensuring the performance of the target task, it adaptively reduces the semantic representation transmission overhead. The target task includes at least one of classification, recognition, reconstruction, clustering and retrieval.
[0117] In this embodiment, the semantic communication basic processing module 100 includes:
[0118] The semantic communication sending end uses convolutional neural networks and deep neural networks to represent the semantic information of the source information.
[0119] The semantic communication sending end performs information encoding on the source information, including semantic encoding, source encoding, and channel encoding;
[0120] Using an artificial intelligence model, the received semantic information is reconstructed to restore the source information;
[0121] The semantic communication receiver performs information decoding on the received semantic information, including semantic decoding, source decoding, and channel decoding.
[0122] In this embodiment, the semantic representation quantization module 200 includes:
[0123] The semantic representation quantification model utilizes artificial intelligence models and maximizes global and local mutual information to enable the semantic information representation to possess both generative and discriminative properties.
[0124] An asymmetric quantization algorithm is used to map the quantized semantic symbols to constellation points with amplitude and phase intervals, in order to adapt to the RF hardware design in practical digital communication systems.
[0125] In this embodiment, the adaptive variable-length semantic representation coding module 300 includes:
[0126] The adaptive variable-length semantic representation coding model is based on the current channel state, including the signal-to-noise ratio and noise variance at the current moment;
[0127] The adaptive variable-length semantic representation coding model is based on target task requirements, including semantic representation pruning threshold and target task performance threshold.
[0128] The apparatus of the above embodiments is used to implement a corresponding multi-task-oriented semantic communication method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0129] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a multi-task-oriented semantic communication method as described in any of the above embodiments.
[0130] Figure 6 This illustration shows a more specific hardware structure diagram of an electronic device provided in this embodiment. The device may include: a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, memory 420, input / output interface 430, and communication interface 440 are interconnected internally via the bus 450.
[0131] The processor 410 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0132] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.
[0133] Input / output interface 430 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0134] The communication interface 440 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).
[0135] Bus 450 includes a pathway for transmitting information between various components of the device (e.g., processor 410, memory 420, input / output interface 430, and communication interface 440).
[0136] It should be noted that although the above-described device only shows the processor 410, memory 420, input / output interface 430, communication interface 440, and bus 450, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0137] The electronic devices described above are used to implement a corresponding multi-task-oriented semantic communication method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0138] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the present invention also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a multi-task-oriented semantic communication method as described in any of the above embodiments.
[0139] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0140] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute a multi-task-oriented semantic communication method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0141] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the invention is limited to these examples; within the framework of the invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of the different aspects of the embodiments of the invention as described above, which are not provided in detail for the sake of brevity.
[0142] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of the invention, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of the invention, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of the invention will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of the invention, it will be apparent to those skilled in the art that the embodiments of the invention may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0143] Although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0144] The embodiments of this invention are intended to cover all such substitutions, modifications, and variations falling within the scope of the claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this invention should be included within the scope of protection of this invention.
Claims
1. A semantic communication method for multi-task applications, wherein, include: A semantic communication framework based on maximizing mutual information is constructed. The semantic communication framework based on maximizing mutual information includes a semantic communication sender and a semantic communication receiver. The semantic communication sender performs semantic information representation, information encoding, and information transformation processing on the source information. The semantic communication receiver performs semantic information recovery, information decoding, and task reasoning processing on the received semantic information. A semantic representation quantization model is constructed. The semantic representation quantization model performs preset precision quantization on the semantic information representation at the semantic communication sending end. The semantic symbols obtained after quantization are represented by the number of bits, and the semantic symbols obtained after quantization are mapped to constellation points. An adaptive variable-length semantic representation coding model is constructed. Based on the current channel state and the target task requirements, the adaptive variable-length semantic representation coding model selectively activates specified dimensions of the semantic information representation sequentially and continuously. Under the premise of ensuring the performance of the target task, the semantic representation transmission overhead is adaptively reduced. The target task includes at least one of classification, recognition, reconstruction, clustering and retrieval.
2. The semantic communication method for multi-task applications according to claim 1, wherein, The semantic communication sending end uses convolutional neural networks and deep neural networks to represent the semantic information of the source information. The semantic communication sending end performs information encoding on the source information, including semantic encoding, source encoding, and channel encoding.
3. The semantic communication method for multi-task applications according to claim 1, wherein, The semantic representation quantification model utilizes artificial intelligence models and maximizes global and local mutual information to enable the semantic information representation to possess both generative and discriminative properties.
4. The semantic communication method for multi-task applications according to claim 1, wherein, During the semantic information recovery process of the received semantic information, the semantic communication receiving end uses an artificial intelligence model to reconstruct the received semantic information and recover the source information. The semantic communication receiver performs information decoding on the received semantic information, including semantic decoding, source decoding, and channel decoding.
5. A semantic communication method for multi-task applications according to claim 1, wherein, In the process of mapping the quantized semantic symbols to constellation points, an asymmetric quantization algorithm is used to map the quantized semantic symbols to constellation points with amplitude and phase intervals, so as to adapt to the radio frequency hardware design in actual digital communication systems.
6. A semantic communication method for multi-task applications according to claim 1, wherein, The adaptive variable-length semantic representation coding model is based on the current channel state, including the signal-to-noise ratio and noise variance at the current moment; The adaptive variable-length semantic representation coding model is based on target task requirements, including semantic representation pruning threshold and target task performance threshold.
7. A semantic communication device for multi-task applications, wherein, include: The semantic communication basic processing module is used to construct a semantic communication basic framework based on maximizing mutual information. The semantic communication basic framework based on maximizing mutual information includes a semantic communication sender and a semantic communication receiver. The semantic communication sender performs semantic information representation, information encoding, and information transformation processing on the source information. The semantic communication receiver performs semantic information recovery, information decoding, and task reasoning processing on the received semantic information. The semantic representation quantization module is used to construct a semantic representation quantization model. The semantic representation quantization model performs preset precision quantization on the semantic information representation at the semantic communication sending end. The semantic symbols obtained after quantization are represented by the number of bits, and the semantic symbols obtained after quantization are mapped to constellation points. An adaptive variable-length semantic representation coding module is used to construct an adaptive variable-length semantic representation coding model. The adaptive variable-length semantic representation coding model selectively activates specified dimensions of the semantic information representation sequentially and continuously based on the current channel state and target task requirements. Under the premise of ensuring the performance of the target task, it adaptively reduces the semantic representation transmission overhead. The target task includes at least one of classification, recognition, reconstruction, clustering and retrieval.
8. A semantic communication device for multi-task applications according to claim 7, wherein, In the semantic communication basic processing module: The semantic communication sending end uses convolutional neural networks and deep neural networks to represent the semantic information of the source information. The semantic communication sending end performs information encoding on the source information, including semantic encoding, source encoding, and channel encoding; Using an artificial intelligence model, the received semantic information is reconstructed to restore the source information; The semantic communication receiver performs information decoding on the received semantic information, including semantic decoding, source decoding, and channel decoding. In the semantic representation quantization module: The semantic representation quantification model utilizes artificial intelligence models and maximizes global and local mutual information to enable the semantic information representation to possess both generative and discriminative properties. An asymmetric quantization algorithm is used to map the quantized semantic symbols to constellation points with amplitude and phase intervals, in order to adapt to the radio frequency hardware design in practical digital communication systems. In the adaptive variable-length semantic representation encoding module: The adaptive variable-length semantic representation coding model is based on the current channel state, including the signal-to-noise ratio and noise variance at the current moment; The adaptive variable-length semantic representation coding model is based on target task requirements, including semantic representation pruning threshold and target task performance threshold.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the program, it implements a multi-task-oriented semantic communication method as described in any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute a multi-task-oriented semantic communication method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multitask-oriented voice semantic communication method, device and system
CN116884404A
Information-aware graph contrastive learning
US20220383108A1