Data transmission method and apparatus, and electronic device, medium and program product
Patent Information
- Application Number
- PCT/CN2026/082989
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-12
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026082989_01102026_PF_FP_ABST
Abstract
Description
Data transmission methods, apparatus, electronic devices, media and software products
[0001] Cross-references to related applications
[0002] This disclosure claims priority to Chinese Patent Application No. 202510370852.8, filed in China on March 27, 2025, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of data transmission technology, specifically to a data transmission method, apparatus, electronic device, medium, and program product. Background Technology
[0004] To address the constraints of high dynamics, long distances, time-varying characteristics, and limited resources in satellite-to-ground coverage scenarios, efficient information extraction, reliable information recovery, and seamless network coverage extension based on joint source-channel coding are required. Designing an efficient and flexible adaptive network resource allocation model for satellite-to-ground transmission network architectures with multi-task parallelism in Artificial Intelligence (AI), and exploring technical solutions for efficient information extraction and reliable information reconstruction, are of great significance in enhancing satellite-to-ground transmission capabilities.
[0005] AI multi-task communication in related technologies is mainly based on joint source-channel coding. A common approach is to design a joint source-channel encoder at the transmitting end to learn and extract deep features from the raw data provided by the mobile device, such as images, videos, audio, and text, in the latent space, and then transmit these deep features via a wireless channel. Correspondingly, at the receiving end, a joint source-channel decoder is designed to receive these deep features, which are affected by the wireless channel, on the server side, and a corresponding task execution network is designed to obtain the task results, which are then sent back to the mobile device. The specific network structure mainly uses Convolutional Neural Networks (CNNs) and Transformers.
[0006] Since multi-task communication based on joint source-channel coding and decoding in related technologies has no channel awareness, designing transmission depth features with fixed dimensions may lead to problems such as degraded task performance or wasted communication overhead. Summary of the Invention
[0007] This disclosure provides a data transmission method, apparatus, electronic device, medium, and program product, which are beneficial for improving task performance and reducing communication overhead.
[0008] In a first aspect, embodiments of this disclosure provide a data transmission method applied to a first communication device, the method comprising:
[0009] The first deep feature vector of the data to be transmitted is obtained by encoding the data to be transmitted through the first neural network model in the joint source-channel encoder;
[0010] The first depth feature vector is transmitted to the second communication device; wherein...
[0011] The second communication device is used to perform data recovery and visual recognition on the first deep feature vector through the second neural network model in the joint source-channel decoder, so as to obtain the recovered data and visual recognition information;
[0012] The first neural network model and the second neural network model are a joint model obtained by joint training based on the objective function.
[0013] Secondly, embodiments of this disclosure provide a data transmission method applied to a second communication device, the method comprising:
[0014] Receive a first depth feature vector sent by a first communication device, wherein the first depth feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on a first neural network model in a joint source-channel encoder;
[0015] The first deep feature vector is used for data recovery and visual recognition based on the second neural network model in the joint source-channel decoder to obtain the recovered data and visual recognition information. The first neural network model and the second neural network model are joint models obtained by joint training based on the objective function.
[0016] Thirdly, embodiments of this disclosure provide a data transmission apparatus applied to a first communication device, the data transmission apparatus comprising:
[0017] The encoding module is used to encode the data to be transmitted using the first neural network model in the joint source-channel encoder to obtain the first deep feature vector of the data to be transmitted.
[0018] A transmission module is used to transmit the first depth feature vector to a second communication device; wherein,
[0019] The second communication device is used to perform data recovery and visual recognition on the first deep feature vector through the second neural network model in the joint source-channel decoder, so as to obtain the recovered data and visual recognition information;
[0020] The first neural network model and the second neural network model are a joint model obtained by joint training based on the objective function.
[0021] Fourthly, embodiments of this disclosure provide a data transmission apparatus applied to a second communication device, the data transmission apparatus comprising:
[0022] The receiving module is used to receive a first depth feature vector sent by the first communication device, wherein the first depth feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on the first neural network model in the joint source-channel encoder;
[0023] The processing module is used to perform data recovery and visual recognition on the first deep feature vector based on the second neural network model in the joint source-channel decoder, so as to obtain the recovered data and visual recognition information. The first neural network model and the second neural network model are joint models obtained by joint training based on the objective function.
[0024] Fifthly, embodiments of this disclosure provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the data transmission method as described in the first and second aspects above.
[0025] In a sixth aspect, embodiments of this disclosure provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the data transmission method as described in the first and second aspects above.
[0026] In a seventh aspect, embodiments of this disclosure provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the data transmission method as described in the first and second aspects above.
[0027] In this embodiment of the disclosure, since the first neural network model and the second neural network model are joint models obtained by joint training based on the objective function, the joint model's ability to perceive the channel used to transmit deep feature vectors can be improved during the training process of the first neural network model and the second neural network model, thereby improving task performance and reducing communication overhead. Attached Figure Description
[0028] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments of this disclosure will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 is a flowchart of one of the data transmission methods provided in this embodiment of the present disclosure;
[0030] Figure 2 is a second flowchart of a data transmission method provided in an embodiment of this disclosure;
[0031] Figure 3 is a schematic diagram of a multi-task channel-aware communication system model in an embodiment of this disclosure;
[0032] Figure 4 shows the joint source-channel encoder f in an embodiment of this disclosure. e Structural diagram;
[0033] Figure 5 shows the joint source-channel decoder f in an embodiment of this disclosure. d Structural diagram;
[0034] Figure 6 is a flowchart of a data transmission method provided in an embodiment of this disclosure;
[0035] Figure 7 is a schematic diagram of one of the structures of a data transmission device provided in an embodiment of this disclosure;
[0036] Figure 8 is a second schematic diagram of a data transmission device provided in an embodiment of this disclosure;
[0037] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0038] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0039] Please refer to Figure 1, which is a flowchart illustrating a data transmission method provided in an embodiment of this disclosure, applied to a first communication device. The data transmission method includes the following steps:
[0040] Step 101: Encode the data to be transmitted using the first neural network model in the joint source-channel encoder to obtain the first deep feature vector of the data to be transmitted;
[0041] Step 102: Transmit the first depth feature vector to the second communication device;
[0042] In this embodiment of the disclosure, the second communication device is used to perform data recovery and visual recognition on the first deep feature vector through the second neural network model in the joint source-channel decoder, so as to obtain the recovered data and visual recognition information; the first neural network model and the second neural network model are a joint model obtained by joint training based on the objective function.
[0043] The data to be transmitted can be image data obtained from surveys in various scenarios, and the visual recognition can be various visual tasks in related technologies, such as image classification and image recognition. For ease of understanding, the following uses image classification as an example to further explain the method provided in the embodiments of this disclosure. The visual recognition information is the image category of the data to be transmitted. In some embodiments of this disclosure, the image category can include the following seven categories: environment, computer room, switching power supply, power distribution equipment, rooftop, battery, and seven other categories.
[0044] The first communication device in this embodiment can be a mobile device. The aforementioned joint source-channel encoder can be deployed on the first communication device, or the joint source-channel encoder can also be an independent device jointly deployed with the first communication device, in which case the joint source-channel encoder is communicatively connected to the first communication device. The aforementioned second communication device can be a server. The aforementioned joint source-channel decoder can be deployed on the second communication device, or the joint source-channel decoder can also be an independent device jointly deployed with the second communication device, in which case the joint source-channel decoder is communicatively connected to the second communication device.
[0045] The data recovery in this embodiment specifically includes: a second neural network model recovering the data to be transmitted based on the received depth feature vector, and correspondingly, the recovered data is the image generated by image recovery of the data to be transmitted.
[0046] In this embodiment, since the first neural network model and the second neural network model are joint models obtained by joint training based on the objective function, the joint model's ability to perceive the channel used to transmit deep feature vectors can be improved during the training process of the first neural network model and the second neural network model, thereby improving task performance and reducing communication overhead.
[0047] Optionally, the objective function is a function constructed based on a first parameter, a second parameter, and a third parameter. The first parameter is used to characterize the dimension of the depth feature vector received by the joint source-channel decoder. The second parameter is used to characterize the mutual information between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to characterize the mutual information between the visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder.
[0048] In related technologies, deep learning-based multi-task communication has a drawback compared to traditional separate source-channel coding / decoding communication. Multi-task communication based on joint source-channel coding / decoding lacks channel awareness; related technologies only jointly train the transceiver under fixed channel conditions, setting the first deep feature vector to a fixed-dimensional feature vector. Since the optimal feature vector dimension (i.e., the number of transmitted symbols) cannot be obtained based on channel conditions, this fixed-dimensional feature vector extraction will encounter problems in two situations. Firstly, if the feature vector dimension used as a hyperparameter in the deep learning network during joint training is lower than the optimal feature vector dimension under the channel conditions, the transmitted deep features will have a reduced ability to resist channel influences, and the extracted deep features will contain lost semantic information related to multiple tasks, leading to a decrease in multi-task performance. Secondly, if the feature vector dimension used as a hyperparameter in the deep learning network during joint training is higher than the optimal feature vector dimension under the channel conditions, the extracted feature vector will contain task-irrelevant redundancy, resulting in an increased number of transmitted symbols. With a constant symbol rate, this leads to wasted communication overhead and increased transmission latency.
[0049] It is evident that multi-task communication based on joint source-channel coding and decoding in related technologies lacks channel awareness, and the fixed-dimensional design of transmission depth features leads to degraded task performance or wasted communication overhead. Furthermore, the lack of additional channel state information (CSI) acquisition limits its application to simple communication environments, making it unsuitable for more complex wireless channels.
[0050] It should be noted that the above-mentioned transmission of the first depth feature vector to the second communication device may refer to the following: after the joint source channel encoder has encoded the first depth feature vector, the joint source channel encoder directly transmits the first depth feature vector to the second communication device via wireless channel transmission.
[0051] It is understandable that during the training of the model, the value of the first parameter can change in different iterations. Since the value of the first parameter can change, the dimension of the deep feature vector output by the joint source-channel decoder can change in different iterations. Compared with related technologies that set the deep feature vector output by the joint source-channel decoder as a feature vector of fixed dimension, in this embodiment, since the dimension of the deep feature vector output by the joint source-channel decoder can change, dynamic deep features can be obtained according to the channel environment information conditions, which is beneficial to improving the model's ability to perceive the channel.
[0052] In some embodiments of this disclosure, the objective function used in the joint training of the first neural network model and the second neural network model includes a first parameter, a second parameter, and a third parameter. The first parameter is used to characterize the dimension of the deep feature vector received by the joint source-channel decoder. The second parameter is used to characterize the mutual information between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to characterize the mutual information between the visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder. In this way, during the joint training process, the dimension of the transmitted feature vector can be reduced, while maximizing the mutual information between the received feature vector and the actual ground truth of the multi-task. This results in obtaining the optimal feature vector dimension and balancing the communication perception trade-off, i.e., the trade-off between transmission latency and multi-task performance. This is beneficial for reducing transmission latency while improving task performance.
[0053] In some embodiments of this disclosure, the objective function is used to optimize the first parameter, the second parameter, and the third parameter, wherein the smaller the value of the first parameter, the better the first parameter is; the larger the value of the second parameter, the better the second parameter is; and the larger the value of the third parameter, the better the third parameter is.
[0054] In some embodiments of this disclosure, the objective function is:
[0055] The objective function described above means: by training θ e and θ d To maximize the following function θ e The parameter set of the joint source-channel encoder, θ d The parameter set of the joint source-channel decoder, where λ1 and λ2 are positive hyperparameters of the joint model, respectively. for Expectations For the first parameter, For the second parameter, The third parameter is the one mentioned above, wherein the... Let x be the depth feature vector received by the joint source-channel decoder, y be the data to be transmitted, and y be the visual recognition information output by the joint source-channel decoder.
[0056] In this embodiment, since the first parameter is used to characterize the dimension of the depth feature vector received by the joint source-channel decoder, the smaller the first parameter, the less data is in the depth feature vector, and correspondingly, the lower the network overhead for transmitting the depth feature vector. Therefore, determining the smallest possible first parameter based on the objective function during model training is beneficial for reducing network overhead. Correspondingly, since the second parameter is used to characterize the mutual information between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted, the larger the second parameter, the more image information is contained in the depth information extracted by the second communication device. Therefore, determining the largest possible second parameter based on the objective function during model training is beneficial for optimizing image reconstruction performance. Since the third parameter is used to characterize the mutual information between the visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder, the larger the third parameter, the more visual task-related information is contained in the depth information extracted by the second communication device. Therefore, determining the largest possible third parameter based on the objective function during model training is beneficial for optimizing image classification task performance.
[0057] In some embodiments of this disclosure, before encoding the data to be transmitted based on the first neural network model in the joint source-channel encoder to obtain the first deep feature vector, the method further includes:
[0058] Based on the objective function and the loss function, the joint model is trained to obtain the trained first neural network model and the second neural network model. The loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function, and a fourth sub-loss function. The first sub-loss function is used to minimize the first parameter. The value of is used to characterize the loss between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The value of is used to characterize the discrimination loss of the discriminator. The value of is used to characterize the loss between the visual recognition information output by the joint source-channel decoder and the corresponding real visual result.
[0059] In some embodiments of this disclosure, the first sub-loss function is:
[0060] Where k1, k2, and k3 are constants, and Sig() represents the sigmoid function. σ 2The variance is additive noise, zi is the i-th element of z, z is the depth feature vector obtained after the first neural network model encodes the data to be transmitted, and n is the dimension of z.
[0061] The second sub-loss function is:
[0062] In the formula, E represents the expectation.
[0063] The third sub-loss function is:
[0064] Among them, f dd For the discriminator, L D To identify the loss, E in the formula represents the expectation;
[0065] The fourth sub-loss function is:
[0066] In the formula, E represents the expectation.
[0067] The loss function is:
[0068] Where, θ dd The parameter set of the discriminator is defined as η1 and η2, which are two positive hyperparameters of the joint model.
[0069] It should be noted that the process of training the above joint model may include the following steps:
[0070] Acquire training data, wherein the training data may include: training images and the true visual results of the training images, wherein the true visual results may be the true category of the training images.
[0071] When the first communication device acquires a training image, it can preprocess the training image. The preprocessing means include data augmentation and data broadening. Data augmentation can be performed on the training image by scaling, rotating, or other processing. Data broadening can be performed on the training image by generating an image similar to the training image.
[0072] Then, the preprocessed image data is input into the joint source-channel encoder, and the data to be transmitted is encoded based on the first neural network model in the joint source-channel encoder to obtain the depth feature vector z;
[0073] The joint source-channel encoder transmits the depth feature vector z through a wireless channel to the joint source-channel decoder of the second communication device. The joint source-channel decoder decodes the depth feature vector z to obtain the depth feature vector. in, It is z after being affected by the channel, and its size is the same as z.
[0074] The second neural network model is based on deep feature vectors. Data recovery and visual recognition are performed to obtain the recovered data and predictive visual recognition information.
[0075] Then, the above loss function is constructed based on the recovered data, predicted visual recognition information, training images and the actual visual results of the training images, and the parameters of the joint model are optimized based on the constructed loss function to obtain the optimized joint model.
[0076] Meanwhile, the optimized joint model can be further trained iteratively multiple times based on other training data, following the above method, until the convergence condition is met. The convergence condition can be that the loss value of the loss function is less than a preset value, or that the number of iterations is greater than a preset number.
[0077] To facilitate understanding, the training process of the above joint model will be further explained below with reference to specific embodiments:
[0078] This disclosure addresses the transmission and classification of raw data across multiple modalities, including images. It simulates a general neural network design and transmission scheme under various channel conditions, reducing transmission latency due to reduced communication overhead while ensuring stable performance of parallel AI tasks for image reconstruction and classification. Specifically, it proposes a channel-aware multi-task communication system based on joint source-channel coding. This system adaptively adjusts the dimension of deep features based on channel conditions to select the optimal number of transmitted symbols, without sacrificing task performance or wasting communication resources. Furthermore, it uses reference signals to obtain implicit CSI information to guide the system's multi-task performance under complex channel conditions.
[0079] This disclosure proposes a channel-aware deep joint source-channel coding (CA-DJSCC) multi-task communication system. This system adaptively adjusts the dimension of deep features based on channel conditions to select the optimal number of transmitted symbols, without sacrificing task performance or wasting communication resources. Furthermore, a reference signal is introduced to obtain implicit CSI information, guiding the system's multi-task performance under complex channels. The complete transmission scheme is as follows: the raw data from the first communication device of the mobile device is preprocessed (data augmentation, data broadening, etc.) and input into the joint source-channel encoder to obtain a deep feature vector. Then, the deep feature vector is input into the joint source-channel decoder deployed on the second communication device of the server to obtain the recovered transmitted data and the result after performing the visual task. The system flowchart is shown in Figure 2. In some embodiments of this disclosure, the raw data can be the aforementioned training images.
[0080] The flowchart shown in Figure 2 illustrates three key steps in the detailed design of the CA-DJSCC system: first, the design of a system model and optimization objectives to balance task performance and communication latency, achieving a trade-off between multi-task performance and communication latency; second, the design of a semantic encoder based on a joint source channel; and third, the design of a semantic decoder based on a joint source channel. The joint source channel encoder adaptively adjusts the dimensions of the transmitted semantic information according to channel conditions, while the joint source channel decoder facilitates data recovery and visual reasoning with the help of implicit versions of additional reference signals known to the transceiver.
[0081] (I) System Model and Optimization Objective Design for Task Performance and Communication Delay Balancing
[0082] This disclosure proposes a multi-task-oriented channel-aware communication scheme based on a joint source-channel codec, where channel awareness refers to the trade-off between multi-task performance and communication latency based on specific channel conditions. The system model is shown in Figure 3.
[0083] Where, x∈R H×W×C The original data is represented by H, W, and C, which are the length, width, and number of channels of the image data, respectively. Taking this embodiment as an example, x is a survey and design image, including seven categories: environment, computer room, switching power supply, power distribution equipment, rooftop, battery, and others. The signal-to-noise ratio (SNR) is the signal-to-noise ratio of the additive noise of the channel, and its definition is as follows:
[0084] Where P z It is the power of the depth characteristic z of the input wireless channel, i.e., the signal power; P n That is noise power. fe It is a joint source-channel encoder. r∈R H×W×1 This is an additional reference signal added to the joint source-channel encoder, and it is known to both the first and second communication devices. z∈C n It is an adaptive-dimensional deep feature vector extracted by the joint source-channel encoder after fusing the original data x, reference signal r, and channel condition guidance factor SNR in the deep feature space, where n is the dimension of z, and the value of n is different for different SNR and x, expressed by the following formula: z = f e (x,r,SNR;θ e (2)
[0085] Where, θ e It is f e The parameter set. This is z after being affected by the channel, and its size is the same as z, which can be expressed by the following formula:
[0086] In some embodiments of this disclosure, the second communication device is used to perform data recovery and visual recognition on the first deep feature vector using a second neural network model in a joint source-channel decoder, to obtain recovered data and visual recognition information, including:
[0087] The second communication device is used to decrypt the first deep feature vector through a joint source-channel decoder to obtain a second deep feature vector, wherein the second deep feature vector is the first deep feature vector after being affected by the channel, and the dimension of the first deep feature vector is the same as the dimension of the second deep feature vector.
[0088] The second communication device is used to perform data recovery and visual recognition on the second deep feature vector through the second neural network model to obtain the recovered data and visual recognition information.
[0089] The second communication device is used to decrypt the first deep feature vector using a joint source-channel decoder to obtain the second deep feature vector, which can be achieved based on the following formula:
[0090] in, z is the second depth feature vector, z is the first depth feature vector, and h∈C n It is the channel response, n∈C n It is additive noise. d It is a joint source-channel decoder. This is the data recovered by the second communication device, and its size is the same as x. It is the result of a visual task performed by the second communication device. The size of the reference signal is related to the task performance. In this case, Because there are seven categories of image labels, namely the aforementioned "Environment, Computer Room, Switching Power Supply, Power Distribution Equipment, Rooftop, Battery, and Others," they can be represented by the following formula:
[0091] The embodiments disclosed herein can be applied to other multitasking tasks, including image classification, image retrieval, image restoration, semantic segmentation, and object detection. Appropriate updates to the network structure can also perform corresponding tasks on raw data such as text, audio, and video.
[0092] The goal of the CA-DJSCC proposed in this disclosure is to optimize the dimension of transmitted deep features under dynamically changing channel conditions, thereby improving the performance of data recovery and visual inference while minimizing transmission latency. To achieve this channel-aware tradeoff, the dimension of z is optimized and... The mutual information between the objective function and the actual truth values of multiple tasks characterizes the transmission delay and task performance. The optimization objective function is expressed as follows:
[0093] in, express The expectation, D() is to obtain The dimension operation, λ1 and λ2 are two positive hyperparameters. The first term... That is, the average value of the depth feature dimension n obtained by the second communication device. The smaller this value, the smaller the number of transmitted symbols, and the lower the communication latency; the second term This is the mutual information between the actual ground truth data and the depth features reconstructed from the image by the second communication device. The larger this mutual information is, the more image information is contained in the depth information extracted by the second communication device, which is used to optimize the performance of image reconstruction; the third item This refers to the mutual information between the actual ground truth labels and the depth features recovered by the second communication device. The larger the mutual information, the more visual task-related information is contained in the depth information extracted by the second communication device, which is used to optimize the performance of the image classification task. Through end-to-end joint training and tuning of hyperparameters λ1 and λ2, a trade-off between communication latency and multi-task performance can be achieved to obtain the optimal transmission depth feature dimension and achieve the lowest transmission latency under the requirements of multi-task performance.
[0094] As shown in Equation (5), the proposed CA-DJSCC aims to reduce the dimensionality of the transmitted feature vectors while maximizing the mutual information between the received feature vectors and the actual ground truth of the multi-task, thereby obtaining the optimal feature vector dimensionality and balancing the communication-aware trade-off, namely the trade-off between transmission latency and multi-task performance. To achieve the optimization objective of (5), a loss function is designed, including three parts: dimensionality adjustment, image reconstruction, and classification inference.
[0095] (1) Dimensional adjustment: The goal is to minimize the transmitted deep feature vector. The dimension of . The corresponding loss function is derived from equation (5), and is expressed as:
[0096] Where k1, k2, and k3 are constants, and Sig() represents the sigmoid function. σ 2 For additive noise variance, z i Let z be the i-th element.
[0097] (2) Image reconstruction: Deep feature vector exist The maximization process maps to the reconstructed image. Typically, Maximizing this corresponds to minimizing the MSE loss function, which measures the similarity between the actual ground truth image and the restored image, and is expressed by the following formula:
[0098] Because this embodiment of the disclosure utilizes generative adversarial networks to obtain... Following the training strategy of generative adversarial networks, adversarial training of the generator and discriminator requires an additional discriminative loss, expressed by the following formula:
[0099] The generator fdg is trained to minimize the recovery loss. and identification loss L D To confuse the recognition results as much as possible, the discriminator fdd is trained to maximize the discriminative loss L. D To distinguish the restored image from the original image as much as possible. Furthermore, it is necessary to minimize... To restore the reference signal.
[0100] (3) Classification reasoning: deep feature vectors exist The maximization process maps to visual reasoning labels. Typically, Maximizing this corresponds to minimizing the cross-entropy loss function, which measures the similarity between the actual ground truth label and the predicted label in a visual classification inference task, and is expressed by the following formula:
[0101] The loss function obtained by weighting equations (7), (8), (9), and (10) is expressed as follows:
[0102] η1 and η2 are two positive hyperparameters.
[0103] (II) Design of a Semantic Encoder Based on a Joint Source Channel
[0104] Optionally, the step of encoding the data to be transmitted using the first neural network model in the joint source-channel encoder to obtain the first deep feature vector of the data to be transmitted includes:
[0105] The data to be transmitted, the preset reference signal, and the signal-to-noise ratio are fused to obtain first fused characteristic information. The first communication device and the second communication device both store the preset reference signal. The signal-to-noise ratio is the signal-to-noise ratio of the additive noise of the channel used to transmit the first depth feature vector.
[0106] Generate an initial semantic information vector corresponding to the first fusion characteristic information, and generate a semantic importance vector corresponding to the first fusion characteristic information, wherein the initial semantic information vector is used to characterize the semantics of the first fusion characteristic information, and the semantic importance vector is used to characterize the semantic importance of each piece of information in the first fusion characteristic information;
[0107] The target semantic importance vector is obtained by multiplying the semantic importance vector with an upper triangular matrix composed of 0s and 1s, where the upper triangular matrix has a dimension of n0.
[0108] A mask vector corresponding to the target semantic importance vector is generated based on a threshold, wherein the i-th element in the mask vector is 1 if the i-th element in the target semantic importance vector is greater than the threshold; and the i-th element in the mask vector is 0 if the i-th element in the target semantic importance vector is less than or equal to the threshold, where i = 1, ..., n0;
[0109] The first depth feature vector is obtained by multiplying the initial semantic information vector and the mask vector.
[0110] Among them, the joint source channel encoder is one of the two core modules in the CA-DJSCC designed in this embodiment. Its functions are as follows: (1) Introducing a reference signal to obtain implicit CSI for the joint source channel decoder of the second communication device; (2) Designing an adaptive dimension pruning method, which can adaptively change the dimension of the transmission depth feature according to different original data and channel conditions, thereby adaptively adjusting the number of transmission symbols, allocating the best communication overhead under certain channel conditions for multi-task communication, achieving a trade-off between communication delay and multi-task performance, and maintaining good AI multi-task performance while reducing communication overhead; (3) When transmission resources are limited, the system can adjust the transmission overhead to provide communication services at the cost of sacrificing some multi-task performance.
[0111] Joint source-channel encoder fe The structure is shown in Figure 4.
[0112] First, an initial semantic information vector is extracted, and then its dimensions are selectively pruned based on the inherent semantic importance. Specifically, the original data x, the reference signal r, and the signal-to-noise ratio (SNR) are fused in the latent space to generate the initial semantic information vector z. 0 And semantically important vector ξ. 0 Both ξ and z have a predetermined dimension n0, which is greater than the dimension n of the final extracted semantic information. Physically, ξ measures z. 0 The impact of each element in the equation on multitasking performance. Consider ξ and the semantic importance threshold ξ. thr The embodiments disclosed herein can be used in z 0 Elements with lower semantic importance are pruned, and finally, z of different dimensions are obtained based on the channel conditions.
[0113] Joint source-channel encoder f e The network structure and workflow are shown in Figure 3. First, the original data x and the reference signal r are merged on the third axis and input into the deep network layer (1) as transmission information. At the same time, the signal-to-noise ratio (SNR) is input into the deep network layer (2) as a channel condition. This is to integrate the semantic information from the original image and the reference signal and incorporate the channel condition into the dimensionality clipping of the semantic information, i.e., the transmission depth features. In addition, the semantic importance vector ξ measures the degree of multi-task relevance in the semantic information. It is only related to the transmission context and the channel condition, which are used to guide dimensionality adjustment and multi-task adjustment. Therefore, by end-to-end joint training of the multi-task execution part in the decoder, it can be learned through the fusion features of the original data x, the reference signal r, and the signal-to-noise ratio (SNR). Therefore, the latent feature vectors obtained from layer (1) and layer (2) are merged on the first axis to obtain a new feature map, and then input into the deep network layer (3) and the deep network layer (4) respectively to obtain the initial semantic information vector z0 and the semantic importance vector ξ respectively. By multiplying the output of layer(3) by an upper triangular n0-dimensional matrix consisting of 1s and 0s, the elements in ξ are in reverse order. Using the obtained ξ, layer(4) is used to generate z0, as expressed by the following equation:
[0114] Where ξ i It is the i-th element of ξ, i = 1, ..., n0. v and v are the last input of size n from layer(4). inThe output is a stacked parameter matrix of a linear layer of size n0 and a stacked input feature vector. For ease of computation, the weight matrix and bias vector are stacked along the second axis, and the input feature vectors are stacked by 1 to match the augmented form of the weight matrix and bias. yes The vector in the i-th row is used. Simultaneously, a mask vector m is generated from ξ based on the threshold ξthr. The i-th element of m is determined by ξi and ξthr. If ξ > ξthr, it is set to 1; otherwise, it is set to 0. Finally, the elements of m and z0 are multiplied together to obtain z.
[0115] In terms of network structure, fully connected layers and residual blocks are mainly used. The convolutional kernels generally use 3×3 kernels with a stride of 1 and a padding size of 1 to more comprehensively traverse the map features and extract more precise feature maps. If the computing power is limited or the parameter deployment pressure of the mobile device is high, a larger convolutional kernel can be replaced to reduce the number of calculations and save costs. Except for layer (4), the activation layers all use simple ReLU to extract nonlinear features. Layer (4) uses tanh to limit the element size of the depth feature vector. In addition, the normalization method of BatchNorm is used to speed up the training speed and normalize the size of the transmission feature vector, that is, the signal power is 1, so as to more conveniently calculate the relationship between the signal-to-noise ratio and the variance of the channel additive noise. The specific structure of the network is shown in the table below:
[0116] Table 1 Network Structure Parameters
[0117] This part combines the principles of signal-to-noise ratio and semantic importance to extract deep features of different dimensions and transmit them under different channel conditions. This joint source-channel encoder starts from the original image, guided by channel conditions, and dynamically crops deep features with low multi-task relevance based on semantic importance, achieving optimal dimension selection and a trade-off between performance and communication latency. Furthermore, it consists of deep network structures capable of joint training and backpropagation.
[0118] (III) Design of a semantic decoder based on a joint source channel
[0119] Joint source channel decoding is another core module in the CA-DJSCC designed in this embodiment. Its functions are as follows: (1) By introducing a reference signal, the known reference signal at both ends is recovered at the recovery end. Since the reference signal and the original data are affected by the same channel, the implicit channel state information can be learned through a deep neural network to guide the execution of multiple tasks. This approach is similar to channel estimation using pilot signals. Its advantage is that it is integrated into a neural network that can be jointly trained and backpropagated. No additional channel estimation is required. It also has an irreplaceable effect in some scenarios where perfect CSI cannot be obtained through channel estimation. (2) Generative adversarial networks are introduced, which can improve the performance of data recovery in multiple tasks. The discriminator in the generative adversarial network can be discarded after training. Therefore, it does not need to be deployed in a second communication device, which can reduce deployment overhead and resources.
[0120] Joint source-channel decoder f d The structure is shown in Figure 5.
[0121] The network structure and workflow of the joint source-channel decoder are shown in Figure 5. This involves receiving semantic information affected by the channel. Recovery of reference signal and from The implicit CSI, denoted by Δr, is obtained from the residual between the original reference signal r and the original reference signal r. Then, Δr is compared with... The extracted deep features are fused, and under the guidance of a reference signal, an enhanced version of the received semantic information, ZF, is obtained. Finally, an adversarial generative network is introduced to obtain high-quality image reconstruction and classification reasoning from ZF to perform visual tasks.
[0122] Specifically, firstly, using f dr from China Resumption Then from The residual between r and Δr is calculated. Since the reference signal r and the original data experience the same channel conditions, Δr contains an implicit version of CSI. Then, Δr and Δr are processed through layer (5) from... The extracted deep features are merged on the third axis to obtain an enhanced version z of the received semantic information. f Guided by the implicit version of CSI obtained from Δr, z f It can better perform multitasking. Finally, z f Input to generator f dg and discriminator f dd The GAN and reasoning module f di In order to obtain the recovered data And the results of visual reasoning.dd It is an auxiliary discriminator that exists independently of the decoder, so f dd The parameter set is not included in θ d In the middle, use θ dd express.
[0123] In terms of network structure, fully connected layers and ResBlock residual blocks are mainly used. The convolutional kernels are generally 3×3 with a stride of 1 and padding size of 1. The activation layers all use simple ReLU, and BatchNorm normalization is employed to accelerate training. The specific network structure is shown in the table below:
[0124] Table 2 Network Structure Parameters
[0125] Where 512×512×3 is the size of the survey training image in this embodiment of the disclosure, and the size can be changed according to the dataset size and modality. di The output size 7 represents the number of categories, which can be adjusted according to the task.
[0126] Through the above workflow, the implicit version of CSI obtained through reference signal recovery and residual calculation enhances the reception capability of deep features, providing effective guidance for data recovery and visual reasoning. Furthermore, all of these methods consist of deep network structures capable of joint training and backpropagation.
[0127] The method provided in this disclosure also has at least the following beneficial effects:
[0128] Multi-task communication based on joint source-channel coding and decoding in related technologies lacks channel awareness and is prone to performance degradation or communication overhead waste due to its fixed-dimensional design of transmission depth features. Furthermore, the lack of additional CSI acquisition limits its application to simple communication environments, making it unsuitable for more complex communication channels. The embodiments of this disclosure adaptively adjust the dimensions of the depth features based on channel conditions to select the optimal number of transmitted symbols, without sacrificing task performance or wasting communication resources. Moreover, implicit CSI information obtained using reference signals guides the system's multi-task performance under complex channels.
[0129] For 6G satellite-to-ground coverage scenarios facing constraints such as high dynamics, long distances, time-varying characteristics, and limited resources, achieving efficient information extraction, reliable information recovery, and seamless network coverage extension based on semantic communication is an effective approach. This disclosure addresses the transmission and classification of raw data across multiple modalities, including images. It simulates a general neural network design and transmission scheme under non-ideal channel conditions, improving image reconstruction and feature fusion performance. While reducing communication overhead, it maintains good data reception and recovery performance and AI task capabilities, expanding efficient information extraction and reliable information recovery schemes for semantic communication in satellite-to-ground transmission. This enhances satellite-to-ground transmission capabilities and has broad application prospects.
[0130] Please refer to Figure 6. This embodiment of the disclosure also provides a data transmission method applied to a second communication device, the method comprising:
[0131] Step 601: Receive the first depth feature vector sent by the first communication device, wherein the first depth feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on the first neural network model in the joint source-channel encoder;
[0132] Step 602: Perform data recovery and visual recognition on the first deep feature vector based on the second neural network model in the joint source-channel decoder to obtain the recovered data and visual recognition information. The first neural network model and the second neural network model are joint models obtained by joint training based on the objective function.
[0133] This implementation method is a method on the second communication device side corresponding to the above embodiments. Its specific implementation process corresponds to the above embodiments and has corresponding beneficial effects. To avoid repetition, it will not be described again here.
[0134] Please refer to Figure 7, which illustrates a data transmission device 700 provided in an embodiment of this disclosure, applied to a first communication device. The data transmission device 700 includes:
[0135] The encoding module 701 is used to encode the data to be transmitted through the first neural network model in the joint source-channel encoder to obtain the first depth feature vector of the data to be transmitted.
[0136] Transmission module 702 is used to transmit the first depth feature vector to the second communication device; wherein,
[0137] The second communication device is used to perform data recovery and visual recognition on the first deep feature vector through the second neural network model in the joint source-channel decoder, so as to obtain the recovered data and visual recognition information;
[0138] The first neural network model and the second neural network model are a joint model obtained by joint training based on the objective function.
[0139] Optionally, the objective function is a function constructed based on a first parameter, a second parameter, and a third parameter. The first parameter is used to characterize the dimension of the depth feature vector received by the joint source-channel decoder. The second parameter is used to characterize the mutual information between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to characterize the mutual information between the visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder.
[0140] Optionally, the objective function is used to optimize the first parameter, the second parameter, and the third parameter, wherein the smaller the value of the first parameter, the better the first parameter is; the larger the value of the second parameter, the better the second parameter is; and the larger the value of the third parameter, the better the third parameter is.
[0141] Optionally, the objective function is:
[0142] Where, θ e The parameter set of the joint source-channel encoder, θ d The parameter set of the joint source-channel decoder, where λ1 and λ2 are positive hyperparameters of the joint model, respectively. for Expectations For the first parameter, For the second parameter, The third parameter is the one mentioned above, wherein the... Let x be the deep feature vector received by the joint source-channel decoder, y be the data to be transmitted, and y be the visual recognition information output by the joint source-channel decoder. The joint model is a joint model formed by combining the first neural network model and the second neural network model.
[0143] Optionally, the data transmission device 700 further includes:
[0144] The training module is used to train the joint model based on the objective function and the loss function to obtain the trained first neural network model and the second neural network model. The loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function, and a fourth sub-loss function. The first sub-loss function is used to minimize the value of the first parameter. The second sub-loss function is used to characterize the loss between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third sub-loss function is used to characterize the discrimination loss of the discriminator. The fourth sub-loss function is used to characterize the loss between the visual recognition information output by the joint source-channel decoder and the corresponding real visual result.
[0145] Optionally, the first sub-loss function is:
[0146] Where k1, k2, and k3 are constants, and Sig() represents the sigmoid function. σ 2 The variance is additive noise, zi is the i-th element of z, z is the depth feature vector obtained after the first neural network model encodes the data to be transmitted, and n is the dimension of z.
[0147] The second sub-loss function is:
[0148] The third sub-loss function is:
[0149] Among them, f dd For the discriminator, L D To identify the loss;
[0150] The fourth sub-loss function is:
[0151] The loss function is:
[0152] Where, θ dd The parameter set of the discriminator is defined as η1 and η2, which are two positive hyperparameters of the joint model.
[0153] Optionally, the second communication device is used to perform data recovery and visual recognition on the first deep feature vector using a second neural network model in the joint source-channel decoder, to obtain recovered data and visual recognition information, including:
[0154] The second communication device is used to decrypt the first deep feature vector through a joint source-channel decoder to obtain a second deep feature vector, wherein the second deep feature vector is the first deep feature vector after being affected by the channel, and the dimension of the first deep feature vector is the same as the dimension of the second deep feature vector.
[0155] The second communication device is used to perform data recovery and visual recognition on the second deep feature vector through the second neural network model to obtain the recovered data and visual recognition information.
[0156] Optionally, the encoding module 701 is specifically used to: perform feature fusion on the data to be transmitted, the preset reference signal and the signal-to-noise ratio to obtain first fusion characteristic information, wherein the first communication device and the second communication device both store the preset reference signal, and the signal-to-noise ratio is the signal-to-noise ratio of the additive noise of the channel used to transmit the first depth feature vector;
[0157] Generate an initial semantic information vector corresponding to the first fusion characteristic information, and generate a semantic importance vector corresponding to the first fusion characteristic information, wherein the initial semantic information vector is used to characterize the semantics of the first fusion characteristic information, and the semantic importance vector is used to characterize the semantic importance of each piece of information in the first fusion characteristic information;
[0158] The target semantic importance vector is obtained by multiplying the semantic importance vector with an upper triangular matrix composed of 0s and 1s, where the upper triangular matrix has a dimension of n0.
[0159] A mask vector corresponding to the target semantic importance vector is generated based on a threshold, wherein the i-th element in the mask vector is 1 if the i-th element in the target semantic importance vector is greater than the threshold; and the i-th element in the mask vector is 0 if the i-th element in the target semantic importance vector is less than or equal to the threshold, where i = 1, ..., n0;
[0160] The first depth feature vector is obtained by multiplying the initial semantic information vector and the mask vector.
[0161] It should be noted that the data transmission device 700 provided in this embodiment is a device capable of executing the data transmission method described in FIG1. All implementations of the data transmission method in the above embodiments are applicable to this device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0162] Please refer to Figure 8, which illustrates a data transmission device 800 provided in an embodiment of this disclosure, applied to a second communication device. The data transmission device 800 includes:
[0163] The receiving module 801 is used to receive a first depth feature vector sent by the first communication device, wherein the first depth feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on the first neural network model in the joint source-channel encoder;
[0164] The processing module 802 is used to perform data recovery and visual recognition on the first deep feature vector based on the second neural network model in the joint source-channel decoder, to obtain the recovered data and visual recognition information, wherein the first neural network model and the second neural network model are a joint model obtained by joint training based on the objective function.
[0165] It should be noted that the data transmission device 800 provided in this embodiment is a device capable of executing the data transmission method described in FIG6. All implementations of the above data transmission method embodiments are applicable to this device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0166] Specifically, as shown in Figure 9, this disclosure also provides an electronic device, including a bus 901, a transceiver 902, an antenna 903, a bus interface 904, a processor 905, and a memory 906.
[0167] Processor 905, used for:
[0168] The first deep feature vector of the data to be transmitted is obtained by encoding the data to be transmitted through the first neural network model in the joint source-channel encoder;
[0169] The first depth feature vector is transmitted to the second communication device; wherein...
[0170] The second communication device is used to perform data recovery and visual recognition on the first deep feature vector through the second neural network model in the joint source-channel decoder, so as to obtain the recovered data and visual recognition information;
[0171] The first neural network model and the second neural network model are a joint model obtained by joint training based on the objective function.
[0172] In Figure 9, a bus architecture (represented by bus 901) is shown. Bus 901 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 905 and memory represented by memory 906. Bus 901 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 904 provides an interface between bus 901 and transceiver 902. Transceiver 902 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 905 is transmitted over a wireless medium via antenna 903, which further receives data and transmits it to processor 905.
[0173] Processor 905 manages bus 901 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 906 can be used to store data used by processor 905 during operation.
[0174] Optionally, the processor 905 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).
[0175] Optionally, the objective function is a function constructed based on a first parameter, a second parameter, and a third parameter. The first parameter is used to characterize the dimension of the depth feature vector received by the joint source-channel decoder. The second parameter is used to characterize the mutual information between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to characterize the mutual information between the visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder.
[0176] Optionally, the objective function is used to optimize the first parameter, the second parameter, and the third parameter, wherein the smaller the value of the first parameter, the better the first parameter is; the larger the value of the second parameter, the better the second parameter is; and the larger the value of the third parameter, the better the third parameter is.
[0177] Optionally, the objective function is:
[0178] Where, θ e The parameter set of the joint source-channel encoder, θ d The parameter set of the joint source-channel decoder, where λ1 and λ2 are positive hyperparameters of the joint model, respectively. for Expectations For the first parameter, For the second parameter, The third parameter is the one mentioned above, wherein the... Let x be the deep feature vector received by the joint source-channel decoder, y be the data to be transmitted, and y be the visual recognition information output by the joint source-channel decoder. The joint model is a joint model formed by combining the first neural network model and the second neural network model.
[0179] Optionally, the processor 905 is further configured to:
[0180] Based on the objective function and the loss function, the joint model is trained to obtain the trained first neural network model and the second neural network model. The loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function, and a fourth sub-loss function. The first sub-loss function is used to minimize the value of the first parameter. The second sub-loss function is used to characterize the loss between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third sub-loss function is used to characterize the discrimination loss of the discriminator. The fourth sub-loss function is used to characterize the loss between the visual recognition information output by the joint source-channel decoder and the corresponding real visual result.
[0181] Optionally, the first sub-loss function is:
[0182] Where k1, k2, and k3 are constants, and Sig() represents the sigmoid function. σ 2 The variance is additive noise, zi is the i-th element of z, z is the depth feature vector obtained after the first neural network model encodes the data to be transmitted, and n is the dimension of z.
[0183] The second sub-loss function is:
[0184] The third sub-loss function is:
[0185] Among them, fdd For the discriminator, L D To identify the loss;
[0186] The fourth sub-loss function is:
[0187] The loss function is:
[0188] Where, θ dd The parameter set of the discriminator is defined as η1 and η2, which are two positive hyperparameters of the joint model.
[0189] Optionally, the second communication device is used to perform data recovery and visual recognition on the first deep feature vector using a second neural network model in the joint source-channel decoder, to obtain recovered data and visual recognition information, including:
[0190] The second communication device is used to decrypt the first deep feature vector through a joint source-channel decoder to obtain a second deep feature vector, wherein the second deep feature vector is the first deep feature vector after being affected by the channel, and the dimension of the first deep feature vector is the same as the dimension of the second deep feature vector.
[0191] The second communication device is used to perform data recovery and visual recognition on the second deep feature vector through the second neural network model to obtain the recovered data and visual recognition information.
[0192] Optionally, the processor 905 is specifically used for:
[0193] The data to be transmitted, the preset reference signal, and the signal-to-noise ratio are fused to obtain first fused characteristic information. The first communication device and the second communication device both store the preset reference signal. The signal-to-noise ratio is the signal-to-noise ratio of the additive noise of the channel used to transmit the first depth feature vector.
[0194] Generate an initial semantic information vector corresponding to the first fusion characteristic information, and generate a semantic importance vector corresponding to the first fusion characteristic information, wherein the initial semantic information vector is used to characterize the semantics of the first fusion characteristic information, and the semantic importance vector is used to characterize the semantic importance of each piece of information in the first fusion characteristic information;
[0195] The target semantic importance vector is obtained by multiplying the semantic importance vector with an upper triangular matrix composed of 0s and 1s, where the upper triangular matrix has a dimension of n0.
[0196] A mask vector corresponding to the target semantic importance vector is generated based on a threshold, wherein the i-th element in the mask vector is 1 if the i-th element in the target semantic importance vector is greater than the threshold; and the i-th element in the mask vector is 0 if the i-th element in the target semantic importance vector is less than or equal to the threshold, where i = 1, ..., n0;
[0197] The first depth feature vector is obtained by multiplying the initial semantic information vector and the mask vector.
[0198] It should be noted that the electronic device provided in this embodiment is a device capable of executing the above-described data transmission method. Therefore, all implementations of the above-described data transmission method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.
[0199] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described data transmission method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0200] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described data transmission method embodiments and achieves the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0201] This disclosure also provides a computer program product, including computer instructions. When executed by a processor, these computer instructions implement the various processes of the above-described data transmission method embodiments and achieve the same technical effects. To avoid repetition, further details are omitted here.
[0202] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0203] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0204] The embodiments of this disclosure have been described above with reference to the accompanying drawings. However, this disclosure is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this disclosure without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this disclosure.
Claims
1. A data transmission method applied to a first communication device, the method comprising: encoding, by a first neural network model in a joint source-channel encoder, to-be-transmitted data to obtain a first deep feature vector of the to-be-transmitted data; transmitting the first deep feature vector to a second communication device; wherein the second communication device is configured to perform data recovery and visual recognition on the first deep feature vector by a second neural network model in a joint source-channel decoder to obtain recovered data and visual recognition information; the first neural network model and the second neural network model are joint models obtained based on a target function.
2. The method of claim 1, wherein, the target function is a function constructed based on a first parameter, a second parameter and a third parameter, the first parameter characterizes a dimension number of a deep feature vector received by the joint source-channel decoder, the second parameter is used to characterize mutual information between recovered data output by the joint source-channel decoder and corresponding to-be-transmitted data, and the third parameter is used to characterize mutual information between visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder.
3. The method of claim 2, wherein, The value of the first parameter of the target function is negatively correlated with the communication time delay, the value of the second parameter is positively correlated with the image reconstruction performance, and the value of the third parameter is positively correlated with the image classification performance.
4. The method of any one of claims 2-3, wherein, The objective function is: wherein θ e is a parameter set of the joint source-channel encoder, and θ d is a parameter set of the joint source-channel decoder, and λ1 and λ2 are positive hyperparameters of the joint model, respectively, and For the expectation of the user, for the first parameter, for the second parameter, for the third parameter, wherein the For the deep feature vector received by the joint source-channel decoder, x is the to-be-transmitted data, and y is the visual recognition information output by the joint source-channel decoder.
5. The method of claim 4, wherein, Before the to-be-transmitted data is encoded by the first neural network model in the joint source-channel encoder to obtain the deep feature vector, the method further comprises: training the joint model based on the target function and a loss function to obtain the trained first neural network model and the second neural network model, wherein the loss function comprises a first sub-loss function, a second sub-loss function, a third sub-loss function and a fourth sub-loss function, the first sub-loss function is used to minimize the value of the first parameter, the second sub-loss function is used to characterize the loss between the recovered data output by the joint source-channel decoder and the corresponding to-be-transmitted data, the third sub-loss function is used to characterize the discrimination loss of the discriminator, and the fourth sub-loss function is used to characterize the loss between the visual recognition information output by the joint source-channel decoder and the corresponding real visual result.
6. The method of claim 5, wherein, The first sub-loss function is: where k1, k2, and k3 are constants, and Sig( ) denotes a sigmoid function, σ 2 is the additive noise variance, zi is the i-th element of z, z is the deep feature vector obtained after the first neural network model encodes the data to be transmitted, and n is the dimension number of z. The second sub-loss function is: The third sub-loss function is: wherein f dd is a discriminator, L D is a discrimination loss; The fourth sub-loss function is: The loss function is: where θ dd is a parameter set of discriminator, and η1 and η2 are two positive hyperparameters of the joint model, respectively.
7. The method of claim 1, wherein, The second communication device is configured to perform data recovery and visual recognition on the first deep feature vector by the second neural network model in the joint source-channel decoder to obtain recovered data and visual recognition information, comprising: the second communication device is configured to decrypt the first deep feature vector by a joint source-channel decoder to obtain a second deep feature vector, wherein the second deep feature vector is the first deep feature vector affected by a channel, and the dimension number of the first deep feature vector is the same as that of the second deep feature vector; The second communication device is configured to perform data recovery and visual recognition on the second deep feature vector by using the second neural network model, to obtain recovered data and visual recognition information.
8. The method of claim 1, wherein, The first neural network model in the joint source-channel encoder is used to encode the to-be-transmitted data, to obtain the first deep feature vector of the to-be-transmitted data, including: performing feature fusion on the to-be-transmitted data, a preset reference signal and a signal-to-noise ratio to obtain first fusion characteristic information, wherein the first communication device and the second communication device both store the preset reference signal, and the signal-to-noise ratio is a signal-to-noise ratio of additive noise of a channel used to transmit the first deep feature vector; generating an initial semantic information vector corresponding to the first fusion characteristic information, and generating a semantic importance vector corresponding to the first fusion characteristic information, wherein the initial semantic information vector is used to represent semantics of the first fusion characteristic information, and the semantic importance vector is used to represent semantic importance of each item of information in the first fusion characteristic information; calculating a product of the semantic importance vector and an upper triangular matrix composed of 0 and 1 to obtain a target semantic importance vector, wherein a dimension number of the upper triangular matrix is n0; generating a mask vector corresponding to the target semantic importance vector according to a threshold value, wherein in a case where an i-th element in the target semantic importance vector is greater than the threshold value, an i-th element in the mask vector is 1; in a case where the i-th element in the target semantic importance vector is less than or equal to the threshold value, the i-th element in the mask vector is 0, and the i = 1, …, n0; calculating a product of the initial semantic information vector and the mask vector to obtain the first deep feature vector.
9. A data transmission method applied to a second communication device, the method comprising: receiving a first deep feature vector sent by a first communication device, wherein the first deep feature vector is a vector obtained by encoding to-be-transmitted data based on a first neural network model in a joint source-channel encoder by the first communication device; performing data recovery and visual recognition on the first deep feature vector based on a second neural network model in a joint source-channel decoder to obtain recovered data and visual recognition information, wherein the first neural network model and the second neural network model are joint models obtained by joint training based on an objective function.
10. A data transmission device applied to a first communication device, the data transmission device comprising: an encoding module configured to encode to-be-transmitted data by using a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the to-be-transmitted data; a transmission module configured to transmit the first deep feature vector to a second communication device; wherein the second communication device is configured to perform data recovery and visual recognition on the first deep feature vector by using a second neural network model in a joint source-channel decoder to obtain recovered data and visual recognition information; the first neural network model and the second neural network model are joint models obtained by joint training based on the objective function. 11.A data transmission apparatus applied to a second communication device, the apparatus comprising: a receiving module configured to receive a first deep feature vector transmitted by a first communication device, wherein the first deep feature vector is obtained by encoding data to be transmitted by the first communication device based on a first neural network model in a joint source-channel encoder; a processing module configured to perform data recovery and visual recognition on the first deep feature vector based on a second neural network model in a joint source-channel decoder, to obtain recovered data and visual recognition information, wherein the first neural network model and the second neural network model are jointly trained based on an objective function.
12. An electronic device comprising: a processor, a memory, and a program stored in the memory and executable in the processor, the program, when executed by the processor, implements the steps of the data transmission method according to any one of claims 1 to 9. 13.A computer readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implements the steps of the data transmission method according to any one of claims 1 to 9. 14.A computer program product comprising computer instructions, the computer instructions, when executed by a processor, implements the steps of the data transmission method according to any one of claims 1 to 9.