Data transmission method and device, electronic equipment, medium and program product
By using the joint training neural network model in the joint source channel encoder and decoder of AI multitasking communication, dynamically adjusting the dimensions of depth features, solving the problem of insufficient channel perception capabilities in the prior art, improving task performance and reducing communication overhead.
Patent Information
- Application Number
- CN202510370852.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-24
AI Technical Summary
Existing AI multitasking communications are based on joint source channel encoding, but lack channel awareness, which may lead to task performance degradation or waste of communication overhead when transmitting deep features.
By using a neural network model based on the objective function in the joint source channel encoder and decoder, the data to be transmitted is encoded and decoded, and the dimensions of depth features are dynamically adjusted to adapt to channel conditions.
Improve task performance and reduce communication overhead, enhance the model's channel perception ability, and adapt to complex channel conditions.
Smart Images

Figure CN120200653A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data transmission, and in particular, to a data transmission method, apparatus, electronic device, medium, and program product. Background Art
[0002] In view of the limited transmission conditions such as high dynamics, long distance, time-varying, and limited resources in the satellite-ground coverage scenario, it is necessary to efficiently extract information, reliably recover information, and expand seamless network coverage based on joint source-channel coding. Designing an efficient and flexible adaptive network resource allocation model for the satellite-ground transmission network architecture for parallel artificial intelligence (AI) multi-tasks and exploring technical solutions for efficient information extraction and reliable information reconstruction are of great significance for enhancing satellite-ground transmission capabilities.
[0003] Existing AI multi-task communication is mainly based on joint source-channel coding. A common method is to design a joint source-channel encoder at the sending end to learn and extract the original data provided by the mobile device, such as images, videos, voices, texts, etc. into deep features in the latent space, and then transmit this deep feature through the wireless channel. Correspondingly, a joint source-channel decoder is designed at the receiving end to receive this deep feature affected by the wireless channel at the server side, and a corresponding task execution network is designed to obtain the task result, and then the result is transmitted back to the mobile device. The specific network structure is mainly convolutional neural networks (CNNs) and transformers.
[0004] Since the multi-task communication based on joint source-channel coding and decoding in the related art has no perception ability for the channel, the fixed-dimension design of the transmitted deep feature may lead to problems such as task performance degradation or communication overhead waste. Summary of the Invention
[0005] Embodiments of this application provide a data transmission method, apparatus, electronic device, medium, and program product, which are beneficial to improving task performance and reducing communication overhead.
[0006] In a first aspect, an embodiment of this application provides a data transmission method, which is applied to a first communication device, and the method includes:
[0007] Encoding the data to be transmitted through a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the data to be transmitted;
[0008] Transmitting the first deep feature vector to a second communication device; where
[0009] The second communication device is configured to perform data recovery and visual recognition on the first deep feature vector through a second neural network model in a joint source-channel decoder, so as to obtain the recovered data and visual recognition information;
[0010] The first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0011] In a second aspect, an embodiment of the present application provides a data transmission method, which is applied to a second communication device. The method includes:
[0012] Receiving a first deep feature vector sent by a first communication device, where the first deep feature vector is a vector obtained by the first communication device encoding data to be transmitted through a first neural network model in a joint source-channel encoder;
[0013] Performing data recovery and visual recognition on the first deep feature vector based on a second neural network model in a joint source-channel decoder, so as to obtain the recovered data and visual recognition information, where the first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0014] In a third aspect, an embodiment of the present application provides a data transmission device, which is applied to a first communication device. The data transmission device includes:
[0015] An encoding module, configured to encode data to be transmitted through a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the data to be transmitted;
[0016] A transmission module, configured to transmit the first deep feature vector to a second communication device; where
[0017] The second communication device is configured to perform data recovery and visual recognition on the first deep feature vector through a second neural network model in a joint source-channel decoder, so as to obtain the recovered data and visual recognition information;
[0018] The first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0019] In a fourth aspect, an embodiment of the present application provides a data transmission device, which is applied to a second communication device. The data transmission device includes:
[0020] A receiving module, configured to receive a first deep feature vector sent by a first communication device, where the first deep feature vector is a vector obtained by the first communication device encoding data to be transmitted through a first neural network model in a joint source-channel encoder;
[0021] A processing module, configured to perform data recovery and visual recognition on the first depth feature vector based on a second neural network model in a joint source-channel decoder, to obtain the recovered data and visual recognition information, where the first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0022] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the data transmission method described in the first aspect and the second aspect above are implemented.
[0023] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data transmission method described in the first aspect and the second aspect above are implemented.
[0024] In a seventh aspect, an embodiment of the present application provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps of the data transmission method described in the first aspect and the second aspect above are implemented.
[0025] In the embodiments of the present application, since the first neural network model and the second neural network model are a joint model jointly trained based on an objective function, therefore, in the process of training the first neural network model and the second neural network model, the perception ability of the joint model for the channel used to transmit the depth feature vector can be improved, which is beneficial to improving the task performance and reducing the communication overhead. Description of the Drawings
[0026] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 is one of the flowcharts of a data transmission method provided by an embodiment of the present application;
[0028] Figure 2 is the second of the flowcharts of a data transmission method provided by an embodiment of the present application;
[0029] Figure 3 is a schematic diagram of a multi-task channel-aware communication system model in an embodiment of the present application;
[0030] Figure 4 is the schematic structural diagram of the joint source-channel encoder f in the embodiments of the present application e ;
[0031] Figure 5 is the schematic structural diagram of the joint source-channel decoder f in the embodiments of the present application d ;
[0032] Figure 6 is the third flowchart of a data transmission method provided by the embodiments of the present application
[0033] Figure 7 is one of the schematic structural diagrams of a data transmission device provided by the embodiments of the present application
[0034] Figure 8 is the second schematic structural diagram of a data transmission device provided by the embodiments of the present application
[0035] Figure 9 is the schematic structural diagram of an electronic device provided by the embodiments of the present application Specific Embodiments
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application
[0037] Please refer to Figure 1 , Figure 1 which is the schematic flowchart of a data transmission method provided by the embodiments of the present application, applied to a first communication device. The data transmission method includes the following steps
[0038] Step 101: Encode the data to be transmitted through the first neural network model in the joint source-channel encoder to obtain the first deep feature vector of the data to be transmitted
[0039] Step 102: Transmit the first deep feature vector to a second communication device
[0040] Among them, in the embodiments of the present application, the second communication device is used to perform data recovery and visual recognition on the first deep feature vector through the second neural network model in the joint source-channel decoder to obtain the recovered data and visual recognition information; the first neural network model and the second neural network model are joint models jointly trained based on the objective function
[0041] Among them, the data to be transmitted can be image data obtained through surveys in various scenarios, and the visual recognition can be various visual tasks in related technologies, such as visual tasks like image classification and image recognition. For the sake of easy understanding, in the following, taking the visual recognition as image classification as an example, the method provided by the embodiments of the present application will be further explained. Among them, the visual recognition information is the image category of the data to be transmitted. In some embodiments of the present application, the image category may include the following seven categories: environment, computer room, switching power supply, power distribution equipment, rooftop, battery, and others.
[0042] The first communication device in the embodiments of the present application may be a mobile device. The joint source-channel encoder may be deployed in the first communication device, or the joint source-channel encoder may also be an independent device jointly deployed with the first communication device. At this time, the joint source-channel encoder is communicatively connected to the first communication device. The second communication device may be a server. The joint source-channel decoder may be deployed in the second communication device, or the joint source-channel decoder may also be an independent device jointly deployed with the second communication device. At this time, the joint source-channel decoder is communicatively connected to the second communication device.
[0043] The data recovery in the embodiments of the present application specifically includes: the second neural network model restores the data to be transmitted according to the received deep feature vector. Correspondingly, the restored data is the image generated by performing image restoration on the data to be transmitted.
[0044] In this embodiment, since the first neural network model and the second neural network model are a joint model jointly trained based on an objective function, during the training process of the first neural network model and the second neural network model, the perception ability of the joint model for the channel used to transmit the deep feature vector can be improved, which is conducive to improving the task performance and reducing the communication overhead.
[0045] Optionally, the objective function is a function constructed based on a first parameter, a second parameter, and a third parameter. The first parameter is used to represent the dimensionality of the deep feature vector received by the joint source-channel decoder. The second parameter is used to represent the mutual information between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to represent: the mutual information between the visual recognition information output by the joint source-channel decoder and the restored data output by the joint source-channel decoder.
[0046] In the related art, there is a defect in multi-task communication based on deep learning compared to traditional separate source-channel coding and decoding communication. Multi-task communication based on joint source-channel coding has no channel awareness ability. In the related art, only the transceiver is jointly trained under fixed channel conditions, and the above-mentioned first deep feature vector is set as a feature vector with a fixed dimension. Since the optimal feature vector dimension number, that is, the number of transmission symbols, cannot be obtained according to the channel conditions, in both cases, problems will occur in the extraction of this fixed-dimension feature vector. One is that the feature vector dimension number, which is a hyperparameter in the deep learning network during joint training, is lower than the optimal feature vector dimension number under this channel condition. This will lead to a deterioration in the ability of the transmitted deep features to resist channel influence, and the loss of semantic information of multiple tasks in the extracted deep features, resulting in a decline in multi-task performance. The other is that the feature vector dimension number, which is a hyperparameter in the deep learning network during joint training, is higher than the optimal feature vector dimension number under this channel condition. This will lead to the retention of task-irrelevant redundancy in the extracted feature vector, resulting in an increase in the number of transmitted symbols. At a constant symbol rate, this will lead to a waste of communication overhead and an increase in transmission delay.
[0047] It can be seen that multi-task communication based on joint source-channel coding in the related art has no channel awareness ability, and the fixed-dimension design of the transmitted deep features leads to a decline in task performance or a waste of communication overhead. In addition, the design without obtaining additional channel state information (Channel State Information, CSI) results in its being only applicable to simple communication environments and unable to adapt to more complex wireless channels.
[0048] It should be noted that the transmission of the first deep feature vector to the second communication device can refer to: after the joint source-channel encoder encodes to obtain the first deep feature vector, the joint source-channel encoder directly uses the wireless channel to transmit the first deep feature vector to the second communication device.
[0049] It can be understood that during the training process of the model, the value of the first parameter can change during different iterations. Since the value of the first parameter can change, that is, during different iterations, the dimension number of the deep feature vector output by the joint source-channel decoder can change. Compared with setting the deep feature vector output by the joint source-channel decoder as a feature vector with a fixed dimension in the related art, in the embodiment of the present application, since the dimension number of the deep feature vector output by the joint source-channel decoder can change, dynamic deep features can be obtained according to the channel environment information conditions, which is beneficial to improving the model's channel awareness ability.
[0050] In some embodiments of the present application, since the objective function used in the process of jointly training the first neural network model and the second neural network model includes a first parameter, a second parameter, and a third parameter, the first parameter is used to characterize the number of dimensions of the deep feature vector received by the joint source-channel decoder, the second parameter is used to characterize the mutual information between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted, and the third parameter is used to characterize: the mutual information between the visual recognition information output by the joint source-channel decoder and the restored data output by the joint source-channel decoder. Thus, in the process of joint training, the dimension of the transmitted feature vector can be reduced, and at the same time, the mutual information between the received feature vector and the multi-task ground truth can be maximized, so as to obtain the feature vector dimension with the optimal dimension and balance the communication-sensing trade-off, that is, the trade-off between transmission delay and multi-task performance, which is beneficial to reducing the transmission delay while improving the task performance.
[0051] In some embodiments of the present application, the objective function is used to optimize the first parameter, the second parameter, and the third parameter. Among them, the smaller the value of the first parameter, the better the first parameter; the larger the value of the second parameter, the better the second parameter; the larger the value of the third parameter, the better the third parameter.
[0052] In some embodiments of the present application, the objective function is:
[0053]
[0054] Among them, the meaning of the above objective function is: by training θ e and θ d , to maximize the following function θ e is the parameter set of the joint source-channel encoder, the θ d is the parameter set of the joint source-channel decoder, the λ1 and the λ2 are respectively positive hyperparameters of the joint model, and the is the expectation of, is the first parameter, is the second parameter, is the third parameter, where the is the deep feature vector received by the joint source-channel decoder, x is the data to be transmitted, and y is the visual recognition information output by the joint source-channel decoder.
[0055] In this embodiment, since the first parameter is used to represent the number of dimensions of the depth feature vector received by the joint source-channel decoder, the smaller the first parameter, the less the amount of data in the depth feature vector. Correspondingly, the network overhead for transmitting this depth feature vector is also smaller. Therefore, by determining as small a first parameter as possible based on the objective function during the training of the model, it is beneficial to reduce the network overhead. Correspondingly, since the second parameter is used to represent the mutual information between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted, the larger the second parameter, the more image information is included in the depth information extracted by the second communication device. Therefore, by determining as large a second parameter as possible based on the objective function during the training of the model, it is beneficial to optimize the performance of image reconstruction. Since the third parameter is used to represent the mutual information between the visual recognition information output by the joint source-channel decoder and the restored data output by the joint source-channel decoder, the larger the third parameter, the more visual task-related information is included in the depth information extracted by the second communication device. Therefore, by determining as large a third parameter as possible based on the objective function during the training of the model, it is beneficial to optimize the performance of the image classification task.
[0056] In some embodiments of the present application, before encoding the data to be transmitted based on the first neural network model in the joint source-channel encoder to obtain the first depth feature vector, the method further includes:
[0057] Based on the objective function and the loss function, training the joint model to obtain the trained first neural network model and the second neural network model, where the loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function, and a fourth sub-loss function, and the first sub-loss function is used to minimize the value of the first parameter The second sub-loss function is used to represent the loss between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted, the third sub-loss function is used to represent the discrimination loss of the discriminator, and the fourth sub-loss function is used to represent the loss between the visual recognition information output by the joint source-channel decoder and the corresponding true visual result.
[0058] In some embodiments of the present application, the first sub-loss function is:
[0059]
[0060] where k1, k2, and k3 are constants, and Sig(·) represents the sigmoid function. σ 2is the additive noise variance, zi is the i-th element of z, where z is the deep feature vector obtained by encoding the data to be transmitted by the first neural network model, and n is the dimension number of z;
[0061] The second sub-loss function is:
[0062]
[0063] where E in the formula represents expectation.
[0064] The third sub-loss function is:
[0065]
[0066] where f dd is the discriminator, L D is the discriminator loss, and E in the formula represents expectation;
[0067] The fourth sub-loss function is:
[0068]
[0069] where E in the formula represents expectation.
[0070] The loss function is:
[0071]
[0072] where θ dd is the parameter set of the discriminator, and η1 and η2 are two positive hyperparameters of the joint model respectively.
[0073] It should be noted that the process of training the above joint model may include the following steps:
[0074] Obtain training data, where the training data may include: training images and the true visual results of the training images, where the true visual results may be the true categories of the training images.
[0075] When the first communication device obtains the training images, it can preprocess the training images. The preprocessing means include data enhancement and data augmentation and other means. The data increase can be operations such as scaling and rotating the training images, and the data augmentation can be generating images similar to the training images.
[0076] Then, input the preprocessed image data into the joint source-channel encoder, and encode the data to be transmitted based on the first neural network model in the joint source-channel encoder to obtain the deep feature vector z;
[0077] The joint source-channel encoder transmits the deep feature vector z to the joint source-channel decoder of the second communication device through a wireless channel. The joint source-channel decoder decodes the deep feature vector z to obtain the deep feature vector wherein is z affected by the channel, and has the same size as z.
[0078] The second neural network model performs data recovery and visual recognition based on the deep feature vector to obtain the recovered data and predicted visual recognition information.
[0079] Then, based on the recovered data, predicted visual recognition information, training images, and the true visual results of the training images and other related data, the above loss function is constructed, and the parameters of the joint model are optimized based on the constructed loss function to obtain an optimized joint model.
[0080] Meanwhile, based on other training data, the optimized joint model can be iteratively trained multiple times according to the above method until the convergence condition is met. The convergence condition may be that the loss value of the loss function is less than a preset value, or the number of iterations is greater than a preset number.
[0081] For ease of understanding, the following further explains the training process of the above joint model in combination with specific embodiments:
[0082] The embodiment of this application is directed to the transmission and classification recognition tasks of raw data of multiple modalities such as images, simulates the general neural network design and transmission scheme under different channel conditions, and while reducing the transmission delay caused by the reduction of communication overhead, ensures the stable performance of the parallel AI tasks of image reconstruction and classification. Specifically, a channel-aware multi-task communication system based on joint source-channel coding is proposed, which can adaptively adjust the dimension of deep features based on channel conditions to select the optimal number of transmission symbols, neither losing task performance nor causing waste of communication resources. In addition, reference signals are used to obtain implicit CSI information to guide the multi-task performance of the system under complex channels.
[0083] The embodiment of this application proposes a channel-aware multi-task communication system based on joint source-channel coding (Channel Aware Deep Joint Source-Channel Coding, CA-DJSCC), which can adaptively adjust the dimension of deep features according to channel conditions to select the optimal number of transmission symbols, without sacrificing task performance or wasting communication resources. In addition, a reference signal is introduced to obtain implicit CSI information to guide the multi-task performance of the system under complex channels. The complete transmission scheme is as follows: The original data of the first communication device of the mobile device is preprocessed (data enhancement, data augmentation, etc.) and then input into the joint source-channel encoder to obtain a deep feature vector. Then, the deep feature vector is input into the joint source-channel decoder deployed on the second communication device of the server to obtain the restored transmitted data and the results after performing visual tasks. The system flowchart is as Figure 2 shown. Among them, in some embodiments of this application, the original data may be the above-mentioned training images.
[0084] In Figure 2 the flowchart shown, the CA-DJSCC system has three key design steps in detail. One is the design of the system model and optimization objective for balancing task performance and communication delay, achieving a trade-off between multi-task performance and communication delay. The second is the design of the semantic encoder based on joint source-channel, and the third is the design of the semantic decoder based on joint source-channel. The joint source-channel encoder adaptively adjusts the dimension of the transmitted semantic information according to channel conditions, while the joint source-channel decoder facilitates data recovery and visual reasoning with the help of an implicit version of the additional reference signal known to the transceiver.
[0085] (1) Design of the system model and optimization objective for balancing task performance and communication delay
[0086] The embodiment of this application proposes a multi-task-oriented channel-aware communication scheme based on joint source-channel codec, where channel awareness means achieving a trade-off between multi-task performance and communication delay according to specific channel conditions. The system model is as Figure 3 shown.
[0087] Among them, \(x\in R\) H×W×C is the original data, and \(H\), \(W\), and \(C\) are the length, width, and number of channels of the image data respectively. Taking the embodiment of this application as an example, \(x\) is a survey and design image, including seven categories: environment, computer room, switch power supply, power distribution equipment, roof, battery, and others. The signal-to-noise ratio (SNR) is the signal-to-noise ratio of the additive noise of the channel, and its definition is as follows:
[0088]
[0089] Among them, \(P\)z is the power of the deep feature z transmitted into the wireless channel, i.e., the signal power; P n is the noise power. f e is the joint source-channel encoder. r ∈ R H×W×1 is a reference signal additionally added to the joint source-channel encoder, which is known to both the first communication device and the second communication device. z ∈ C n is a deep feature vector of an adaptive dimension extracted by the joint source-channel encoder after fusing the original data x, the reference signal r, and the channel condition guiding factor SNR in the deep feature space, where n is the number of dimensions of z, and n is different for different values of SNR and x, and is represented by the following formula:
[0090] z = f e (x, r, SNR; θ e ) (2)
[0091] where, θ e is the parameter set of f e . is z affected by the channel, with the same size as z, and can be represented by the following formula:
[0092] In some embodiments of the present application, the second communication device is configured to perform data recovery and visual recognition on the first deep feature vector through a second neural network model in the joint source-channel decoder, and obtain the recovered data and visual recognition information, including:
[0093] The second communication device is configured to decrypt the first deep feature vector through the joint source-channel decoder to obtain a second deep feature vector, where the second deep feature vector is the first deep feature vector affected by the channel, and the number of dimensions of the first deep feature vector is the same as that of the second deep feature vector;
[0094] The second communication device is configured to perform data recovery and visual recognition on the second deep feature vector through the second neural network model, and obtain the recovered data and visual recognition information.
[0095] Among them, the second communication device is configured to decrypt the first deep feature vector through the joint source-channel decoder to obtain the second deep feature vector, which can be implemented based on the following formula:
[0096]
[0097] where, is the second deep feature vector, z is the first deep feature vector, h ∈ C n is the channel response, n ∈ C n is the additive noise. f dis a joint source-channel decoder. is the data recovered by the second communication device, with the size being consistent with x. is the result of the visual task performed by the second communication device. The size of the reference signal is related to the task performance. In this case, Since there are a total of seven types of image tags, namely the above-mentioned seven types of tags: "environment, computer room, switch power supply, power distribution equipment, rooftop, battery, and others". It can be expressed by the following formula:
[0098]
[0099] The embodiments of this application can be applied to other multi-tasks, including image classification, image retrieval, image restoration, semantic segmentation, and object detection. By appropriately updating the network structure, corresponding tasks can also be performed on raw data such as text, audio, and video.
[0100] The goal of the CA-DJSCC proposed in the embodiments of this application is to optimize the number of dimensions of the transmitted deep features under different dynamically changing channel conditions, improve the performance of data recovery and visual reasoning, and minimize the transmission delay. To achieve this channel-aware trade-off, the transmission delay and task performance are characterized by the number of dimensions of z and the mutual information between the actual ground truth of the multi-task. The optimization objective function is expressed as follows:
[0101]
[0102] where, represents the expectation of, D(·) is the operation to obtain the number of dimensions, and λ1 and λ2 are two positive-valued hyperparameters. The first term is the average value of the number of dimensions n of the deep features obtained by the second communication device. The smaller this value is, that is, the smaller the number of transmitted symbols, the lower the communication delay; the second term is the mutual information between the actual ground truth data and the deep features of the image reconstruction by the second communication device. The larger this mutual information is, it means that the deep information extracted by the second communication device contains more image information, which is used to optimize the performance of image reconstruction; the third term is the mutual information between the actual ground truth label and the deep features recovered by the second communication device. The larger the mutual information is, it means that the deep information extracted by the second communication device contains more information related to the visual task, which is used to optimize the performance of the image classification task. Through end-to-end joint training and tuning of the hyperparameters λ1 and λ2, the communication delay and multi-task performance can be balanced to obtain the optimal number of dimensions of the transmitted deep features and achieve the lowest transmission delay under the multi-task performance requirements.
[0103] As shown in Equation (5), the goal of the proposed CA-DJSCC is to reduce the dimension of the transmitted feature vectors while maximizing the mutual information between the received feature vectors and the multi-task ground truth, so as to obtain the feature vector dimension with the optimal dimension and balance the communication-sensing trade-off, that is, the trade-off between transmission delay and multi-task performance. To achieve the optimization goal of (5), a loss function is designed, including three parts: dimension adjustment, image reconstruction, and classification inference:
[0104] (1) Dimension adjustment: The goal is to minimize the dimension of the transmitted deep feature vectors . The corresponding loss function is deduced in Equation (5) and is expressed as:
[0105]
[0106] where k1, k2, and k3 are constants, Sig(·) represents the sigmoid function, σ 2 is the additive noise variance, and zi i is the i-th element of z.
[0107] (2) Image reconstruction: The deep feature vectors are mapped into the reconstructed images during the maximization process of . Generally, the maximization of corresponds to the minimization of the MSE loss function, which can measure the similarity between the ground truth image and the restored image and is expressed by the following formula:
[0108]
[0109] Since the embodiments of this application use a generative adversarial network to obtain and follow the training strategy of the generative adversarial network, additional discriminative losses are required for the adversarial training of the generator and the discriminator, which are expressed by the following formula:
[0110]
[0111] where the generator fdg is trained to minimize the restoration loss and the discriminative loss L D to confuse the recognition results as much as possible, and the discriminator fdd is trained to maximize the discriminative loss L D to distinguish the restored image and the original image as much as possible. In addition, it is also necessary to minimize to restore the reference signal.
[0112] (3) Classification inference: The deep feature vectors are mapped into the visual inference labels during the maximization process of . Generally, Maximization corresponds to minimization of the cross - entropy loss function, which measures the similarity between the actual true - value labels and the predicted labels in a visual classification inference task, and is expressed as follows:
[0113]
[0114] The loss function obtained through the weighted sum of equations (7), (8), (9) and (10) is expressed as:
[0115]
[0116] where η1 and η2 are two positive hyperparameters.
[0117] (II) Design of the semantic encoder based on joint source - channel
[0118] Optionally, encoding the data to be transmitted by the first neural network model in the joint source - channel encoder to obtain the first deep feature vector of the data to be transmitted includes:
[0119] Performing feature fusion on the data to be transmitted, a preset reference signal, and the signal - to - noise ratio, to obtain first fusion characteristic information, where the first communication device and the second communication device both store the preset reference signal, and the signal - to - noise ratio is the signal - to - noise ratio of the additive noise of the channel for transmitting the first deep feature vector;
[0120] Generating an initial semantic information vector corresponding to the first fusion characteristic information, and generating a semantic importance vector corresponding to the first fusion characteristic information, where the initial semantic information vector is used to characterize the semantics of the first fusion characteristic information, and the semantic importance vector is used to characterize the semantic importance of each item of information in the first fusion characteristic information;
[0121] Calculating the product of the semantic importance vector and an upper - triangular matrix composed of 0s and 1s to obtain a target semantic importance vector, where the dimension number of the upper - triangular matrix is n0;
[0122] Generating a mask vector corresponding to the target semantic importance vector according to a threshold, where when the i - th element in the target semantic importance vector is greater than the threshold, the i - th element in the mask vector is 1; when the i - th element in the target semantic importance vector is less than or equal to the threshold, the i - th element in the mask vector is 0, and i = 1, …, n0;
[0123] Calculating the product of the initial semantic information vector and the mask vector to obtain the first deep feature vector.
[0124] Among them, the joint source-channel encoder is one of the two core modules in the CA-DJSCC designed in the embodiments of this application, and its functions are as follows: (1) Introduce a reference signal to obtain implicit CSI for the joint source-channel decoder of the second communication device; (2) Design an adaptive dimension pruning method, which can adaptively change the dimension of the transmitted deep features according to different original data and channel conditions, so as to adaptively adjust the number of transmitted symbols, allocate the best communication overhead under certain channel conditions for multi-task communication, achieve the trade-off between communication delay and multi-task performance, and maintain good AI multi-task performance while reducing communication overhead. (3) Under the condition of limited transmission resources, it can adjust the transmission overhead through the system to provide communication services at the cost of sacrificing part of the multi-task performance.
[0125] The joint source-channel encoder f e has a structure as Figure 4 shown.
[0126] First, extract the initial semantic information vector, and then selectively prune its dimension according to the inherent semantic importance. Specifically, fuse the original data x, the reference signal r, and the signal-to-noise ratio SNR in the latent space to generate the initial semantic information vector z 0 and the semantic importance vector ξ. Both z 0 and ξ have a predetermined dimension n0, which is greater than the dimension n of the finally extracted semantic information. Physically, ξ measures the influence of each element in z 0 on the multi-task performance. Considering ξ and the semantic importance threshold ξ thr , the embodiments of this application can prune the elements with lower semantic importance in z 0 and finally obtain z with different dimensions according to the channel conditions.
[0127] The network structure and working process of the joint source-channel encoder f e are as Figure 3As shown in the figure. First, the original data \(x\) and the reference signal \(r\) are merged on the third axis as the transmission information and input into the deep network layer(1). At the same time, the signal-to-noise ratio SNR, as the channel condition, is input into the deep network layer(2). This is to integrate the semantic information from the original image and the reference signal and incorporate the channel condition into the dimension pruning of the semantic information, that is, the transmitted deep features. In addition, the semantic importance vector \(\xi\) measures the degree of multi-task correlation in the semantic information. It is only related to the transmission context and the channel condition, and the latter is used to guide the dimension adjustment and multi-task adjustment. Therefore, through the end-to-end joint training of the multi-task execution part in the decoder, it can be learned through the fused features of the original data \(x\), the reference signal \(r\), and the signal-to-noise ratio SNR. Therefore, the latent feature vectors obtained from layer(1) and layer(2) are merged on the first axis to obtain a new feature map, which is respectively input into the deep network layer(3) and the deep network layer(4) to obtain the initial semantic information vector \(z_0\) and the semantic importance vector \(\xi\). By multiplying the output of layer(3) by an \(n_0\)-dimensional upper triangular matrix composed of 1 and 0, the elements in \(\xi\) are in reverse order. Using the obtained \(\xi\), layer(4) is used to generate \(z_0\), which is expressed by the following formula:
[0128]
[0129] where \(\xi\) i is the \(i\)-th element of \(\xi\), \(i = 1,\ldots,n_0\), and \(v\) are respectively the stacked parameter matrix and the stacked input feature vector of the linear layer with the last input size of \(n\) in and the output size of \(n_0\). For the convenience of calculation, the weight matrix and the bias vector are stacked along the second axis, and the input feature vector is stacked 1 to match the augmented form of the weight matrix and the bias. is the vector of the \(i\)-th row of
[0130] In terms of the network structure, fully connected layers and residual blocks are mainly used. Among them, the convolutional kernels generally adopt a 3×3 kernel, a stride of 1, and a padding size of 1 to traverse the graph features more comprehensively and extract more precise feature maps. If the computing power is limited or the parameter deployment pressure on mobile devices is high, larger convolutional kernels can be replaced to reduce the number of calculations and save overhead. Except for layer(4), simple ReLu is used in the activation layer to extract non-linear features, and tanh is used in layer(4) to limit the element size of the deep feature vector. In addition, the normalization method of BatchNorm is adopted to accelerate the training speed, and at the same time, the size of the transmitted feature vector can be normalized, that is, the signal power is obtained as 1, so as to more conveniently calculate the relationship between the signal-to-noise ratio and the channel additive noise variance. The specific structure of the network is shown in the following table:
[0131] Table 1 Network Structure Parameters
[0132]
[0133]
[0134] This part combines the principles of signal-to-noise ratio and semantic importance, and can extract deep features of different dimensions for transmission under different channel conditions. This joint source-channel encoder can start from the original image, be guided by the channel conditions, and dynamically prune the deep features with low correlation to multi-tasks through semantic importance, so as to achieve the selection of the optimal dimension and the trade-off between performance and communication delay. And they are all composed of deep network structures that can be jointly trained and backpropagated.
[0135] (3) Design of the semantic decoder based on joint source-channel
[0136] Joint source-channel decoding is another core module in the CA-DJSCC designed in the embodiments of this application, and its functions are as follows: (1) By introducing the reference signal, the reference signal known to both ends is restored at the recovery end. Since the reference signal and the original data are affected by the same channel, the implicit channel state information can be learned through the deep neural network to guide the execution of multi-tasks. This approach is similar to using pilot signals for channel estimation. The advantage is that it is incorporated into the neural network that can be jointly trained and backpropagated, and no additional channel estimation is required. And it has an irreplaceable effect in some scenarios where perfect CSI cannot be obtained through channel estimation; (2) The generative adversarial network is introduced, which can improve the performance of data recovery in multi-tasks. The discriminator in the generative adversarial network can be discarded after training, so it does not need to be deployed in the second communication device, which can reduce the deployment overhead and resources.
[0137] Joint source-channel decoder f d has the structure as Figure 5 shown.
[0138] The network structure and working process of the joint source-channel decoder are as follows Figure 5 As shown. Recover the reference signal from the received semantic information affected by the channel and obtain the implicit CSI from the residual between and the original reference signal r, denoted as Δr. Subsequently, fuse Δr with the deep features extracted from Under the guidance of introducing the reference signal, obtain the enhanced version zf of the received semantic information. Finally, an adversarial generation network is introduced to obtain high-quality image reconstruction from zf and a classification inference to perform visual tasks. Specifically, first use f
[0139] to recover from dr from and then calculate Δr from the residual between and r. Since the reference signal r and the original data experience the same channel conditions, Δr contains an implicit version of the CSI. Then, merge Δr and the deep features extracted from through layer(5) on the third axis to obtain the enhanced version z of the received semantic information . Under the guidance of the implicit version of the CSI obtained from Δr, z f can perform better performance multi-tasks. Finally, input z f into the GAN and inference module f f composed of the generator f dg and the discriminator f dd to obtain the restored data di and the visual inference result. f is an auxiliary discriminator independent of the decoder, so the parameter set of f dd is not included in θ dd , denoted as θ d . dd
[0140] In terms of network structure, fully connected layers and residual blocks ResBlock are mainly used, where the convolutional kernels generally use 3×3 kernels, a stride of 1 and a padding size of 1. The activation layers all use simple ReLu, and in addition, the normalization method of BatchNorm is adopted to speed up the training speed. The specific structure of the network is shown in the following table:
[0141] Table 2 Network structure parameters
[0142] Module Input Dimension Output Dimension layer(5) 128 512×512×1 <![CDATA[f dr > 128 512×512×1 <![CDATA[f dg > 512×512×2 512×512×3 <![CDATA[f dd > 512×512×2 2 <![CDATA[f di > 512×512×2 7
[0143] Among them, 512×512×3 is the size of the exploration training image in the embodiment of the present application, and the size can be changed according to the size of the data set and the modality, f di The output size 7 of is the number of classification categories, which can be adjusted according to the task.
[0144] Through the above work process, the implicit version of the CSI obtained by the recovery of the reference signal and the calculation of the residual enhances the reception ability of the depth features, providing effective guidance for data recovery and visual reasoning. And all are composed of deep network structures that can be jointly trained and backpropagated.
[0145] The method provided by the embodiment of the present application at least further has the following beneficial effects:
[0146] The existing multi-task communication based on joint source-channel coding has no channel perception ability, and it is easy to cause a decline in task performance or waste of communication overhead due to its fixed-dimension design for transmitting depth features. In addition, the design without additionally obtaining CSI results in its being only applicable to simple communication environments and unable to adapt to more complex communication channels. The embodiment of the present application can adaptively adjust the dimension of the depth features based on the channel conditions to select the optimal number of transmission symbols, neither losing task performance nor causing waste of communication resources. In addition, implicit CSI information is obtained using the reference signal to guide the multi-task performance of the system under complex channels.
[0147] Aiming at the limited transmission conditions such as high dynamics, long distance, time-varying, and resource limitation in the 6G satellite-ground coverage scenario, realizing efficient information extraction, reliable information recovery, and seamless network coverage expansion based on semantic communication is an effective path. The embodiment of the present application faces the transmission and classification recognition tasks of raw data of multiple modalities such as images, simulates the general neural network design and transmission scheme under non-ideal channel conditions, improves the performance of image reconstruction and feature fusion, maintains good data reception recovery and AI task performance while reducing communication overhead, expands the efficient information extraction and reliable information recovery scheme for satellite-ground transmission semantic communication to enhance satellite-ground transmission capabilities, and has broad application prospects.
[0148] Please refer to Figure 6 , the embodiment of the present application also provides a data transmission method, which is applied to a second communication device, and the method includes:
[0149] Step 601, receiving a first depth feature vector sent by a first communication device, where the first depth feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on a first neural network model in a joint source-channel encoder;
[0150] Step 602: Based on the second neural network model in the joint source-channel decoder, perform data recovery and visual recognition on the first deep feature vector to obtain the recovered data and visual recognition information, where the first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0151] This implementation is the method on the side of the second communication device corresponding to the above embodiment. Its specific implementation process corresponds to the above embodiment and has corresponding beneficial effects. To avoid repetition, it will not be elaborated here.
[0152] Please refer to Figure 7 , Figure 7 A data transmission device 700 provided by an embodiment of the present application is applied to a first communication device. The data transmission device 700 includes:
[0153] An encoding module 701, configured to encode data to be transmitted through a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the data to be transmitted;
[0154] A transmission module 702, configured to transmit the first deep feature vector to a second communication device; where
[0155] The second communication device is configured to perform data recovery and visual recognition on the first deep feature vector through a second neural network model in a joint source-channel decoder to obtain the recovered data and visual recognition information;
[0156] The first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0157] Optionally, the objective function is a function constructed based on a first parameter, a second parameter, and a third parameter. The first parameter is used to characterize the number of dimensions of the deep feature vector received by the joint source-channel decoder. The second parameter is used to characterize the mutual information between the recovered data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to characterize: the mutual information between the visual recognition information output by the joint source-channel decoder and the recovered data output by the joint source-channel decoder.
[0158] Optionally, the objective function is used to optimize the first parameter, the second parameter, and the third parameter. Among them, the smaller the value of the first parameter, the better the first parameter; the larger the value of the second parameter, the better the second parameter; the larger the value of the third parameter, the better the third parameter.
[0159] Optionally, the objective function is:
[0160]
[0161] Among them, θ e is the parameter set of the joint source-channel encoder, and the θ d is the parameter set of the joint source-channel decoder. The λ1 and the λ2 are respectively positive hyperparameters of the joint model. The is the expectation of, is the first parameter, is the second parameter, is the third parameter. Among them, the is the deep feature vector received by the joint source-channel decoder, x is the data to be transmitted, y is the visual recognition information output by the joint source-channel decoder, and the joint model is a joint model formed by combining the first neural network model and the second neural network model.
[0162] Optionally, the data transmission device 700 further includes:
[0163] A training module, configured to train the joint model based on the objective function and the loss function, to obtain the trained first neural network model and the second neural network model. Among them, the loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function, and a fourth sub-loss function. The first sub-loss function is used to minimize the value of the first parameter. The second sub-loss function is used to characterize the loss between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted. The third sub-loss function is used to characterize the discrimination loss of the discriminator. The fourth sub-loss function is used to characterize the loss between the visual recognition information output by the joint source-channel decoder and the corresponding true visual result.
[0164] Optionally, the first sub-loss function is:
[0165]
[0166] Among them, k1, k2, and k3 are constants, and Sig(·) represents the sigmoid function. σ 2 is the additive noise variance, zi is the i-th element of z, the z is the deep feature vector obtained by encoding the data to be transmitted by the first neural network model, and n is the dimension number of z;
[0167] The second sub-loss function is:
[0168]
[0169] The third sub-loss function is as follows:
[0170]
[0171] where f dd is the discriminator, and L D is the discriminator loss;
[0172] The fourth sub-loss function is as follows:
[0173]
[0174] The loss function is as follows:
[0175]
[0176] where θ dd is the parameter set of the discriminator, and the η1 and the η2 are respectively two positive hyperparameters of the joint model.
[0177] Optionally, the second communication device is configured to perform data recovery and visual recognition on the first depth feature vector through a second neural network model in the joint source-channel decoder, and obtain the recovered data and visual recognition information, including:
[0178] The second communication device is configured to decrypt the first depth feature vector through the joint source-channel decoder to obtain a second depth feature vector, where the second depth feature vector is the first depth feature vector affected by the channel, and the dimensionality of the first depth feature vector is the same as that of the second depth feature vector;
[0179] The second communication device is configured to perform data recovery and visual recognition on the second depth feature vector through the second neural network model, and obtain the recovered data and visual recognition information.
[0180] Optionally, the encoding module 701 is specifically configured to: perform feature fusion on the data to be transmitted, a preset reference signal, and the signal-to-noise ratio to obtain first fusion characteristic information, where the first communication device and the second communication device both store the preset reference signal, and the signal-to-noise ratio is the signal-to-noise ratio of the additive noise of the channel for transmitting the first depth feature vector;
[0181] generate an initial semantic information vector corresponding to the first fusion characteristic information, and generate a semantic importance vector corresponding to the first fusion characteristic information, where the initial semantic information vector is used to represent the semantics of the first fusion characteristic information, and the semantic importance vector is used to represent the semantic importance of each item of information in the first fusion characteristic information;
[0182] Calculate the product of the semantic importance vector and an upper triangular matrix composed of 0s and 1s to obtain a target semantic importance vector, where the number of dimensions of the upper triangular matrix is n0;
[0183] Generate a mask vector corresponding to the target semantic importance vector according to a threshold. Specifically, when the i-th element in the target semantic importance vector is greater than the threshold, the i-th element in the mask vector is 1; when the i-th element in the target semantic importance vector is less than or equal to the threshold, the i-th element in the mask vector is 0, where i = 1,..., n0;
[0184] Calculate the product of the initial semantic information vector and the mask vector to obtain the first depth feature vector.
[0185] It should be noted that the data transmission device 700 provided in the embodiments of the present application is a device capable of executing Figure 1 the data transmission method. All implementation manners in the above data transmission method embodiments are applicable to this device, and all can achieve the same or similar beneficial effects. To avoid repeated description, this embodiment will not be elaborated.
[0186] Please refer to Figure 8 , Figure 8 A data transmission device 800 provided in the embodiments of the present application, which is applied to a second communication device. The data transmission device 800 includes:
[0187] A receiving module 801, configured to receive a first depth feature vector sent by a first communication device, where the first depth feature vector is a vector obtained by the first communication device encoding data to be transmitted based on a first neural network model in a joint source-channel encoder;
[0188] A processing module 802, configured to perform data recovery and visual recognition on the first depth feature vector based on a second neural network model in a joint source-channel decoder to obtain the recovered data and visual recognition information, where the first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0189] It should be noted that the data transmission device 800 provided in the embodiments of the present application is a device capable of executing Figure 6 the data transmission method. All implementation manners in the above data transmission method embodiments are applicable to this device, and all can achieve the same or similar beneficial effects. To avoid repeated description, this embodiment will not be elaborated.
[0190] Specifically, refer to Figure 9As shown in the figure, an embodiment of the present application further provides an electronic device, including a bus 901, a transceiver 902, an antenna 903, a bus interface 904, a processor 905, and a memory 906.
[0191] The processor 905 is configured to:
[0192] Encode the data to be transmitted through a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the data to be transmitted;
[0193] Transmit the first deep feature vector to a second communication device; wherein,
[0194] The second communication device is configured to perform data recovery and visual recognition on the first deep feature vector through a second neural network model in a joint source-channel decoder to obtain the recovered data and visual recognition information;
[0195] The first neural network model and the second neural network model are a joint model jointly trained based on an objective function.
[0196] In Figure 9 , the bus architecture (represented by the bus 901), the bus 901 may include any number of interconnected buses and bridges, and the bus 901 links various circuits including one or more processors represented by the processor 905 and a memory represented by the memory 906 together. The bus 901 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface 904 provides an interface between the bus 901 and the transceiver 902. The transceiver 902 may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on a transmission medium. The data processed by the processor 905 is transmitted on a wireless medium through the antenna 903. Further, the antenna 903 also receives data and transmits the data to the processor 905.
[0197] The processor 905 is responsible for managing the bus 901 and general processing, and may also provide various functions including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 906 may be used to store data used by the processor 905 when performing operations.
[0198] Optionally, the processor 905 may be a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or a Complex Programmable Logic Device (CPLD).
[0199] Optionally, the objective function is a function constructed based on a first parameter, a second parameter, and a third parameter. The first parameter is used to characterize the number of dimensions of the depth feature vector received by the joint source-channel decoder. The second parameter is used to characterize the mutual information between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted. The third parameter is used to characterize the mutual information between the visual recognition information output by the joint source-channel decoder and the restored data output by the joint source-channel decoder.
[0200] Optionally, the objective function is used to optimize the first parameter, the second parameter, and the third parameter. Among them, the smaller the value of the first parameter, the better the first parameter; the larger the value of the second parameter, the better the second parameter; the larger the value of the third parameter, the better the third parameter.
[0201] Optionally, the objective function is:
[0202]
[0203] where θ e is the parameter set of the joint source-channel encoder, the θ d is the parameter set of the joint source-channel decoder, λ1 and λ2 are respectively positive hyperparameters of the joint model, the is the expectation of, is the first parameter, is the second parameter, is the third parameter, where the is the depth feature vector received by the joint source-channel decoder, x is the data to be transmitted, y is the visual recognition information output by the joint source-channel decoder, and the joint model is a joint model formed by combining the first neural network model and the second neural network model.
[0204] Optionally, the processor 905 is further configured to:
[0205] Based on the objective function and the loss function, the joint model is trained to obtain the trained first neural network model and the second neural network model, where the loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function, and a fourth sub-loss function. The first sub-loss function is used to minimize the value of the first parameter. The second sub-loss function is used to characterize the loss between the restored data output by the joint source-channel decoder and the corresponding data to be transmitted. The third sub-loss function is used to characterize the discrimination loss of the discriminator. The fourth sub-loss function is used to characterize the loss between the visual recognition information output by the joint source-channel decoder and the corresponding true visual result.
[0206] Optionally, the first sub-loss function is:
[0207]
[0208] where k1, k2, and k3 are constants, and Sig(·) represents the sigmoid function. σ 2 is the additive noise variance, zi is the i-th element of z, where z is the deep feature vector obtained by encoding the data to be transmitted by the first neural network model, and n is the dimensionality of z.
[0209] The second sub-loss function is:
[0210]
[0211] The third sub-loss function is:
[0212]
[0213] where f dd is the discriminator, and L D is the discrimination loss.
[0214] The fourth sub-loss function is:
[0215]
[0216] The loss function is:
[0217]
[0218] where θ dd is the parameter set of the discriminator, and η1 and η2 are two positive hyperparameters of the joint model respectively.
[0219] Optionally, the second communication device is configured to perform data recovery and visual recognition on the first deep feature vector through a second neural network model in a joint source-channel decoder, and obtain the recovered data and visual recognition information, including:
[0220] The second communication device is configured to decrypt the first deep feature vector through the joint source-channel decoder to obtain a second deep feature vector, where the second deep feature vector is the first deep feature vector affected by the channel, and the number of dimensions of the first deep feature vector is the same as that of the second deep feature vector;
[0221] The second communication device is configured to perform data recovery and visual recognition on the second deep feature vector through the second neural network model to obtain the recovered data and visual recognition information.
[0222] Optionally, the processor 905 is specifically configured to:
[0223] Perform feature fusion on the data to be transmitted, a preset reference signal, and a signal-to-noise ratio to obtain first fusion characteristic information, where both the first communication device and the second communication device store the preset reference signal, and the signal-to-noise ratio is the signal-to-noise ratio of the additive noise of the channel for transmitting the first deep feature vector;
[0224] Generate an initial semantic information vector corresponding to the first fusion characteristic information, and generate a semantic importance vector corresponding to the first fusion characteristic information, where the initial semantic information vector is used to represent the semantics of the first fusion characteristic information, and the semantic importance vector is used to represent the semantic importance of each item of information in the first fusion characteristic information;
[0225] Calculate the product of the semantic importance vector and an upper triangular matrix composed of 0s and 1s to obtain a target semantic importance vector, where the number of dimensions of the upper triangular matrix is n0;
[0226] Generate a mask vector corresponding to the target semantic importance vector according to a threshold, where when the i-th element in the target semantic importance vector is greater than the threshold, the i-th element in the mask vector is 1; when the i-th element in the target semantic importance vector is less than or equal to the threshold, the i-th element in the mask vector is 0, and i = 1,..., n0;
[0227] Calculate the product of the initial semantic information vector and the mask vector to obtain the first deep feature vector.
[0228] It should be noted that the electronic device provided in the embodiments of the present application is a device capable of executing the above data transmission method. Therefore, all implementation manners in the embodiments of the above data transmission method are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not be described in detail.
[0229] An embodiment of the present invention further provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements each process of the above data transmission method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0230] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above data transmission method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0231] An embodiment of the present application further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above data transmission method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0232] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element.
[0233] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0234] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A data transmission method, characterized in that: Applied to a first communication device, the method comprises: Encoding the data to be transmitted by a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the data to be transmitted; The first depth feature vector is transmitted to a second communication device; wherein, The second communication device is used to perform data recovery and visual recognition on the first depth feature vector through a second neural network model in a joint source channel decoder to obtain recovered data and visual recognition information; The first neural network model and the second neural network model are joint models obtained by joint training based on an objective function.
2. The method according to claim 1, characterized in that The objective function is a function constructed based on the first parameter, the second parameter and the third parameter. Characterizes the number of dimensions of the depth feature vector received by the joint source channel decoder, the second parameter is used to characterize the mutual information between the recovered data output by the joint source channel decoder and the corresponding data to be transmitted, and the third parameter is used to characterize: the mutual information between the visual recognition information output by the joint source channel decoder and the recovered data output by the joint source channel decoder.
3. The method according to claim 2, characterized in that The value of the first parameter of the objective function is negatively correlated with the communication delay, the value of the second parameter is positively correlated with the image reconstruction performance, and the value of the third parameter is positively correlated with the image classification performance.
4. The method according to any one of claims 2 to 3, characterized in that The objective function is: Among them, θ e is the parameter set of the joint source channel encoder, the θ d is the parameter set of the joint source channel decoder, the λ1 and the λ2 are positive hyperparameters of the joint model, respectively. for expectations, is the first parameter, is the second parameter, is the third parameter, wherein the is the depth feature vector received by the joint source-channel decoder, x is the data to be transmitted, and y is the visual recognition information output by the joint source-channel decoder.
5. The method according to claim 4, characterized in that Before encoding the data to be transmitted based on the first neural network model in the joint source channel encoder to obtain a deep feature vector, the method further includes: Based on the objective function and the loss function, the joint model is trained to obtain the trained first neural network model and the second neural network model, wherein the loss function includes a first sub-loss function, a second sub-loss function, a third sub-loss function and a fourth sub-loss function, the first sub-loss function is used to minimize the value of the first parameter, the second sub-loss function is used to characterize the loss between the recovered data output by the joint source channel decoder and the corresponding data to be transmitted, the third sub-loss function is used to characterize the identification loss of the discriminator, and the fourth sub-loss function is used to characterize the loss between the visual recognition information output by the joint source channel decoder and the corresponding real visual result.
6. The method according to claim 5, characterized in that The first sub-loss function is: Among them, k1, k2 and k3 are constants, Sig(·) represents the sigmoid function, σ 2 is the additive noise variance, zi is the i-th element of z, said z is the deep feature vector obtained after the first neural network model encodes the data to be transmitted, and n is the dimension number of z; The second sub-loss function is: The third sub-loss function is: Among them, f dd is the discriminator, L D To identify the loss; The fourth sub-loss function is: The loss function is: Among them, θ dd is the parameter set of the discriminator, and the η1 and the η2 are two positive hyperparameters of the joint model.
7. The method according to claim 1, characterized in that The second communication device is used to perform data recovery and visual recognition on the first depth feature vector through a second neural network model in a joint source channel decoder to obtain recovered data and visual recognition information, including: The second communication device is used to decrypt the first depth feature vector through a joint source-channel decoder to obtain a second depth feature vector, wherein the second depth feature vector is the first depth feature vector affected by the channel, and the number of dimensions of the first depth feature vector is the same as the number of dimensions of the second depth feature vector; The second communication device is used to perform data recovery and visual recognition on the second deep feature vector through the second neural network model to obtain restored data and visual recognition information.
8. The method according to claim 1, characterized in that The encoding of the data to be transmitted by the first neural network model in the joint source channel encoder to obtain a first deep feature vector of the data to be transmitted includes: Performing feature fusion on the data to be transmitted, a preset reference signal, and a signal-to-noise ratio to obtain first fusion characteristic information, wherein both the first communication device and the second communication device store the preset reference signal, and the signal-to-noise ratio is a signal-to-noise ratio of additive noise of a channel used to transmit the first depth feature vector; Generate an initial semantic information vector corresponding to the first fused characteristic information, and generate a semantic importance vector corresponding to the first fused characteristic information, wherein the initial semantic information vector is used to represent the semantics of the first fused characteristic information, and the semantic importance vector is used to represent the semantic importance of each information in the first fused characteristic information; Calculate the product of the semantically important vector and an upper triangular matrix consisting of 0s and 1s to obtain a target semantically important vector, where the dimension of the upper triangular matrix is n0; Generate a mask vector corresponding to the target semantically important vector according to a threshold, wherein, when the i-th element in the target semantically important vector is greater than the threshold, the i-th element in the mask vector is 1; when the i-th element in the target semantically important vector is less than or equal to the threshold, the i-th element in the mask vector is 0, where i=1, ..., n0; The product of the initial semantic information vector and the mask vector is calculated to obtain the first depth feature vector.
9. A data transmission method, characterized in that: Applied to a second communication device, the method comprises: Receiving a first deep feature vector sent by a first communication device, wherein the first deep feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on a first neural network model in a joint source channel encoder; Based on the second neural network model in the joint source channel decoder, data recovery and visual recognition are performed on the first depth feature vector to obtain restored data and visual recognition information, wherein the first neural network model and the second neural network model are a joint model obtained by joint training based on an objective function.
10. A data transmission device, characterized in that: Applied to a first communication device, the data transmission device comprises: An encoding module, used for encoding the data to be transmitted by using a first neural network model in a joint source-channel encoder to obtain a first deep feature vector of the data to be transmitted; A transmission module, used to transmit the first depth feature vector to a second communication device; wherein, The second communication device is used to perform data recovery and visual recognition on the first depth feature vector through a second neural network model in a joint source channel decoder to obtain recovered data and visual recognition information; The first neural network model and the second neural network model are joint models obtained by joint training based on an objective function.
11. A data transmission device, characterized in that: Applied to a second communication device, the device comprising: A receiving module, configured to receive a first deep feature vector sent by a first communication device, wherein the first deep feature vector is a vector obtained by the first communication device encoding the data to be transmitted based on a first neural network model in a joint source channel encoder; A processing module is used to perform data recovery and visual recognition on the first depth feature vector based on a second neural network model in a joint source channel decoder to obtain recovered data and visual recognition information, wherein the first neural network model and the second neural network model are a joint model obtained by joint training based on an objective function.
12. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the data transmission method according to any one of claims 1 to 9 are implemented.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data transmission method according to any one of claims 1 to 9 are implemented.
14. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the data transmission method according to any one of claims 1 to 9.