A Source-Channel Joint Encoding / Decoding Method for Cross-Mode Communication Systems
By employing a cross-modal source-channel joint encoding and decoding method, fusing image and tactile features, and utilizing Wasserstein generative adversarial networks and knowledge distillation techniques, the problem of variable transmission environments in multimodal communication is solved, achieving efficient and robust signal transmission.
Patent Information
- Application Number
- CN202411075763.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-08-07
AI Technical Summary
Existing cross-modal communication systems cannot effectively handle the transmission of multimodal data, especially in situations with variable transmission environments, and cannot guarantee signal transmission quality and robustness.
A source-channel joint encoding and decoding method based on a cross-modal communication system is designed. Image and tactile features are extracted and fused through a cross-modal source encoder, channel transmission is performed using a cross-modal channel encoder and decoder, and the image signal is reconstructed at the receiving end. Wasserstein generative adversarial network and knowledge distillation techniques are used to improve the image reconstruction quality.
It achieves efficient processing and transmission of multimodal data, overcomes the cliff effect in traditional encoding and decoding methods, and improves the robustness and quality of signal transmission.
Smart Images

Figure CN119011843B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cross-modal image signal reconstruction technology, specifically to a source-channel joint encoding and decoding method based on a cross-modal communication system. Background Technology
[0002] With the development of wireless communication technology and the satisfaction people have gained from traditional multimedia services such as audio and video, there will be a greater pursuit of multi-sensory immersive experiences. Therefore, multimodal services integrating audio, video, and tactile signals will gradually become mainstream in various scenarios (remote operation, online education, e-health, digital twins). For example, in remote education, audio and video services integrating tactile perception and feedback can solve the problem of ineffective remote practical teaching, improving student learning outcomes and immersive learning experiences. In virtual world interaction, tactile perception can bring users a more realistic interactive experience.
[0003] However, due to the significant differences in transmission requirements between audio / video and haptic feedback in multimodal services—high throughput is required for audio / video, while haptic feedback is more sensitive to low latency and high reliability—traditional audio / video and haptic communication methods cannot meet these demands. Therefore, cross-modal communication has emerged. Cross-modal communication fully utilizes the potential correlations between modalities and the common semantic information to establish a universal cross-modal flow scheduling scheme, and leverages AI methods to achieve efficient multimodal flow, processing, and recovery. It has also achieved initial success in application scenarios such as remote acupuncture teaching, teleoperation, and remote throat swab administration.
[0004] On the other hand, to further improve transmission efficiency, the Joint Source-Channel Coding (JSCC) method maps image pixel values to channel input symbols instead of using traditional separate coding, achieving joint optimization of the source and channel. Experiments show that the joint method outperforms communication methods based on separate schemes and can overcome the cliff effect when facing deteriorating channel environments. This method is increasingly becoming a focus of attention in academia and industry. Compared with traditional bit-based communication, by highly abstracting the semantic information in the source, data redundancy can be further compressed, improving network information transmission capabilities; and by reducing the amount of transmitted data, network information processing latency can be reduced, making it more robust to severe channels. However, JSCC is only for single-mode communication and cannot handle multi-mode situations. In summary, taking haptic-image cross-modal communication as an example, the challenges faced by existing cross-modal communication are as follows: there is a lack of a concrete and feasible framework. On the one hand, the existing communication transmission framework only focuses on single-modal transmission, such as haptic streams or audiovisual streams; on the other hand, the transmission environment faced by cross-modal communication is variable. Therefore, it is necessary to design a cross-modal transmission framework that comprehensively considers the source and channel to ensure signal transmission quality and robustness. Summary of the Invention
[0005] In view of the above-mentioned problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by this invention is to integrate a cross-modal source encoder, a channel encoder, a channel decoder, and a source decoder to achieve efficient processing and transmission of multimodal data.
[0007] To address the aforementioned technical problems, this invention provides the following technical solution: a source-channel joint encoding and decoding method based on a cross-modal communication system, comprising:
[0008] Image features and tactile features corresponding to the image signal and tactile signal to be transmitted are extracted by a cross-modal source encoder, and the image features and tactile features are fused to obtain fused features;
[0009] Design a cross-modal channel encoder and a cross-modal channel decoder. Input the fused features into the cross-modal channel encoder to convert them into a bit stream suitable for wireless transmission. After the channel transmission process is completed, the channel decoder maps the received bit stream back to the fused features.
[0010] The image signal is reconstructed from the fused features by decoding the fused features using a cross-modal source decoder.
[0011] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, the cross-modal source encoder includes a tactile feature extraction module, an image feature extraction module, and a modal fusion module;
[0012] Design a loss function to update the relevant parameters of the cross-modal source encoder; design a loss function, and use the cross-entropy loss function for image feature classification, tactile feature classification, and fused feature classification, expressed as follows:
[0013]
[0014] in, This indicates that the tactile feature extraction network was used for each step. Image feature extraction and modal fusion Output image features, tactile features, and fusion features, y j Labels representing objects, These represent tactile feature extraction networks. Image feature extraction Modal fusion and cross-modal source encoder integrated network Parameters;
[0015] By minimizing the loss function L sPreliminary training of the parameters of the cross-modal source encoder was performed.
[0016] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, the cross-modal channel encoder is represented as follows:
[0017]
[0018] in, These represent the parameters of the channel encoder model;
[0019] The cross-modal channel encoder takes a cross-modal fusion feature f as input and outputs a representation z of the feature mapped to a channel symbol.
[0020] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, the channel is modeled as a neural network layer representing the process of the encoded fused feature z being transmitted through the channel;
[0021] The channel modeling includes additive white Gaussian noise channel and Rayleigh slow fading channel;
[0022] In an AWGN channel, the channel's transfer function is expressed as follows:
[0023] y = z + n
[0024] Where n is the channel noise, which is independent and identically distributed, and σ 2 It is the average noise power;
[0025] The signal-to-noise ratio of the channel is By changing σ 2 To change the signal-to-noise ratio of the channel;
[0026] The transmission function of the Rayleigh slow fading channel is:
[0027] y = hz + n
[0028] Where h is the channel gain. n is the channel noise. h and n respectively obey the parameters as follows: and Different normal distributions.
[0029] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, wherein: the cross-modal channel decoder, after transmission through the channel, receives a complex-valued signal affected by noise interference. The estimated value mapped back to the original feature input by the channel decoder
[0030] Corresponding to the channel encoder, the channel decoder consists of a series of one-dimensional transposed convolutions, parameterized ReLU (PReLU) activation functions, and inverse normalization layers;
[0031] The channel decoder is represented as follows:
[0032]
[0033] in, These represent the parameters of the channel decoder model;
[0034] The final channel codec is jointly designed and optimized by minimizing the cross-modal fusion feature f and the fusion feature estimated after transmission. The mean square error between the two is used to jointly optimize the channel codec parameters, which is expressed as follows:
[0035]
[0036] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, wherein: the cross-modal source decoder, at the receiving end, extracts from fused features... The corresponding image signal is reconstructed;
[0037] Image reconstruction is achieved using the Wasserstein generative adversarial network method. The discriminator D adopts the PatchGAN design, and the generator G generates the image by fusing features. The discriminator D is responsible for distinguishing between the real image v and the generated image.
[0038] The loss function is defined as follows:
[0039]
[0040] in,
[0041] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, the key distribution information of the generated image is extracted from the teacher model and applied to the student model using a knowledge distillation technique.
[0042] We selected the VGG16 model pre-trained on the ImageNet dataset as the teacher model and the generator G as the student model.
[0043] Generator G from fused features To image In the generation process, each layer of the generator represents information from different levels of the image. Therefore, the feature information of the l-th layer is represented as... Represents the parameters of the l-th layer;
[0044] When an image is input into the VGG16 network, layers of different depths extract information at different levels. Deeper layers often represent deeper information, so the feature information of the m-th layer is represented as follows: Let m represent the parameters of the m-th layer, and let the hierarchical constraint loss be expressed as:
[0045]
[0046] The layers in the generator correspond to certain layer structures in VGG16, all having the same feature output size. The hierarchical constraint loss is calculated, and the loss function for the entire image generation is expressed as follows:
[0047] L genv =L Gv +δL layerloss ,
[0048] Here, δ is a hyperparameter used to balance the two loss functions.
[0049] As a preferred embodiment of the source-channel joint encoding and decoding method based on a cross-modal communication system described in this invention, the method includes: training and optimization based on a joint optimization algorithm, which includes a first stage and a second stage.
[0050] In the first stage, the cross-modal source encoder network parameters are pre-trained by minimizing the loss function. The cross-modal source encoder and the cross-modal source decoder are then combined for adversarial training against the discriminator. The loss function is calculated and the network parameters are updated by gradient until the model converges.
[0051] The second stage involves optimizing the loss function and performing a gradient descent algorithm on the parameters of the channel encoder and channel decoder until the entire system reaches a stable state.
[0052] After the above training process, the optimal cross-modal source encoder, cross-modal source decoder, channel encoder, and cross-modal channel decoder are obtained.
[0053] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the source-channel joint encoding and decoding method based on a cross-modal communication system as described above.
[0054] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the source-channel joint encoding and decoding method based on a cross-modal communication system as described above.
[0055] The beneficial effects of this invention are:
[0056] The present invention provides a source-channel joint encoding and decoding method based on a cross-modal communication system, which utilizes cross-modal information to achieve better signal feature extraction. The designed joint encoding and decoding method overcomes the cliff effect faced by traditional encoding and decoding methods. It can achieve better cross-modal source generation results. Attached Figure Description
[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a framework diagram of a source-channel joint encoding and decoding method based on a cross-modal communication system, provided as an embodiment of the present invention.
[0059] Figure 2 This is a schematic diagram illustrating the hierarchical constraints of a source-channel joint encoding / decoding method based on a cross-modal communication system, provided as an embodiment of the present invention.
[0060] Figure 3 The image shows the result of cross-modal image signal reconstruction based on a source-channel joint encoding and decoding method of a cross-modal communication system provided in one embodiment of the present invention, compared with other comparative methods.
[0061] Figure 4 The image shows the reconstruction results of cross-modal image signals under different channel conditions, based on a source-channel joint encoding and decoding method for a cross-modal communication system provided in an embodiment of the present invention. Detailed Implementation
[0062] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0063] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0064] Example 1, referring to Figures 1-2As an embodiment of the present invention, a source-channel joint encoding and decoding method based on a cross-modal communication system is provided, comprising:
[0065] Step 1: Design a cross-modal source encoder. Extract the image features and tactile features corresponding to the image signal and tactile signal to be transmitted, respectively, and fuse the image features and tactile features to obtain fused features.
[0066] Step 1-1: Design a cross-modal source encoder Including tactile feature extraction Image feature extraction and modal fusion There are three modules. Here, v represents the image signal, h represents the tactile signal, and θ represents the module parameters. The main purpose of the tactile feature extraction module and the image feature extraction module is to extract features within each modality. In the image feature extraction module, a VGG16 network pre-trained on the ImageNet dataset is used, but the fully connected layers of the VGG16 network are not included. To better utilize category information, two more fully connected layers are added after the VGG16 model, one of which serves as the feature output t of the image signal. v The last layer is used for image label classification. The output of the last fully connected layer is (None, 6), and the output of the second-to-last layer is (None, x), where the value of x was chosen as 640 after parameter experimentation. It is worth noting that during training, the parameters of the VGG16 network are frozen, and only the parameters of the last two fully connected layers are updated.
[0067] Conversely, the tactile feature extraction module and the modality fusion module employ fully connected layers of different sizes. The tactile feature extraction module takes tactile sequences as input, and therefore designs different fully connected layers to extract tactile features from the tactile data. Similar to the image module, the last layer is responsible for tactile label classification, while the penultimate layer serves as the tactile feature output t. g The output of the last fully connected layer is (None, 6), and the output of the second-to-last layer is (None, x), where the value of x was chosen as 640 after parameter experimentation.
[0068] After obtaining the features of the image and tactile signals separately, they are concatenated along the channel dimension and input into the modality fusion module for fusion. This module consists of multiple fully connected layers of different sizes. Similar to the first two modules, the output of the penultimate layer is selected as the fused feature f, while the last layer is used for label classification. The output of the last fully connected layer is (None, 6), and the output of the penultimate layer is (None, x), where the value of x was chosen as 1280 after parameter experimentation.
[0069] Steps 1-2: Accurately designing the loss function is crucial for updating the relevant parameters of the cross-modal source encoder. Therefore, three different loss functions were designed, corresponding to the three aforementioned modules, using the cross-entropy loss function for image classification, tactile feature classification, and fused feature classification. This method helps to jointly update the parameters of the cross-modal signal encoder. Details are as follows:
[0070]
[0071] in, This indicates that the tactile feature extraction network was used for each step. Image feature extraction and modal fusion The output includes image features, tactile features, and fused features. j A label representing an object. These represent tactile feature extraction networks. Image feature extraction Modal fusion and cross-modal source encoder integrated network The parameters. Among them, By minimizing the loss function L s Initial training was performed on the parameters of the cross-modal source encoder. To improve communication performance, joint training with the cross-modal decoder will be conducted next.
[0072] Step 2: Design a cross-modal channel codec, which includes a cross-modal channel encoder and a cross-modal channel decoder. After obtaining the fused features in Step 1, these features are input into the channel encoder and converted into a bit stream suitable for wireless transmission. After the channel transmission process is completed, the channel decoder maps the received bit stream back to the fused features.
[0073] Step 2-1: Design a cross-modal channel encoder. Its input is the cross-modal fusion feature f obtained in Step 1, and its output is the representation z of this feature mapped to the channel symbol. The specific channel encoder structure consists of a one-dimensional convolutional layer, a PReLU activation function, and a normalization layer. This design helps capture the statistical characteristics of the signal and effectively maps signal features to channel transmission symbols. The size of the one-dimensional convolutional layer is set to (128-256-512-z), where z represents the channel output size set in the experiment. The combination of one-dimensional convolution and nonlinear activation functions allows the model to learn the complex relationship between the signal and the channel symbol. The introduction of the normalization layer ensures that the average power of the signal meets specific power constraints. Where * denotes the complex conjugate operation. The channel encoder model is represented as: These represent the parameters of the channel encoder model.
[0074] Step 2-2: To represent the transmission of the encoded fused feature z through the channel, the channel is modeled as a series of untrainable but differentiable neural network layers. Consider two commonly used channel models: Additive White Gaussian Noise (AWGN) channel and Rayleigh slow fading channel. In an AWGN channel, the channel transfer function can be expressed as: y = z + n. n is the channel noise, which is independent and identically distributed, and... σ 2 This is the average noise power. The signal-to-noise ratio of the channel is... By changing σ 2 This is used to change the signal-to-noise ratio of the channel. The transfer function of a Rayleigh slow fading channel is: y = hz + n, where h is the channel gain. n is the channel noise. h and n respectively obey the parameters as follows: and Different normal distributions.
[0075] Steps 2-3: Design a cross-modal channel decoder to handle complex-valued signals that are affected by noise after transmission through the channel. The estimated value mapped back to the original feature input by the channel decoder Corresponding to the channel encoder, the channel decoder consists of a series of one-dimensional transposed convolutions, parameterized ReLU (PReLU) activation functions, and inverse normalization layers. The size of the one-dimensional convolutional layers is set to (5^12 - 256 - 128 - 1). The mathematical model of the channel decoder can be expressed as: These represent the parameters of the channel decoder model.
[0076] Steps 2-4: The final channel codec is jointly designed and optimized by minimizing the cross-modal fusion feature f and the fusion feature estimated after transmission. The mean square error between the two values is used to jointly optimize the channel codec parameters. This process ensures that the codec system can maintain signal accuracy and reliability even when faced with channel distortion and noise interference.
[0077]
[0078] Where N represents the total number of signals, and J represents the signal being processed.
[0079] Step 3: Design a cross-modal source decoder. Reconstruct the image signal from the fused features by decoding the fused features.
[0080] Step 3-1: Design a cross-modal source decoder. At the receiver, it is necessary to extract features from the fused data. The corresponding image signal is reconstructed. Image reconstruction is achieved using a Wasserstein Generative Adversarial Network (WGAN). The discriminator D employs a PatchGAN design. The generator G generates the image by fusing features. The discriminator D is responsible for distinguishing between the real image v and the generated image. The generator G needs to produce images that are closer to reality, while the discriminator D needs a stronger ability to distinguish between real and fake images. Therefore, the entire training process is an adversarial process between D and G. The loss function is defined here as:
[0081]
[0082] in,
[0083] Step 3-2: Furthermore, considering the difference between the real image signal v and the generated image signal... Inconsistencies in data distribution within high-dimensional pixel space can lead to blurred image signals when relying solely on reconstruction loss. To address this challenge, a knowledge distillation technique is employed to extract key distribution information from the teacher model and apply it to the student model, thereby improving the accuracy of image signal generation. A VGG16 model pre-trained on the ImageNet dataset is selected as the teacher model, and the generator G is used as the student model. Figure 2 As shown, specifically, the generator G fuses features To image In the generation process, each layer of the generator represents information from different levels of the image. Therefore, the feature information of the l-th layer is represented as... Let represent the parameters of layer l. During the image input to the VGG16 network, layers of different depths extract information at different levels; deeper layers often represent deeper information. Therefore, the feature information of layer m is represented as... Let represent the parameters of the m-th layer. Then, the hierarchical constraint loss is defined as:
[0084]
[0085] The layers in the generator correspond to certain layer structures in VGG16, all having the same feature output size, which facilitates the calculation of the hierarchical constraint loss. Finally, the loss function for the entire image generation is:
[0086] L genv =L Gv +δL layerloss ,
[0087] Here, δ is a hyperparameter used to balance the two loss functions.
[0088] Specifically, the generator consists of seven 2D convolutional layers, with BatchNormalization performed after each layer, and the activation function set to leaky_relu to better handle the non-linear relationships of features. The discriminator, on the other hand, comprises five 2D convolutional layers, also with BatchNormalization after each layer and leaky_relu as the activation function. For knowledge distillation, we chose a VGG16 model pre-trained on ImageNet as the teacher model, and the generator as the student model. In this way, we distill the knowledge from layers 1, 4, 11, 17, and 18 of the VGG16 model into layers 2, 3, 4, 6, and 7 of the generator to improve the accuracy of image generation.
[0089] Step 4 includes the following steps:
[0090] Step 4-1: Perform the first stage of image generation training. In this stage, the cross-modal source encoder network parameters are first pre-trained by minimizing the loss function. Then, starting from this, the cross-modal source encoder and cross-modal source decoder are combined and adversarial training is performed, simultaneously with adversarial training against the discriminator. The loss function is calculated and the network parameters are updated with gradients until the model converges.
[0091] Step 4-11: Initialize the cross-modal source codec network parameters Cross-modal source decoder parameters θ gv ,θ dv .
[0092] Step 4-12, according to L s Update the network parameters of the cross-modal source encoder
[0093]
[0094] Step 4-13, according to L Dv Update θ dv :
[0095]
[0096] Step 4-14, according to L genv renew θ gv :
[0097]
[0098] Steps 4-15: In summary, the training of the cross-modal source codec network is achieved.
[0099]
[0100] Step 4-2: The second phase focuses on training the channel encoder and channel decoder. This process optimizes the loss function and performs gradient descent on the parameters of the channel encoder and channel decoder until the entire system reaches a stable state. Specifically, the goal of this training phase is to minimize the impact of channel distortion on signal quality and ensure signal integrity during transmission.
[0101] Step 4-21, according to L mse renew
[0102]
[0103] Step 4-22: Optimize the channel codec.
[0104]
[0105] After the above training process, the optimal cross-modal source encoder, cross-modal source decoder, and channel encoder-decoder are obtained.
[0106] Example 2, an embodiment of the present invention, differs from the previous embodiment in that:
[0107] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0108] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0109] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0110] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0111] Example 3, referring to Figure 3 and Figure 4 As an embodiment of the present invention, a source-channel joint encoding and decoding method based on a cross-modal communication system is provided. To verify the beneficial effects of the present invention, a simulation experiment is conducted for scientific demonstration.
[0112] This embodiment uses the AU dataset, which contains multimodal information for 63 objects, including image, kinematic, and tactile data (audio vibrations). These objects are mainly toys and household items. Some objects have similar shapes and colors in images, while in tactile senses they have similar materials and textures but different colors. These objects reflect the different effects of images and touch on perception.
[0113] The existing methods are as follows:
[0114] Existing Method 1: Disco-AC GAN Model: This method uses DiscoGAN to learn the transformation between ground images and spectrograms. By utilizing trained DX and DY as discriminators to impose dual constraints on the generated results, and adding an auxiliary classification layer to the original discriminator, this approach allows the model to consider class labels during the generation process, thereby generating images that better match the characteristics of specific classes.
[0115] Existing Method 2: Cross-modal vision-touch mutual generation: A generative model based on an improved CGAN, generating images according to the semantic features encoded by touch. This model involves intermodal correlation learning to measure the correlation between the tactile modality and the image modality.
[0116] The third existing method is based on CGAN and a residual-fusion (RF) module, trained using additional feature matching and perceptual loss. The residual-fusion module helps the model capture and utilize deep features of the input data, while the perceptual loss can be used to ensure that the generated data is highly similar to the real data in the image, thereby achieving mutual generation between the image and the tactile spectrum.
[0117] This invention: The method of this embodiment.
[0118] The experiment used the Structure Similarity Index Measure (SSIM) and Peak Signal-to-Noise Ratio (PSNR) as evaluation metrics to assess the effectiveness of cross-modal generation. The experimental results are shown in Table 1.
[0119] Table 1 Comparison of Experimental Results
[0120] Method 1 Method 2 Method 3 This invention Average PSNR 30.961 26.360 23.613 27.96 AverageSSIM 0.890 0.8193 0.7979 0.856 PSNR standard deviation 7.187 5.6333 5.0238 2.87
[0121] From Table 1 and Figure 3 It can be seen that our proposed method has significant advantages compared to the state-of-the-art methods mentioned above.
[0122] Simultaneously, this invention also conducted relevant channel experiments to verify the robustness of the model to channel variations. Experimental results are as follows: Figure 4 As shown in the figure, the experimental results under different training SNR conditions were analyzed in detail. It is noteworthy that the system can still successfully recover images even when the test SNR is lower than the training SNR, i.e., when the channel conditions are poor. Under the same test SNR, the closer the training SNR is to the test SNR, the better the image quality recovered by the system. This indicates that the more consistent the training and test channel conditions are, the better the system performance. As the channel conditions improve, the difference in image recovery performance under different training SNRs gradually decreases, reflecting the model's adaptability to changes in the channel environment. This shows that with improved channel conditions, the image quality recovered by a high training SNR will be significantly improved.
[0123] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A source-channel joint encoding and decoding method based on a cross-modal communication system, characterized in that, include: Image features and tactile features corresponding to the image signal and tactile signal to be transmitted are extracted by a cross-modal source encoder, and the image features and tactile features are fused to obtain fused features; Design a cross-modal channel encoder and a cross-modal channel decoder. Input the fused features into the cross-modal channel encoder to convert them into a bit stream suitable for wireless transmission. After the channel transmission process is completed, the channel decoder maps the received bit stream back to the fused features. The fused features are decoded by a cross-modal source decoder, and the image signal is reconstructed from the fused features; The cross-modal source encoder includes a tactile feature extraction module, an image feature extraction module, and a modal fusion module; Design a loss function to update the relevant parameters of the cross-modal source encoder; design a loss function, and use the cross-entropy loss function for image feature classification, tactile feature classification, and fused feature classification, expressed as follows: in, This indicates that the tactile feature extraction network was used for each step. Image feature extraction and modal fusion Output image features, tactile features, and fusion features, y j Labels representing objects, These represent tactile feature extraction networks. Image feature extraction Modal fusion and cross-modal source encoder integrated network Parameters; By minimizing the loss function L s Preliminary training of cross-modal source encoder parameters was performed; The cross-modal channel decoder, after transmission through the channel, receives a complex-valued signal that is affected by noise interference. The estimated value mapped back to the original feature input by the channel decoder Corresponding to the channel encoder, the channel decoder consists of a series of one-dimensional transposed convolutions, parameterized ReLU (PReLU) activation functions, and inverse normalization layers; The channel decoder is represented as follows: in, These represent the parameters of the channel decoder model; The final channel encoder and channel decoder are jointly designed and optimized by minimizing the cross-modal fusion feature f and the fusion feature estimated after transmission. The mean square error between the two is used to jointly optimize the channel codec parameters, which is expressed as follows:
2. The source-channel joint encoding and decoding method based on a cross-modal communication system as described in claim 1, characterized in that: The cross-modal channel encoder is represented as follows: in, These represent the parameters of the channel encoder model; The cross-modal channel encoder takes cross-modal fusion features f as input and outputs a representation z of the cross-modal fusion features f mapped to channel symbols.
3. The source-channel joint encoding and decoding method based on a cross-modal communication system as described in claim 2, characterized in that: The channel is modeled as a neural network layer representing the process of the encoded fused feature z being transmitted through the channel; The channel modeling includes additive white Gaussian noise channel and Rayleigh slow fading channel; In an AWGN channel, the channel's transfer function is expressed as follows: y = z + n Where n is the channel noise, which is independent and identically distributed, and σ 2 It is the average noise power; The signal-to-noise ratio of the channel is By changing σ 2 To change the signal-to-noise ratio of the channel; The transmission function of the Rayleigh slow fading channel is: y = hz + n Where h is the channel gain. n is the channel noise. h and n respectively obey the parameters as follows: and Different normal distributions.
4. The source-channel joint encoding and decoding method based on a cross-modal communication system as described in claim 3, characterized in that: The cross-modal source decoder, at the receiving end, extracts features from the fused data. The corresponding image signal is reconstructed; Image reconstruction is achieved using the Wasserstein generative adversarial network method. The discriminator D adopts the PatchGAN design, and the generator G generates the image by fusing features. The discriminator D is responsible for distinguishing between the real image v and the generated image. The loss function is defined as follows: in, 5. The source-channel joint encoding and decoding method based on a cross-modal communication system as described in claim 4, characterized in that: The key distribution information of the generated image is extracted from the teacher model using the knowledge distillation technique and applied to the student model. The VGG16 model pre-trained on the ImageNet dataset was chosen as the teacher model, and the generator G was chosen as the student model. Generator G from fused features To image In the generation process, each layer of the generator represents information from different levels of the image. Therefore, the feature information of the l-th layer is represented as... Represents the parameters of the l-th layer; When an image is input into the VGG16 network, layers of different depths extract information at different levels. Deeper layers often represent deeper information, so the feature information of the m-th layer is represented as follows: Let m represent the parameters of the m-th layer, and let the hierarchical constraint loss be expressed as: The layers in the generator correspond to certain layer structures in VGG16, all having the same feature output size. The layer constraint loss is calculated, and the loss function for the entire image generation is expressed as follows: L genv =L Gv +δL layerloss , Here, δ is a hyperparameter used to balance the two loss functions.
6. The source-channel joint encoding and decoding method based on a cross-modal communication system as described in claim 5, characterized in that: Training optimization is performed based on a joint optimization algorithm, which includes a first stage and a second stage. In the first stage, the cross-modal source encoder network parameters are pre-trained by minimizing the loss function. The cross-modal source encoder and the cross-modal source decoder are then combined for adversarial training against the discriminator. The loss function is calculated and the network parameters are updated by gradient until the model converges. The second stage involves optimizing the loss function and performing a gradient descent algorithm on the parameters of the channel encoder and channel decoder until the entire system reaches a stable state. After the above training process, the optimal cross-modal source encoder, cross-modal source decoder, channel encoder, and cross-modal channel decoder are obtained.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the source-channel joint encoding and decoding method based on any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the source-channel joint encoding and decoding method based on any one of claims 1 to 6 for a cross-modal communication system.
Citation Information
Patent Citations
Visual-auditory cross-modal object material retrieval method and system
CN108520758A
Cross-modal image generation method and device based on audio-tactile signal fusion
CN113627482A