A haptic signal reconstruction method based on joint source-channel coding
By using a source-channel joint coding method, tactile and image signal features are extracted and fused, and the channel codec is optimized. This solves the problems of low latency, high reliability, and channel noise in tactile signal transmission, and achieves efficient and robust tactile signal reconstruction.
Patent Information
- Application Number
- CN202411075761.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-08-07
AI Technical Summary
In existing cross-modal communication, the requirements for low latency and high reliability in the transmission of tactile signals have not been fully met, and the impact of channel noise has been insufficiently studied, resulting in a decline in data quality.
A source-channel joint coding method is designed. The source encoder extracts and fuses tactile and image signal features, the channel encoder-decoder learns channel changes, and the time-domain and frequency-domain loss functions are combined to optimize the reconstruction of tactile signals.
It achieves efficient compression and robust transmission of tactile signals, reduces data volume while maintaining information integrity and accuracy, and adapts to stable and reliable transmission under various channel conditions.
Smart Images

Figure CN119011084B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of tactile signal generation, in particular to a tactile signal reconstruction method based on joint coding of signal source and channel. BACKGROUND
[0002] With the advancement of technology and the growing demand of users, traditional audio-visual sensory experience has been unable to meet the pursuit of users for immersive and interactive communication methods, thus prompting multimedia services to develop to a higher level. However, existing researches mainly focus on the processing and transmission of single modal signals, and fail to fully adapt to the diversified transmission needs of heterogeneous multi-modal signals. Cross-modal services thus emerge as the times require, which provides users with a new mode of communication and experience by integrating visual, auditory and tactile sensory information. In particular, cross-modal services dominated by touch, which provide direct and real interactive feedback, are regarded as an important development direction of future communication technology. Current research on cross-modal services dominated by touch still faces some challenges.
[0003] On the one hand, tactile signals have high transmission requirements, need low-delay and high-reliability transmission conditions, and have high sensitivity. Tactile signals have high requirements for transmission quality, otherwise it will directly affect the immersive experience or the accuracy of remote control. Tactile signal transmission requires low latency, because the human tactile system responds very quickly, we can almost feel the touch or pressure change instantly. In order to reproduce this instantaneity in remote interaction or virtual environment, the transmission of tactile signals must be fast enough to match or approach the speed of human perception. The low-latency and high-reliability characteristics of tactile signals pose a challenge to cross-modal services.
[0004] On the other hand, the current cross-modal signal reconstruction scheme has not fully studied the influence of channel noise, and needs to further consider the robustness of signal reconstruction under different channel conditions. In actual communication scenarios, tactile signals need to be encoded at the signal source end, transmitted in the channel, and decoded and reconstructed at the receiving end. This process involves the interaction of multiple factors, including semantic fusion of cross-modal signals, transmission characteristics of the channel, and reconstruction strategies at the receiving end, etc.
[0005] Existing research has also made some progress in the cross-modal fusion and generation of haptic signals. In the literature "Information recovery technology for cross-modal communication", Xu et al. proposed an information recovery scheme for the problem of loss or noise pollution of wireless channels that multi-modal data (such as audio, video, and haptic signals) may encounter during transmission. The scheme uses existing multi-modal data at the receiving end to achieve information recovery through same-modal one-to-one search, cross-modal one-to-one search, cross-modal one-to-many search, and other methods. Wei et al. proposed an audio-visual assisted haptic signal reconstruction method based on cloud-edge collaboration in the literature "Haptic Signal Reconstruction for Cross-Modal Communications". For the case of insufficient data received by the edge node, a large audio-visual auxiliary database stored in the central cloud is used to obtain and transmit knowledge and migrate to the edge node. Then, the edge node is used to mine the inherent semantic consistency and correlation between modalities, and finally the required haptic signal reconstruction is achieved. By considering the received audio-visual signals and linking them with the potential correlation of the haptic modality, the reconstruction of the haptic signal is achieved.
[0006] However, these models are mostly limited to one-way conversion at the sending end or receiving end, without considering end-to-end design of the entire communication process, and lack simulation of transmission in actual channels. In actual situations, it is easy to be disturbed by noise, resulting in a decline in the quality of the transmitted data. This will have an adverse effect on multi-modal data streams, especially haptic streams, which require low latency and high reliability. In addition, how to design effective data compression and coding technology to significantly reduce the transmission data volume of haptic signals while ensuring the quality of haptic perception is a difficult problem faced by haptic-based cross-modal communication. SUMMARY
[0007] In view of the above problems, the present application is proposed.
[0008] Therefore, the technical problem solved by the present application is: in order to meet the requirements of haptic signal communication, a source codec is designed to fully exploit the inherent semantic correlation and consistency between modalities, fully utilize the haptic information hidden in the visual modality, and fuse the semantic information of different modalities with each other to significantly compress the source signal and reduce the data volume. On the other hand, in order to improve the robustness to channel changes and reduce the delay of signals through the channel, a channel codec is designed to learn the channel changes in the communication process. Finally, a loss function is designed at the decoding end from the time domain and the frequency domain, respectively, to better recover the haptic signal.
[0009] To solve the above technical problems, the present application provides the following technical solutions: a haptic signal reconstruction method based on joint encoding of source and channel, comprising:
[0010] inputting the haptic signal and the image signal into a source encoder, the source encoder comprising modal feature extraction and modal feature fusion;
[0011] extracting features of the haptic signal and the image signal through a classification task, respectively by a haptic feature extraction network and an image feature extraction network; inputting the extracted haptic features and image features into a modal feature fusion network to perform modal feature fusion, and obtaining fusion features about the target object;
[0012] inputting the fusion features into a channel encoder to obtain code words suitable for channel transmission, and after channel transmission, inputting the code words into a channel decoder to recover the received fusion features;
[0013] inputting the recovered fusion features into a haptic decoder, setting loss constraints from time domain and frequency domain angles respectively, optimizing related parameters in the haptic decoder, and realizing reconstruction of the target haptic signal;
[0014] training the source encoder, jointly training a channel encoder and a channel decoder, and training the haptic decoder to obtain optimal haptic reconstruction model parameters;
[0015] inputting the to-be-tested haptic and image signals into the optimal source encoder, extracting features of the haptic signal and the image signal and fusing the features, inputting the obtained fusion features into the haptic decoder after optimal channel encoding, channel transmission and optimal channel decoding, and generating the target haptic signal.
[0016] As a preferred scheme of the haptic signal reconstruction method based on source channel joint coding, the haptic signal and the image signal are inputted into the source encoder, and a haptic and image signal set is constructed wherein h j ,v j ,y j represent the haptic signal and the image signal of the jth object and the corresponding class label, respectively, and the haptic signal and the image signal in the signal set are inputted into the source encoder network;
[0017] The source encoder is represented as which receives the image and the haptic signal, extracts haptic features and image features from the image and the haptic signal, respectively, and fuses the haptic features and the image features to obtain fusion features, wherein the haptic feature extraction the image feature extraction and the modal fusion
[0018] After the image features and the haptic features are obtained, the image features and the haptic features are spliced together in the channel dimension and inputted into the modal fusion module for fusion;
[0019] The related parameters of the source encoder are updated through the classification task of the haptic and image signal set.
[0020] The source encoder parameters are jointly updated using the cross-entropy loss function applied to image classification, tactile classification, and fused feature classification, as shown below.
[0021]
[0022] in, This indicates that the tactile feature extraction network was used for each step. Image feature extraction network and modal fusion The output image features, tactile features, and fusion features, y j Indicates the category label of the current tactile-image signal. These represent tactile feature extraction networks. Image feature extraction network Modal fusion network The parameters, Indicates source encoder integrated network The parameters,
[0023] By minimizing the loss function Update the source encoder parameters.
[0024] As a preferred embodiment of the tactile signal reconstruction method based on source-channel joint coding described in this invention, the recovered fused features include obtaining fused features after passing through the source encoder; and being directly mapped by the channel encoder into a symbol stream for channel transmission.
[0025] After being mapped by the channel encoder network, the symbol stream needs to be sent into the channel for transmission. The real and imaginary parts of the complex vector z input to the channel are mapped onto the I and Q components of the digital signal, respectively, and then transmitted through the communication channel.
[0026] The complex-valued vector z is subject to channel noise interference. End-to-end optimization is performed by modeling the channel as a neural network layer. The channel modeling includes AWGN channels and Rayleigh slow fading channels.
[0027] As a preferred embodiment of the tactile signal reconstruction method based on joint source-channel coding described in this invention, wherein the transfer function of the AWGN channel is expressed as follows:
[0028] y = z + n,
[0029] Where n is the channel noise, which is independent and identically distributed, and σ 2 This is the average noise power, and the channel's signal-to-noise ratio (SNR) is SNR = 10log 10 (1 / σ2 ), by changing 2 the signal-to-noise ratio of the channel;
[0030] The transmission function of the Rayleigh slow fading channel is expressed as,
[0031] y = hz + n,
[0032] where h is the channel gain, n is the channel noise, h and n respectively obey different normal distributions with parameters and
[0033] After transmission through the channel, the complex-valued signal disturbed by noise is mapped by the channel decoder into the estimated value of the fusion feature
[0034] The channel decoder model is expressed as,
[0035]
[0036] where denotes the parameters of the channel decoder model;
[0037] The channel encoder and the channel decoder are jointly designed and optimized, and the parameters of the channel encoder and decoder are jointly updated by minimizing the mean square error of the fusion feature f and the estimated value of the fusion feature after transmission , expressed as,
[0038]
[0039] As a preferred scheme of the haptic signal reconstruction method based on joint source-channel coding according to the present application, the reconstruction of the target haptic signal includes that the haptic decoder reconstructs the haptic signal from the fusion semantic feature, expressed as
[0040] The haptic decoder network adopts multiple fully connected layers and dropout layers of different sizes to extract haptic information from the fusion feature, so as to recover the haptic loss function set as the average absolute error and the average square error between the reconstructed haptic and the real haptic H, expressed as,
[0041]
[0042] Starting from the frequency domain, the short-time Fourier transform (STFT) loss function is set to reflect the frequency domain characteristics of the signal;
[0043] The loss function includes a spectrum convergence loss and log magnitude loss
[0044] The spectrum convergence loss is represented as,
[0045]
[0046] The log magnitude loss is represented as,
[0047]
[0048] where ||·|| and ||·||1 represent F-norm and L1-norm operations, STFT(·) represents an STFT transform, and Num represents the number of elements in the STFT magnitude; F and ||·||1 represent F-norm and L1-norm operations, STFT(·) represents an STFT transform, and Num represents the number of elements in the STFT magnitude;
[0049] The STFT loss function is represented as,
[0050]
[0051] The loss function for haptic signal generation is represented as,
[0052]
[0053] where a, b, and g are hyperparameters.
[0054] As a preferred scheme of the haptic signal reconstruction method based on joint coding of a source channel according to the present application, the optimal haptic reconstruction model parameters are obtained by pre-training the parameters of the source encoder network by minimizing the loss function .
[0055] The parameters of the source encoder network are fixed by minimizing the loss on the channel encoder and the channel decoder , and gradient updates are performed until the model converges.
[0056] The fused semantic features are obtained by the source encoder, and input into the haptic decoder network G h (f; q gh ), and gradient updates are performed on the parameters of the haptic decoder network by minimizing until the model converges.
[0057] As a preferred scheme of the haptic signal reconstruction method based on joint coding of a source channel according to the present application, the pre-training of the parameters of the source encoder network includes initializing the parameters of the source encoder network
[0058] input training signal set set the number of iterations to n1 and the learning rate to μ1;
[0059] According to update the source encoder network parameters denoted as
[0060]
[0061] wherein is the partial derivative of each loss function; until the number of iterations n1 is reached to complete the training;
[0062] If n < n1, n = n + 1, continue the next iteration; otherwise, terminate the iteration; after n1 iterations, the optimized source encoder network parameters are obtained
[0063] The gradient update on the parameters of the channel encoder and the channel decoder includes initializing the channel encoder network parameters and the channel decoder network parameters
[0064] Take the fusion features output by the source encoder as input, set the number of iterations to n2 and the learning rate to μ2;
[0065] Fix the source encoder network parameters Optimize and denoted as
[0066]
[0067] If n < n2, n = n + 1, continue the next iteration; otherwise, terminate the iteration; after n2 iterations, the optimized channel encoder network parameters and the channel decoder network parameters are obtained
[0068] The gradient update on the parameters of the channel encoder includes initializing the channel encoder network parameters θ gh ;
[0069] Take the fusion features output by the source encoder as input, set the number of iterations to n3 and the learning rate to μ3;
[0070] Fix the source encoder network parameters Optimize θ gh , denoted as
[0071]
[0072] If n < n3, n = n + 1, continue the next iteration; otherwise, terminate the iteration; after n3 rounds of iteration, the optimized haptic decoder network parameter θ is obtained gh .
[0073] As a preferred scheme of the haptic signal reconstruction method based on source-channel joint coding according to the present application, the generating target haptic signal comprises that the source encoder extracts key features from the input data and generates fusion features, the fusion features are further processed by the channel encoder, the encoded data is then transmitted through the channel and decoded by the channel decoder to recover the original fusion features, and the haptic decoder network generates the expected haptic signal by using the recovered fusion features to complete the entire haptic signal reconstruction.
[0074] A computer device comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the haptic signal reconstruction method based on source-channel joint coding according to the present application when executing the computer program.
[0075] A computer readable storage medium having a computer program stored thereon, the computer program is executed by a processor to implement the steps of the haptic signal reconstruction method based on source-channel joint coding according to the present application.
[0076] The haptic signal reconstruction method based on source-channel joint coding according to the present application fully utilizes the advantages of multi-modal feature fusion, and effectively fuses different modal semantic information. Not only can it significantly compress the source signal and reduce the data volume, but also can maintain the integrity and accuracy of the information. Further, in order to cope with the influence of channel noise and changes on signal transmission, the present application specially designs a set of channel encoder and decoder. The encoder and decoder can improve the robustness of the signal to channel changes, and ensure that the signal remains stable and reliable under various channel conditions. BRIEF DESCRIPTION OF DRAWINGS
[0077] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0078] Figure 1 A haptic signal reconstruction method based on source-channel joint coding according to an embodiment of the present application is provided.
[0079] Figure 2A network structure schematic diagram of a haptic signal reconstruction method based on source channel joint coding provided for a second embodiment of the present application.
[0080] Figure 3 A haptic signal reconstruction result diagram of a haptic signal reconstruction method based on source channel joint coding provided for a second embodiment of the present application and other comparative methods.
[0081] Figure 4 A model haptic reconstruction result diagram of a haptic signal reconstruction method based on source channel joint coding provided for a second embodiment of the present application under AWGN channel conditions and different training signal-to-noise ratios.
[0082] Figure 5 A model haptic reconstruction result diagram of a haptic signal reconstruction method based on source channel joint coding provided for a second embodiment of the present application under Rayleigh channel conditions and different training signal-to-noise ratios. DETAILED DESCRIPTION
[0083] In order to make the above objectives, characteristics and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.
[0084] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced in other manners different from those described herein, and those skilled in the art can make similar generalizations without departing from the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0085] Embodiment 1, with reference to Figures 1-2 For an embodiment of the present application, a haptic signal reconstruction method based on source channel joint coding is provided, comprising:
[0086] Step 1: input the haptic signal and the image signal into the source encoder, the source encoder includes two parts of modal feature extraction and modal feature fusion, as shown in Figure 2 illustrated. First, through a classification task, the haptic feature extraction network and the image feature extraction network extract the features of the haptic and image signals respectively. Then, the extracted haptic and image features are input into the modal feature fusion network for modal feature fusion to obtain the fusion features about the target object; through the learning of the source encoder, the fusion representation reflecting the key features of the target object can be obtained, and this fusion representation greatly compresses the data amount;
[0087] Step 1-1, for the haptic-image signal set where h j , v j , y j represent the haptic signal and image signal of the jth object and its corresponding class label, respectively, and the haptic signal and image signal pair in the signal set is input into the source encoder;
[0088] Step 1-2, the source encoder is represented as: which receives image and haptic signals, extracts image features and image features from them, respectively, and fuses them to obtain fused features. It mainly includes three modules of haptic feature extraction image feature extraction and modal fusion The first two modules are mainly used to capture intra-modal features, and the last module is used to fuse multiple modalities.
[0089] where the image feature extraction module adopts the VGG16 network pre-trained on the ImageNet dataset (excluding the fully connected layer in the VGG16 network), and in order to better utilize the class information, two fully connected layers are added after the VGG16 model, one as the output of the image feature t v , and the last layer is used for image label classification. Note that during training, the parameters of the VGG16 network are frozen, and only the parameters of the last two fully connected layers are updated
[0090] The haptic feature extraction adopts 4 layers of fully connected layers (1024→512→l h →6). Among them, l h represents the length of the haptic feature, and 6 represents the number of article categories. The last layer performs haptic label classification, and the second to last layer is used as the output of the haptic feature t h . The feature fusion adopts 3 layers of fully connected layers (1280→1280→6), and the last layer is used for fusion feature classification, and the second to last layer is taken as the fusion feature output.
[0091] Step 1-3, after obtaining the image feature and haptic feature, they are spliced together in the channel dimension and input into the modal fusion module for fusion, which is composed of multiple fully connected layers of different sizes. Like the first two modules, the second to last layer is taken as the fusion feature f, and the last layer is used for label classification.
[0092] Step 1-4, update the related parameters of the source encoder through the classification task of the haptic-image signal set. Here, three loss functions are designed to correspond to the above three modules, and the cross-entropy loss functions of image classification, haptic classification and fusion feature classification are used to jointly update the parameters of the source encoder. The specific is as follows:
[0093]
[0094] wherein, represent the image feature, the haptic feature and the fused feature respectively after passing through the image feature extraction network the image feature extraction network and the modal fusion network respectively. j represents the class label to which the current haptic-image signal belongs. represent the parameters of the haptic feature extraction network the image feature extraction network the modal fusion network respectively; in addition, represents the parameters of the source encoder integrated network .
[0095] The source encoder parameters are updated by minimizing the loss function .
[0096] Step 2: The fused feature is input into the channel encoder to obtain the code word suitable for channel transmission. After channel transmission, it enters the channel decoder to recover the accepted fused feature, as shown in Figure 2 . Among them, the channel encoder learns the mapping from the haptic and image fused feature space to the channel transmission symbol space, and the channel decoder learns the mapping from the channel transmission symbol space to the fused feature space. Then, the channel is modeled as a series of untrainable, differentiable neural network layers. Through the joint training of the channel encoder and decoder, the model can learn from the channel to better cope with the influence of channel noise and fading, so that the channel decoder end can better recover the fused feature;
[0097] Step 2-1, After passing through the source encoder, the fused feature is obtained. Then it is directly mapped by the channel encoder into a symbol stream for channel transmission. Specifically, the channel encoder is composed of a series of one-dimensional convolution, parameterized ReLU (PReLU) activation function and normalization layer. One-dimensional convolution and nonlinear activation function can effectively learn the characteristics of the fused feature, and realize the complex nonlinear mapping of the fused feature to the channel transmission symbol stream. Specifically, in the channel decoder, after data normalization, there are 4 layers of one-dimensional convolution, and the activation function is set to parameterized ReLU (PReLU). The last layer of the channel decoder outputs is composed of 2k units, followed by another normalization layer, which performs average power constraint on the signal. The output of the normalization layer is combined into k complex-valued channel input samples to form the encoded signal representation. The channel encoder maps the semantic feature f of length n into a complex-valued vector z of length k. In order to be able to transmit in the physical channel, a normalization layer needs to be set at the last layer of the encoder to normalize the power and meet the average power constraint: denotes complex conjugate operation. The channel encoder model is denoted as: denotes the parameters of the channel encoder model.
[0098] Step 2-2, After mapping through the channel encoder network, the symbol stream needs to be sent to the channel for transmission, where the real and imaginary parts of the complex-valued vector z of the channel input need to be mapped to the I and Q components of the digital signal respectively, and then transmitted through the communication channel. In this process, the complex-valued vector z will be disturbed by channel noise. In order to be able to optimize end-to-end, the channel is therefore modeled as a series of untrainable, differentiable neural network layers. Two widely used channel models are considered: 1) Additive White Gaussian Noise (AWGN) channel, 2) Rayleigh slow fading channel. The transmission function of the AWGN channel is:
[0099] y = z + n,
[0100] n is the channel noise, which is independent and identically distributed, and σ 2 is the average noise power. The signal-to-noise ratio of the channel is SNR = 10log 10 (1 / σ 2 ). The signal-to-noise ratio of the channel is changed by changing σ 2 . The transmission function of the Rayleigh slow fading channel is:
[0101] y = hz + n,
[0102] h is the channel gain, n is the channel noise, h and n respectively obey different normal distributions with parameters and . It should be noted that in order to be able to perform gradient descent and backpropagation during the training of the neural network, the transfer function of the channel must be differentiable.
[0103] Step 2-3, After transmission through the channel, the noisy complex-valued signal is mapped to the estimate of the fused feature The channel decoder is composed of a series of one-dimensional transpose convolution, parameterized ReLU (PReLU) activation function and inverse normalization layer, corresponding to the channel encoder structure. Specifically, the channel decoder combines the real and imaginary parts of the k complex-valued noisy channel output samples into 2k values, and inputs these values into the corresponding 4-layer one-dimensional transpose convolution to gradually recover the estimate of the fused feature. The channel decoder model is denoted as: denotes the parameters of the channel decoder model.
[0104]
[0105] The channel encoder and channel decoder are jointly designed and jointly optimized by minimizing the fused feature f and the estimated fused feature after transmission. Mean square error To jointly update the parameters of the channel codec, where N represents the number of samples;
[0106] Step 3: Input the recovered fused features into the tactile decoder network, and optimize the parameters of the tactile decoder network by setting loss constraints from the time domain and frequency domain perspectives respectively, so as to realize the reconstruction of the target tactile signal;
[0107] Step 3-1: After transmission through the channel, at the receiving end, the corresponding tactile modal signal needs to be reconstructed from the fused features. The tactile decoder needs to reconstruct the tactile signal from the fused semantic features. This is represented as... Since tactile signals are simpler than RGB image signals, the tactile decoder network employs multiple fully connected layers and dropout layers of varying sizes to extract tactile information from the fused features, thereby recovering the tactile information. The tactile loss function is set to reconstruct the tactile information. The mean absolute error (MAE) and mean squared error (MSE) between the actual tactile sensation h and the real tactile sensation h. Expressed as:
[0108]
[0109] Where N represents the number of samples.
[0110] Step 3-2: To better reconstruct the details of the tactile signal, inspired by speech synthesis methods, a Short Time Fourier Transform (STFT) loss function is set from the frequency domain to reflect the frequency domain characteristics of the signal. This makes the characteristics of the reconstructed signal in both the time and frequency domains closer to the real signal. Specifically, the STFT loss function consists of two parts: spectral convergence loss... And logarithmic magnitude loss
[0111]
[0112]
[0113] Among them, ||·|| F The operators ||·||1 represent the F-norm and L1-norm operations. STFT(·) represents the STFT transform, and Num represents the number of elements in the STFT amplitude.
[0114] Spectral convergence loss The main focus is on large-scale amplitude variations to make the generated signal resemble the real signal; logarithmic amplitude loss. The primary focus is on small-scale amplitude variations, aiming to make the generated signal closely resemble the real signal. Therefore, the overall STFT loss function is:
[0115]
[0116] Finally, the loss function for the entire tactile signal generation is:
[0117]
[0118] Here, α, β, and γ are hyperparameters.
[0119] Step 4: First, train the source encoder, then train the channel encoder and channel decoder together, and finally train the haptic decoder to obtain the optimal haptic reconstruction model parameters.
[0120] Step 4-1: First, minimize the loss function. First, pre-train the source encoder network. The parameters. The specific process is as follows:
[0121] Step 4-11: Initialize the source encoder network parameters
[0122] Step 4-12: Input training signal set Set the number of iterations to n1 and the learning rate to μ1;
[0123] Step 4-13, according to Update source encoder network parameters
[0124]
[0125] in, To perform partial derivatives with respect to each loss function; training continues until the number of iterations n1 is reached to complete the training.
[0126] Step 4-14: If n < n1, then jump to step 413, n = n + 1, and continue to the next iteration; otherwise, terminate the iteration.
[0127] Step 4-15: After n1 iterations, the optimized source encoder network parameters are obtained.
[0128] Step 4-2: Fix the source encoder network parameters By minimizing the loss Channel encoder and channel decoder The parameters are updated using gradients until the model converges. The specific process is as follows:
[0129] Step 4-21, initialize the channel encoder network parameters and the channel decoder network parameters
[0130] Step 4-22, input the fusion features output by the source encoder as input, set the number of iterations to n2, and the learning rate to μ2;
[0131] Step 4-23, fix the source encoder network parameters Optimize and
[0132]
[0133] Step 4-24, if n < n2, jump to step 423, n = n + 1, and continue the next iteration; otherwise, terminate the iteration;
[0134] Step 4-25, after n2 iterations, obtain the optimized channel encoder network parameters and the channel decoder network parameters
[0135] Step 4-3, obtain the fusion semantic features through the source encoder, input the haptic decoder network G h (f; θ gh ), and perform gradient update on the haptic decoder network parameters by minimizing until the model converges. The specific process is as follows:
[0136] Step 4-31, initialize the haptic decoder network parameters θ gh ;
[0137] Step 4-32, input the fusion features output by the source encoder as input, set the number of iterations to n3, and the learning rate to μ3;
[0138] Step 4-33, fix the source encoder network parameters Optimize θ gh :
[0139]
[0140] Step 4-34, if n < n3, jump to step 4323, n = n + 1, and continue the next iteration; otherwise, terminate the iteration;
[0141] Step 4-35, after n3 iterations, obtain the optimized haptic decoder network parameters θ gh ;
[0142] Step 5: The optimal model is inputted with the to-be-tested haptic and image signal pair. The optimal model extracts the features of the haptic signal and the image signal and fuses them, and then, after channel transmission, the received fused features are used to generate the target haptic signal.
[0143] Step 5-1: Select the trained optimal source encoder, channel encoder, channel decoder, and haptic decoder network model;
[0144] Step 5-2: Input the to-be-tested haptic-image signal pair into the model. First, pass through the source encoder, which is responsible for extracting key features from the input data and generating fused features. This step not only significantly compresses the original data volume but also provides a more compact data representation for subsequent processing. Then, the fused features are further processed by the channel encoder to enhance the robustness of the data to channel noise during transmission. The encoded data is then transmitted through the channel and decoded by the channel decoder to recover the original fused features. Finally, the haptic decoder uses the recovered fused features to generate the desired haptic signal, thus completing the entire haptic signal reconstruction.
[0145] Embodiment 2, one embodiment of the present application, which is different from the previous embodiment is:
[0146] If the function is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application or the parts that essentially contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0147] The logic and / or steps represented in the flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable instructions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber (optical), and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted, or otherwise processed, and stored in a computer memory in a form that can be later executed by a computer. In this context, a "computer-readable medium" can be any means that can store the program for use by or in connection with the instruction execution system, apparatus, or device.
[0148] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber (optical), and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example via an optical scanner, then compiled, interpreted, or otherwise processed, and stored in a computer memory in a form that can be later executed by a computer.
[0149] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), and / or the like.
[0150] Example 3, with reference to Figures 3-5 For one embodiment of the present application, a haptic signal reconstruction method based on source channel joint coding is provided. In order to verify the beneficial effects of the present application, economic benefit calculation and simulation experiments are used for scientific demonstration.
[0151] The present embodiment uses the AU dataset for experiments, which is collected by Lasse et al. from the Department of Electrical and Computer Engineering, Aarhus University. The dataset contains multimodal data of 63 objects: vision, motion, and touch (audio vibration data). The objects selected in the dataset are mainly toys and household items. Each object contains 3 pictures at different angles and 5 recorded touch vibration signals. Among them, there are objects with similar shapes and colors (visual ambiguity) and objects with similar materials and textures but different colors (touch ambiguity). These objects with visual ambiguity and touch ambiguity will reflect the different effects of vision and touch on perception. The present example selects 62 available objects, and augments 3 images of each object to 5 (for example, randomly adjust the brightness and contrast of the picture, flip the image), and 5 recorded touch vibration signals to form 310 groups of samples. In the dataset, 80% is selected for training, and the remaining 20% is used for testing and performance evaluation.
[0152] First, test the tactile reconstruction ability of the model simply by removing the channel and channel codec, and select the following three tactile reconstruction methods for experimental comparison:
[0153] Existing method one: In the literature “Cross-Modal Data Generation Using Residue-Fusion GAN With Feature-Matching and Perceptual Losses” (authors Cai S, Zhu K, Ban Y, et al), a cGAN (conditional-GAN) structure and a residual-fusion (RF) module are used to train the model using additional feature matching (FM) and perceptual loss to realize the mutual generation between visual pictures and tactile frequency spectrum.
[0154] Existing method two: In the literature “Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired People” (authors Liu H, Guo D, Zhang X, et al), DiscoGAN is used to learn the conversion between images and frequency spectrum. By using the trained DX and DY as discriminators to double constrain the generated results, and adding an auxiliary classification layer to the original discriminator, the category information is needed, and the model can learn the mutual conversion between the vision domain and the touch domain.
[0155] The third existing method: in the document "Learning cross-modal visual-tactile representation using ensembled generative adversarial networks" (author: Li X, Liu H, Zhou J, et al.), based on deep learning, the texture attribute of the image is taken as the input, the category information of the image is obtained, and then it is taken as the selection condition of the GAN, so as to realize the cross-modal conversion of the texture image to the tactile signal through the corresponding GAN training.
[0156] The present application: the method of the embodiment.
[0157] In the experiment, the similarity (SIM) is taken as the evaluation index to evaluate the effect of cross-modal generation. SIM is used to measure the similarity between the real tactile signal and the reconstructed tactile signal, and is defined as:
[0158]
[0159] Wherein, the exponential in the denominator h i (h i ′) is the Chebyshev distance between h i and h i ′ waveform. That is, the maximum value of the distance between two waveform points. The range of SIM is [0, 1]. When the difference between the two is small, SIM is closer to 1, and the two waveforms are more similar, and vice versa.
[0160] Table 1 is the experimental results of the present application
[0161]
[0162] From Table 1 and Figure 3 It can be seen that compared with the above-mentioned most advanced method, the method proposed by us has obvious advantages. All other comparative models input only visual signals, while our model fuses visual and tactile, extracts high-level semantic information of the same object from two modalities, better extracts common information between modalities and unique information within modalities, and thus better restores the tactile signal.
[0163] Secondly, in the case of considering the channel, the model is trained under the condition that the signal-to-noise ratio is 0dB, 5dB, 10dB, 15dB and 20dB, and tested under different test signal-to-noise ratios. Considering the conditions of AWGN channel and Rayleigh channel. Figure 4 The tactile reconstruction effect diagram of the model under the condition of AWGN channel of the present model and different training signal-to-noise ratios. Figure 5 The tactile reconstruction effect diagram of the model under the condition of Rayleigh channel of the present model and different training signal-to-noise ratios.
[0164] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not to limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. A method for reconstructing tactile signals based on joint source-channel coding, characterized in that, include: Tactile signals and image signals are input into a source encoder, which includes modal feature extraction and modal feature fusion. Through a classification task, a tactile feature extraction network and an image feature extraction network extract features from tactile signals and image signals, respectively. The extracted tactile features and image features are then input into a modal feature fusion network for modal feature fusion to obtain fused features about the target object. The fused features are input into the channel encoder to obtain codewords suitable for channel transmission. After channel transmission, they enter the channel decoder to recover the received fused features. The recovered fusion features are input into the tactile decoder. By setting loss constraints from the time domain and frequency domain perspectives respectively, the relevant parameters in the tactile decoder are optimized to achieve the reconstruction of the target tactile signal. The source encoder is trained, the channel encoder and channel decoder are trained together, and the tactile decoder is trained to obtain the optimal tactile reconstruction model parameters. The tactile and image signals to be tested are input into the optimal source encoder to extract and fuse the features of the tactile and image signals. After passing through the optimal channel encoder, channel transmission, and optimal channel decoder, the fused features are input into the tactile decoder to generate the target tactile signal. The fusion features are directly mapped by the channel encoder into a symbol stream for channel transmission; After being mapped by the channel encoder network, the symbol stream needs to be sent into the channel for transmission. The real and imaginary parts of the complex vector z input to the channel are mapped onto the I and Q components of the digital signal, respectively, and then transmitted through the communication channel. The complex-valued vector z models the channel as a series of untrainable, differentiable neural network layers, including AWGN channels and Rayleigh slow fading channels.
2. The tactile signal reconstruction method based on joint source-channel coding as described in claim 1, characterized in that: The input of tactile signals and image signals into the signal source encoder includes constructing a set of tactile and image signals. Where N represents the number of objects, h j ,v j ,y j The tactile signal and image signal of the j-th object, and their corresponding category labels, are respectively input into the source encoder network. The source encoder is represented as It receives images and tactile signals, extracts tactile features and image features respectively, and fuses them to obtain fused features, including tactile feature extraction. Image feature extraction and modal fusion After obtaining the image features and tactile features, they are stitched together in the channel dimension and then input into the modality fusion module for fusion. The relevant parameters of the source encoder are updated through a classification task of tactile and image signal sets; The source encoder parameters are jointly updated using the cross-entropy loss function applied to image classification, tactile classification, and fused feature classification, as shown below. in, This indicates that the tactile feature extraction network was used for each step. Image feature extraction network and modal fusion The output image features, tactile features, and fusion features, y j Indicates the category label of the current tactile-image signal. These represent tactile feature extraction networks. Image feature extraction network Modal fusion network The parameters, Indicates source encoder integrated network The parameters, By minimizing the loss function Update the source encoder parameters.
3. The tactile signal reconstruction method based on joint source-channel coding as described in claim 2, characterized in that: The transmission function of the AWGN channel is expressed as follows: y = z + n, Where n is the channel noise, which is independent and identically distributed, and σ 2 This is the average noise power, and the channel's signal-to-noise ratio (SNR) is SNR = 10log 10 (1 / σ 2 ), by changing σ 2 To change the signal-to-noise ratio of the channel; The transmission function of the Rayleigh slow fading channel is expressed as follows: y = hz + n, Where h is the channel gain. n is the channel noise. h and n respectively obey the parameters as follows: and Different normal distributions; After transmission through the channel, the noise-affected complex-valued signal is decoded by the channel decoder. Mapped to estimates of fused features The channel decoder model is represented as follows: in, These represent the parameters of the channel decoder model; The channel encoder and channel decoder are jointly designed and jointly optimized by minimizing the fused feature f and the estimated fused feature after transmission. Mean square error The parameters of the channel codec are updated jointly, denoted as follows:
4. The tactile signal reconstruction method based on joint source-channel coding as described in claim 3, characterized in that: The reconstruction of the target tactile signal includes the tactile decoder reconstructing the tactile signal from the fused semantic features, represented as follows: Where, θ gh Indicates the parameters of the haptic decoder network; The haptic decoder network employs multiple fully connected layers and dropout layers of varying sizes to extract haptic information from fused features, thereby recovering the haptic information. The loss function is set to reconstruct the haptic feedback. The mean absolute error and mean squared error between the actual tactile sensation H and the real tactile sensation H are expressed as follows: Starting from the frequency domain, a short-time Fourier transform (STFT) loss function is set to reflect the frequency domain characteristics of the signal; The loss function includes spectral convergence loss. And logarithmic magnitude loss The spectral convergence loss is expressed as: The logarithmic magnitude loss is expressed as follows: Among them, ||·|| F The F-norm and L1-norm operations are represented by ||·||1, STFT(·) represents the STFT transform, and Num represents the number of elements in the STFT amplitude. The STFT loss function is expressed as follows: The loss function for generating tactile signals is expressed as follows: Here, α, β, and γ are hyperparameters.
5. The tactile signal reconstruction method based on joint source-channel coding as described in claim 4, characterized in that: The optimal tactile reconstruction model parameters are obtained by minimizing the loss function. First, pre-train the source encoder network. Parameters; Fixed source encoder network parameters By minimizing the loss Channel encoder and channel decoder The parameters are updated using gradients until the model converges; The fused semantic features are obtained through the source encoder and input into the tactile decoder network G. h (f;θ gh ), by minimizing Gradient updates are performed on the parameters of the haptic decoder network until the model converges.
6. The tactile signal reconstruction method based on joint source-channel coding as described in claim 5, characterized in that: The pre-trained source encoder network The parameters include the initialization parameters of the source encoder network. Input training signal set Set the number of iterations to n1 and the learning rate to μ1; according to Update source encoder network parameters Represented as, in, To perform partial derivatives with respect to each loss function; training continues until the number of iterations n1 is reached to complete the training. If n < n1, then n = n + 1, and continue to the next iteration; otherwise, terminate the iteration; after n1 iterations, the optimized source encoder network parameters are obtained. The channel encoder and channel decoder Gradient updates of parameters include initializing the channel encoder network parameters. and channel decoder network parameters The fused features output from the source encoder are used as input, the number of iterations is set to n2, and the learning rate is μ2; Fixed source encoder network parameters optimization and Represented as, If n < n2, then n = n + 1, and continue to the next iteration; otherwise, terminate the iteration; after n2 iterations, the optimized channel encoder network parameters are obtained. and channel decoder network parameters The minimize Gradient updates to the haptic decoder network parameters include initializing the haptic decoder network parameters θ. gh ; The fused features output from the source encoder are used as input, the number of iterations is set to n3, and the learning rate is μ3; Fixed source encoder network parameters Optimize θ gh , represented as If n < n3, then n = n + 1, and continue to the next iteration; otherwise, terminate the iteration; after n3 iterations, the optimized haptic decoder network parameters θ are obtained. gh .
7. The tactile signal reconstruction method based on joint source-channel coding as described in claim 6, characterized in that: The generation of the target tactile signal includes the source encoder extracting key features from the input data and generating fused features. The fused features are further processed by the channel encoder. The encoded data is then transmitted through the channel and decoded by the channel decoder to recover the original fused features. The tactile decoder network uses the recovered fused features to generate the desired tactile signal, thus completing the reconstruction of the entire tactile signal.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the tactile signal reconstruction method based on source-channel joint coding as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the tactile signal reconstruction method based on source-channel joint coding as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Source-channel joint tactile coding method, apparatus and device, and medium
CN115694737A
Visual-tactile signal adaptive reconstruction method based on transfer learning
CN115993888A