A method for haptic signal reconstruction of a signal-to-noise ratio adaptive joint optimization of a signal source and a channel

By combining a cross-modal joint codec and a signal-to-noise ratio adaptive module, the encoding, transmission, and decoding processes of tactile signals are optimized, solving the problem of insufficient signal-to-noise ratio generalization in existing models and achieving stable reconstruction of tactile signals under different channel conditions.

CN119094082BActive Publication Date: 2025-11-07NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411075759.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2025-11-07
Estimated Expiration
2044-08-07

AI Technical Summary

Technical Problem

Existing cross-modal signal reconstruction models have low generalization ability for signal-to-noise ratio and cannot effectively cope with different channel conditions, making tactile signals susceptible to noise interference in actual communication, affecting data quality and stability.

Method used

A cross-modal joint codec design is adopted, combined with a signal-to-noise ratio adaptive module. Through hierarchical feature extraction and adaptive signal-to-noise ratio adjustment, the signal encoding, transmission and decoding processes are optimized to achieve stable reconstruction of tactile signals.

Benefits of technology

This improves the model's robustness to channel variations, ensuring that the tactile signal remains stable and reliable under different signal-to-noise ratio conditions, and significantly enhances the stability and accuracy of signal reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119094082B_ABST
    Figure CN119094082B_ABST
Patent Text Reader

Abstract

The application discloses a kind of signal-to-noise ratio adaptive joint optimization haptic signal reconstruction method of source channel, it is related to haptic signal generation technical field, comprising: the image corresponding to each kind of object and the haptic vibration signal generated by sliding are preprocessed;Cross-modal joint encoder is constructed to generate hierarchical fusion features, cross-modal joint decoder is constructed, and the training data obtained by preprocessing is input into the cross-modal joint encoder and cross-modal joint decoder constructed and is trained;The paired haptic signal to be measured and image signal are input into cross-modal joint encoder and optimal cross-modal joint decoder, and the target haptic signal is reconstructed.The haptic signal reconstruction method of signal-to-noise ratio adaptive joint optimization source channel provided in the application simplifies the encoding training process, better modal feature fusion is carried out;The robustness of model to channel change is improved, so that the model shows higher haptic signal reconstruction stability under a larger range of signal-to-noise ratio change.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of tactile signal generation, in particular to a tactile signal reconstruction method for adaptive joint optimization of signal-to-noise ratio and source channel. BACKGROUND

[0002] With the progress of technology and the continuous evolution of user needs, traditional audio-visual media has been insufficient to meet the growing expectations of users for immersive and interactive experiences, which has driven the evolution of multimedia services to more advanced forms. However, existing research has focused on the analysis and transmission of single modal signals, failing to fully address the complex transmission needs of heterogeneous multi-modal signals. Therefore, cross-modal services have emerged, which create a new mode of communication and experience for users by integrating visual, auditory, and tactile sensory information. In particular, cross-modal services centered on touch are widely considered a key development direction for future communication technologies, as they can provide direct and realistic interactive feedback. However, current research on cross-modal services centered on touch still faces several challenges.

[0003] On the one hand, current cross-modal signal reconstruction schemes have not fully considered the impact of channel noise, and need to further consider the robustness of signal reconstruction under different channel conditions. In actual communication scenarios, tactile signals need to be encoded at the source end, transmitted in the channel, and decoded and reconstructed at the receiving end. This process involves the interaction of multiple factors, including semantic fusion of cross-modal signals, transmission characteristics of the channel, and reconstruction strategies at the receiving end, etc.

[0004] On the other hand, existing signal communication reconstruction models have low generalization to signal-to-noise ratio in actual applications, and need to be trained and deployed according to different channel conditions, occupying more computing and storage resources. Existing models based on deep neural networks rely on data-driven strategies. The optimization process of the entire system is based on an end-to-end loss function that restores accuracy, and the parameters of the model are adjusted through supervised learning. This model training mechanism easily leads to low generalization of the model to signal-to-noise ratio, resulting in good performance of the model at certain specific signal-to-noise ratios during training, but rapid performance decline at other signal-to-noise ratios.

[0005] Existing research has also made some progress in the cross-modal fusion and generation of haptic signals. In the literature "Information recovery technology for cross-modal communication", Xu et al. proposed an information recovery scheme for the problem of loss or noise pollution of wireless channels that multi-modal data (such as audio, video, and haptic signals) may encounter during transmission. The scheme uses existing multi-modal data at the receiving end to achieve information recovery through same-modal one-to-one search, cross-modal one-to-one search, cross-modal one-to-many search, and other methods. Wei et al. proposed an audio-visual assisted haptic signal reconstruction method based on cloud-edge collaboration in the literature "Haptic Signal Reconstruction for Cross-Modal Communications". For the case of insufficient data received by the edge node, a large audio-visual assisted database stored in the central cloud is used to obtain and transmit knowledge and migrate to the edge node. Then, the edge node is used to mine the inherent semantic consistency and correlation between modalities, and finally the required haptic signal reconstruction is realized. By considering the received audio-visual signals and relating them to the potential correlation of the haptic modality, the reconstruction of the haptic signal is realized.

[0006] However, these models are mostly limited to one-way conversion at the sending end or receiving end, without considering end-to-end design of the entire communication process, and lack simulation of transmission in actual channels. In actual situations, it is easy to be disturbed by noise, resulting in a decline in the quality of the transmitted data. This will have an adverse effect on multi-modal data streams, especially haptic streams, which require low latency and high reliability. In addition, for a variety of channel conditions and varying channel signal-to-noise ratios, how to design an effective channel adaptive technology to significantly improve the stability of haptic signal reconstruction is also a difficult problem faced by haptic-based cross-modal communication. SUMMARY

[0007] In view of the above problems, the present application is proposed.

[0008] Therefore, the present application realizes the encoding, transmission and decoding process of the signal based on the deep neural network. On the one hand, through the unified model design of the cross-modal joint encoder and decoder, the encoding training process is simplified,

[0009] At the same time, the unified model also brings a unified haptic-image signal feature representation, which better performs modal feature fusion. On the other hand, through the separately designed signal-to-noise ratio adaptive module, the robustness of the model to channel changes is improved, so that the model shows higher stability in haptic signal reconstruction under a larger range of signal-to-noise ratio changes.

[0010] To solve the above technical problems, the present application provides the following technical solutions: a signal-to-noise ratio adaptive joint optimization haptic signal reconstruction method of a signal source channel, comprising:

[0011] Preprocess the image signal corresponding to each object and the tactile vibration signal generated by sliding based on the tactile-image database;

[0012] Input the preprocessed image and tactile signal into the cross-modal joint encoder, block process the input data, input the hierarchical feature extraction module and the signal-to-noise ratio adaptive module in turn, and obtain the hierarchical fusion feature;

[0013] The hierarchical fusion feature is transmitted through a channel and enters the cross-modal joint decoder, including the hierarchical feature extraction module, the signal-to-noise ratio adaptive module and the signal reconstruction module;

[0014] The hierarchical feature extraction module decodes the tactile low-dimensional feature from the hierarchical fusion feature, the signal-to-noise ratio adaptive module dynamically adjusts the low-dimensional feature according to the channel change, and the signal reconstruction module converts the low-dimensional feature into the tactile frequency spectrum. The tactile frequency spectrum is converted into a tactile sequence through inverse short-time Fourier transform ISTFT, so as to complete the tactile signal reconstruction;

[0015] According to the training data, the cross-modal joint encoder and the cross-modal joint decoder are jointly trained to obtain the optimal model parameters;

[0016] The tactile and image signals to be tested are input into the optimal cross-modal joint encoder to obtain the hierarchical fusion feature, and after channel transmission, the received hierarchical fusion feature is input into the cross-modal joint decoder to generate the target tactile frequency spectrum, and the target tactile signal is obtained through ISTFT transformation.

[0017] As a preferred scheme of the tactile signal reconstruction method of the signal-to-noise ratio adaptive joint optimization signal source channel according to the application, wherein: the preprocessing includes cutting the image according to the HxW size, which is represented as,

[0018]

[0019] Wherein, v j represents the jth picture in the data set with the size of HxW RGB three-channel picture;

[0020] The tactile signal is preprocessed, the short-time Fourier transform STFT of the tactile sequence is performed to obtain the frequency spectrum of the tactile sequence, which is represented as,

[0021]

[0022] The frequency spectrum is passed through and

[0023] The amplitude spectrum and the phase spectrum

[0024] Concatenate the amplitude spectrum and the phase spectrum to obtain the haptic signal

[0025] Concatenate the RGB image and the haptic signal In the channel dimension to five-dimensional data Where n = H x W x 5 represents the size of the image-haptic signal pair, and the one-to-one pairing of the image and the haptic signal is maintained during processing.

[0026] As a preferred scheme of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio of the source channel according to the application, wherein: the block processing includes dividing the data pair into non-overlapping blocks of 4x4 size;

[0027] The feature dimension of each patch is 4x4x5 = 80, and the size of the data tensor is

[0028] The block-processed data is input into a plurality of hierarchical feature extraction modules and a signal-to-noise ratio adaptive module;

[0029] The hierarchical feature extraction module is FE e (·), and the signal-to-noise ratio adaptive module is SD e (·), which are connected after the first and last feature extraction components, respectively;

[0030] Let the number of feature extraction components in the encoder be L, and the input of the jth feature extraction component be t j ,j∈{1,2,...,L}, and the continuous calculation process is represented as,

[0031]

[0032] The haptic-image signal pair d∈D is divided into non-overlapping small blocks in the length and width directions to obtain t1, which is input into the hierarchical feature extraction module to obtain the first layer feature representation The first layer feature representation is input into the signal-to-noise ratio adaptive module to obtain t2, which is input into a plurality of feature extraction components, and the deep layer feature is obtained after the last layer feature extraction component, which is input into the last layer signal-to-noise ratio adaptive layer to obtain the remapped feature e2.

[0033] As a preferred scheme of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio of the source channel according to the application, wherein: the hierarchical feature extraction module FE e (·) includes a merging layer and a Swin Transformer module, and learns the hierarchical features of the haptic-image signal;

[0034] The merging layer is responsible for down-sampling and up-dimensioning to generate features at different levels, while the SwinTransformer block performs hierarchical feature learning;

[0035] The first merging layer connects the features of each group of 2x2 adjacent patches in the channel dimension, reduces the feature resolution by 2 times, and increases the feature dimension by 4 times;

[0036] The full connection layer is applied on the connected features to unify the output dimension to twice the original dimension, i.e., 2C, and the tensor size is

[0037] After passing through the Swin Transformer block and the merging layer multiple times, hierarchical feature representation is obtained;

[0038] The last merging layer of the cross-modal joint encoder is responsible for mapping deep feature information to the same dimension as the channel dimension, and after connecting the features of adjacent patches, a linear layer is applied to transform the output dimension to the channel dimension b, at this time the tensor size is

[0039] As a preferred scheme of the signal-to-noise ratio adaptive joint optimization method of the signal source channel of the haptic signal reconstruction method, wherein: the signal-to-noise ratio adaptive module SD(·) includes global pooling and a feedforward network;

[0040] The global pooling is used to capture the full-size information of the input features, and the local features are summarized as global features, the signal-to-noise ratio information is connected with the pooled features, and the scaling factor is generated by inputting the feedforward network, the scaling factor is multiplied with the output features of the hierarchical feature extraction module FE e (·), and the weighted features are obtained;

[0041] The output of the feature extraction component is o={o1,o2,...,o c},o ipq represents the element at (p,q) in the i-th channel;

[0042] After global average pooling P(·), each feature map is pooled in the channel dimension to extract global information

[0043] The global pooling operation compresses each two-dimensional feature map o i into a single value s i , which represents the average activation strength of the entire feature map, and is represented as,

[0044]

[0045] The global information is spliced with the current SNR value to obtain It is input into the feedforward network, denoted as,

[0046]

[0047] The The original input feature before pooling is multiplied by to obtain the feature adjusted according to the channel condition Since the sizes are different, before multiplication, the data of each channel is copied by factor, so that the size is expanded to HxWxc.

[0048] As a preferred scheme of the haptic signal reconstruction method of adaptive optimization of signal-to-noise ratio and channel of the application, wherein: the cross-modal joint decoder uses an expansion layer to upsample the extracted deep features, the expansion layer reshapes the feature maps of adjacent dimensions and reduces the feature dimension to half of the original dimension, completing the dimension reduction of the features;

[0049] The channel signal is input into a plurality of hierarchical feature extraction modules to obtain features of different levels, and the signal-to-noise ratio adaptive module is connected after the first and last hierarchical feature extraction modules;

[0050] Let the number of hierarchical feature extraction modules in the cross-modal joint decoder be L, and the input of the jth hierarchical feature extraction module according to the data flow direction be t j ,j∈{1,2,...,L}, the continuous calculation process is represented as,

[0051]

[0052]

[0053] The signal transmitted from the channel is input into the hierarchical feature extraction module to obtain the first layer of reduced dimension feature representation t2 is input into the signal-to-noise ratio adaptive module, and after the dimension reduction by the plurality of hierarchical feature extraction modules, the last reduced shallow layer feature is scaled again by a layer of signal-to-noise ratio adaptive layer to obtain the remapped feature e2, which is scaled by a layer of expansion layer by 4 times to convert the feature size to

[0054] The reduced dimension feature is input into the signal reconstruction module, which is mapped from the feature domain to the time-frequency matrix containing the amplitude and phase information of the haptic sequence through a plurality of convolution layers, denoted as,

[0055]

[0056] The multi-level feature input multiple convolutions, keeps the size unchanged, realizes the expansion and contraction of the channel number, and the channel number is expanded and contracted to 2, is mapped to the time-frequency matrix of the amplitude and phase information of the haptic sequence, and the recovered haptic sequence is obtained through ISTFT transformation.

[0057] As a preferred scheme of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio and channel of the source, wherein: the joint training includes selecting Charbonnier loss as the loss function of the optimization process, expressed as,

[0058]

[0059] Wherein, h j represents the haptic signal recovered by the decoder, h j represents the real haptic signal, and epsilon = 10 -3 is a constant.

[0060] The cross-modal joint encoder and the cross-modal joint decoder are jointly trained as a whole network, the gradient is calculated through the Adam optimization algorithm, and the network parameters are updated through back propagation. Its expression is as follows:

[0061]

[0062] Wherein, m t and v t are the unbiased estimates of the first moment and the second moment respectively. g t is the gradient of the parameter, beta1 and beta2 are the decay rates of the first moment estimate and the second moment estimate respectively. mu is the learning rate, and epsilon is a very small constant to ensure that the denominator is not zero. Theta t and theta t+1 are the network parameters before and after updating respectively. The network parameters here refer to the network parameters composed of the cross-modal joint encoder and the cross-modal joint decoder. The learning rate is selected as 1x10-4, the iteration number is 1000 times, and the network parameters are determined after iteration, and the model training is completed.

[0063] As a preferred scheme of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio and channel of the source, wherein: the reconstructed target haptic signal includes a hierarchical fusion feature generated through the cross-modal joint encoder, and the cross-modal joint decoder decodes and reconstructs the spectrum of the target haptic signal after channel transmission, and obtains the expected haptic signal through ISTFT transformation.

[0064] A computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio and channel of the source when executing the computer program.

[0065] A computer readable storage medium, having stored thereon a computer program, the computer program being executed by a processor to implement the steps of the haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel as described above.

[0066] The present application has the following beneficial effects: The haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel provided by the present application fully utilizes the advantages of multi-modal feature fusion, and effectively fuses different modal semantic information. Not only can it significantly compress the signal source, reduce the data volume, but also can maintain the integrity and accuracy of the information. Further, in order to cope with the influence of channel noise and changes on signal transmission, the present application specially designs a set of channel encoder and decoder. The encoder and decoder can improve the robustness of the signal to channel changes, and ensure that the signal remains stable and reliable under various channel conditions. BRIEF DESCRIPTION OF DRAWINGS

[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0068] Figure 1 The overall flowchart of a haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel provided by an embodiment of the present application.

[0069] Figure 2 The network structure diagram of a haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel provided by an embodiment of the present application.

[0070] Figure 3 The Swin Transformer module diagram of a haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel provided by an embodiment of the present application.

[0071] Figure 4 The signal-to-noise ratio adaptive module diagram of a haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel provided by an embodiment of the present application.

[0072] Figure 5 The haptic signal reconstruction result diagram of a haptic signal reconstruction method of adaptive signal-to-noise ratio joint optimization of a signal source channel provided by an embodiment of the present application and other comparative methods.

[0073] Figure 6A haptic spectrum and sequence reconstruction result graph of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a haptic signal source channel is provided for an embodiment of the present application.

[0074] Figure 7 A haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a haptic signal source channel and other method haptic reconstruction stability comparison graphs are provided for an embodiment of the present application. DETAILED DESCRIPTION

[0075] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0076] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details that are set forth in the description, and it is understood that persons having ordinary skill in the art can make and use other implementations of the present application without departing from the scope of the present application. Accordingly, the present application is not limited to the embodiments described herein.

[0077] Embodiment 1

[0078] Reference Figures 1-4 For an embodiment of the present application, a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a haptic signal source channel is provided, comprising:

[0079] Step 1: On the basis of the disclosed haptic-image database, pre-process the image signal corresponding to each object and the haptic vibration signal generated by sliding. Convert the haptic vibration signal generated by sliding into amplitude spectrum and phase spectrum by short-time Fourier transform (STFT), and cut the size of the image signal to align the size of the haptic spectrum;

[0080] Step 1-1, first, pre-process the image signal. Cut the image according to the size of HxW, denoted as: wherein, v j represents the jth picture in the data set with a size of HxW RGB three-channel picture.

[0081] Step 1-2, pre-process the haptic signal, and perform short-time Fourier transform (STFT) on the haptic sequence to obtain the frequency spectrum of the haptic sequence When performing STFT transformation, a suitable window function and window size should be selected, and the "Constant OverLap Add" (COLA) constraint should be met, so as to ensure that there is no information loss when performing STFT transformation, so that the haptic sequence can be restored through ISTFT transformation when restoring the haptic. The frequency spectrum is calculated by the following formula to obtain the amplitude spectrum and the phase spectrum Then the haptic signal is spliced

[0082]

[0083] Finally, the RGB image and the haptic signal are spliced into five-dimensional data in the channel dimension , wherein n = H x W x 5 represents the size of the image-haptic signal pair. During processing, the image and the haptic signal are kept in one-to-one pairing;

[0084] Step 2: input the preprocessed image and haptic signal into the cross-modal joint encoder. First, the input data is divided into blocks, which can reduce the length of the subsequent embedding sequence and more effectively capture long-distance dependencies. Then, the hierarchical feature extraction module and the signal-to-noise ratio adaptive module are input in turn, and finally the hierarchical fusion features are obtained, as shown in Figure 2 . Among them, the hierarchical feature extraction module extracts and fuses the multi-level features of the image-haptic signal, and the signal-to-noise ratio adaptive module dynamically adjusts the multi-level fusion features generated by the hierarchical feature extraction module according to the real-time channel conditions to adapt to the changing channel environment.

[0085] Step 2-1, input the preprocessed image and haptic signal into the cross-modal joint encoder. The input data is divided into blocks, which can reduce the length of the subsequent embedding sequence and more effectively capture long-distance dependencies. In addition, cutting the image into small blocks can preserve local structural information. Specifically, the data pair is divided into 4x4 non-overlapping blocks (patches), and through this division method, the feature dimension of each patch is 4x4x5 = 80. At this time, the size of the data tensor is

[0086] Step 2-2, then input the blocked data into multiple hierarchical feature extraction modules and signal-to-noise ratio adaptive modules. The hierarchical feature extraction module is denoted as FE e (·). Among them, the signal-to-noise ratio adaptive module is denoted as SD e (·), which is connected after the first and last feature extraction components respectively, so as to learn the feature representation under different signal-to-noise ratios of shallow and deep levels, as shown in Figure 2As shown. Specifically, assume that the number of feature extraction components in the encoder is L, and the input of the j-th feature extraction component is t. j ,j∈{1,2,...,L}. The continuous calculation process is as follows:

[0087]

[0088] First, the tactile-image signal pair The feature is divided into non-overlapping blocks in both length and width directions to obtain t1, which is then input into the hierarchical feature extraction module to obtain the first layer of feature representation. Next, the first layer of feature representation is input into the signal-to-noise ratio adaptive module to obtain t2, and then t2 is input into multiple feature extraction components. After passing through the last feature extraction component, deep features are obtained, and then input into the last signal-to-noise ratio adaptive layer to obtain the remapped features e2.

[0089] Steps 2-3, including the hierarchical feature extraction module FE e The (·) module includes a merging layer and a SwinTransformer module to learn hierarchical features of haptic-image signals. The merging layer is primarily responsible for downsampling and dimensionality upsampling to generate features at different levels, while the SwinTransformer block mainly performs hierarchical feature learning. The first merging layer concatenates the features of each 2×2 adjacent patch along the channel dimension. This process downsamples the feature resolution by a factor of 2 and increases the feature dimension by a factor of 4 due to the concatenation operation. Therefore, a fully connected layer is applied to the concatenated features to unify the output dimension to twice the original dimension, i.e., 2C. At this point, the tensor size is... Then, after multiple passes through the SwinTransformer block and merging layer, a hierarchical feature representation is obtained. The final merging layer of the cross-modal co-encoder is responsible for mapping the deep feature information to a dimension with the same channel size. Therefore, unlike the previous merging layers, after concatenating adjacent patch features, a linear layer is applied to transform the output dimension to the channel dimension b. At this point, the tensor size is...

[0090] The Swin Transformer module is crucial for feature extraction; it will be discussed in detail below. Figure 3 As shown, two consecutive SwinTransformer modules are used, each consisting of a normalized (LayerNorm, LN) layer, a multi-head self-attention module, residual connections, and a multilayer perceptron (MLP). The two consecutive SwinTransformer modules employ a window-based multi-head self-attention module (W-MSA) and a displacement window-based multi-head self-attention module (SW-MSA), respectively.

[0091] In the standard ViT (Vision Transformer), global attention is used, which calculates the relationship between the target token and all other tokens, increasing the computational complexity. The Swin Transformer module is based on a shift window. The tensor input into the Swin Transformer is evenly divided into several non-overlapping windows, and the self-attention is calculated only within the local window. Since the number of patches within the window is much smaller than the number of patches in the picture, the amount of calculation is greatly reduced. On the other hand, WMSA only calculates self-attention within the window, and there is a lack of information exchange between non-overlapping windows. In order to improve the model's ability to build global relationships, SW-MSA needs to be constructed to realize cross-window information exchange and increase the receptive field by introducing a shift window.

[0092] The method of shift window will cause a problem. Since the window size and the tensor size are not aligned, after shifting, the number of windows will increase, and the size of the elements in each window will not be uniform. The literature "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" proposes to use a cyclic shift with a mask to keep the number of windows unchanged after shifting and to make the elements uniform.

[0093] Specifically, the method of shift window is realized between two consecutive SwinTransformer modules. The first module performs WMSA operation to calculate the attention within the window, and the second module performs SW-MSA operation to calculate the attention by moving the window. Assuming that the input of the i-th Swin Transformer module is u i , the specific calculation process is as follows:

[0094]

[0095] Among them, LN represents the normalization layer, WMSA represents the window multi-head self-attention module, MLP represents the multi-layer perceptron, and SW-MSA represents the shift window multi-head self-attention module. represents the output of the i-th Swin Transformer module. The calculation process of attention is as follows:

[0096]

[0097] The input token sequence will generate three matrices of query, key and value, denoted as Q, K, V. d represents the dimension of query or key. B represents the relative position bias.

[0098] Step 2-4, the SNR adaptive module SD(·) is the key to ensure the model to adapt to different channel levels, which includes global pooling, feedforward network, as shown in Figure 4 Global pooling is used to capture the full-size information of input features, and local features are summarized as global features. The SNR information is connected with the pooled features, and the feedforward network (FFN) is input to generate scaling factors. These scaling factors are multiplied with the output features of the hierarchical feature extraction module FE e (·), and the weighted features are then used for subsequent encoding and transmission. The purpose of multiplication is to adjust the importance of features according to different SNR conditions, so that the model can more effectively handle channel noise and transmit data.

[0099] First, the output of the hierarchical feature extraction component is o = {o1, o2,..., o c}, o ipq represents the element of the i-th channel at (p, q). After global average pooling P(·), each feature map is pooled in the channel dimension, which can extract global information The global pooling operation compresses each two-dimensional feature map o i in the channel into a single value s i , which represents the average activation strength of the entire feature map.

[0100]

[0101] Then, these global information and the current SNR value are spliced to get which is input into a feedforward network FFN(·). The output of the network is a scaling factor, which determines the weight of each element in the subsequent feature map. The feedforward network consists of multiple linear layers, and the activation function of the last layer is Sigmoid, while the activation function of the remaining layers is Relu. This ensures that the output scaling factor is between 0 and 1.

[0102]

[0103] Finally, multiply with the original input feature before pooling to get the feature adjusted according to the channel condition Since the sizes of the two are different, before multiplication, the data of each channel of factor is copied, so that the size is expanded to HxWxc.

[0104] By the scaling factor predicted according to the feedforward network, the SNR adaptive module can weight and adjust each element in the original feature map. The SNR adaptive module does not change the size of the input feature. Specifically, each element in the feature map is multiplied by the corresponding scaling factor, thereby realizing the recalibration of the feature. If the scaling factor is close to 0, it means that the corresponding feature is not important under the current SNR condition and should be suppressed; if the scaling factor is close to 1, the feature remains unchanged or is enhanced. Unlike the traditional binary mask, the scaling factor is equivalent to a kind of soft mask, and the scaling factor is a value between 0 and 1, which is used to adjust the weight of the feature in proportion, rather than simply switching on and off. This method provides a more delicate and continuous control method, allowing the model to adjust the importance of the feature in a more fine-grained manner according to the current channel condition and image content.

[0105] Step 3: After the hierarchical fusion feature is transmitted through the channel, it enters the cross-modal joint decoder. The cross-modal joint decoder adopts a symmetrical design similar to the structure of the encoder, including a hierarchical feature extraction module, an SNR adaptive module, and a signal reconstruction module, as shown in Figure 2 . Among them, the hierarchical feature extraction module decodes the haptic low-dimensional feature from the hierarchical fusion feature, the SNR adaptive module dynamically adjusts this low-dimensional feature according to the channel change, and finally the signal reconstruction module converts the low-dimensional feature into a haptic spectrum. The haptic spectrum is converted into a haptic sequence through inverse short-time Fourier transform (ISTFT), thereby completing the reconstruction of the haptic signal;

[0106] Step 3-1, the hierarchical fusion feature is input into the cross-modal joint decoder after being transmitted through the channel. Its structure corresponds to the cross-modal joint encoder, which is composed of multiple hierarchical feature extraction modules , an SNR adaptive module SD1 d , and a signal reconstruction module. At this time, the hierarchical feature extraction module is composed of an expansion layer and a SwinTransformer module. Unlike the merging layer used in the encoder, the expansion layer is used in the decoder to upsample the extracted deep features. The expansion layer reshapes the feature maps of adjacent dimensions into larger size feature maps (2x upsampling), and correspondingly reduces the feature dimension to half of the original dimension, completing the dimension reduction of the feature.

[0107] Step 3-2, the channel signal is input into multiple hierarchical feature extraction modules to obtain features at different levels. Similarly, the SNR adaptive module is connected after the first and last hierarchical feature extraction modules.

[0108] Specifically, assuming that the number of hierarchical feature extraction modules in the decoder is L, the input of the jth hierarchical feature extraction module is t jj e {1,2,...,L}. The continuous calculation process is as follows:

[0109]

[0110] Signal transmitted from the channel The input hierarchical feature extraction module obtains a first layer of reduced dimension feature representation Next, the input signal-to-noise ratio adaptive module obtains t2. After that, t2 is reduced in dimension through multiple hierarchical feature extraction modules, and then again through a signal-to-noise ratio adaptive layer, the shallow layer feature after the last dimension reduction is scaled, and the feature e2 that is remapped is obtained. Secondly, through an expansion layer, the feature size is converted to

[0111] Step 3-3, input the reduced dimension feature into the signal reconstruction module, and map from the feature domain to the time-frequency matrix containing the amplitude and phase information of the haptic sequence through multiple convolution layers At this stage, multiple levels of features are input into multiple convolutions, the size is kept unchanged, the number of channels is scaled, and finally the number of channels is scaled to 2, which is accurately mapped to the time-frequency matrix of the amplitude and phase information of the haptic sequence. Finally, the recovered haptic sequence is obtained through ISTFT transformation.

[0112] Step 4, jointly train the cross-modal joint encoder and the cross-modal joint decoder according to the training data, and obtain the optimal model parameters;

[0113] Step 4-1, in the training stage, the model receives the haptic image signal pair and the channel signal-to-noise ratio information for training, and learns the mapping relationship under different signal-to-noise ratio levels. The Charbonnier loss is selected as the loss function of the optimization process. The calculation formula is as shown in the formula:

[0114]

[0115] where h j represents the haptic signal recovered by the decoder, h j represents the real haptic signal, and ε = 10 -3 is a constant. The existence of the constant ε ensures that the gradient is not too small when it approaches zero, and the square root helps to avoid gradient disappearance. The Charbonnier loss function has strong robustness to outliers. Compared with the mean square error (L2 loss), it is not sensitive to large residuals, which makes the model more stable during training and less susceptible to noise, making it easier to converge quickly, and the loss calculation is relatively simple.

[0116] Step 4-2, the cross-modal joint encoder and the cross-modal joint decoder are jointly trained as a whole network, the gradient is calculated through the Adam optimization algorithm, and the network parameters are updated through back propagation. The expression is as follows:

[0117]

[0118] Where, m t and v t are the unbiased estimates of the first and second moments respectively. g t is the gradient of the parameter, beta1 and beta2 are the decay rates of the first and second moment estimates respectively. μ is the learning rate, and ε is a very small constant to ensure that the denominator is not zero. θ t and θ t+1 are the network parameters before and after updating respectively. The network parameters here refer to the network parameters composed of the cross-modal joint encoder and the cross-modal joint decoder. The learning rate is selected as 1x10-4, the iteration number is 1000 times, and the network parameters are determined after the iteration is completed, and the model training is completed.

[0119] Step 5: input the to-be-tested haptic and image signals into the optimal cross-modal joint encoder obtained in the above step to obtain hierarchical fusion features, then transmit the hierarchical fusion features through a channel, input the received hierarchical fusion features into the cross-modal joint decoder to generate a target haptic spectrum, and obtain a target haptic signal through ISTFT transformation.

[0120] Step 5-1, select the optimal cross-modal source encoder, channel codec and haptic decoder network model trained;

[0121] Step 5-2, input the to-be-tested haptic-image signal pair into the model, first, generate a hierarchical fusion feature through the cross-modal joint encoder. Then, after channel transmission, the cross-modal joint decoder is used for decoding, and the spectrum of the target haptic signal is reconstructed. Finally, the expected haptic signal is obtained through ISTFT transformation.

[0122] Embodiment 2, an embodiment of the present application, which is different from the above embodiment is:

[0123] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the technical solutions that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0124] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor or other system that can fetch the instructions from the instruction execution system, apparatus or device and execute the instructions, or in conjunction with these instructions execution systems, apparatus or devices. For the purpose of this specification, the "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by or in connection with an instruction execution system, apparatus or device, or in conjunction with these instruction execution systems, apparatus or devices.

[0125] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electrical devices), a portable computer diskette (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be electronically obtained, for example, by optical scanning of the paper or other medium, followed by editing, interpreting or otherwise processing, if necessary, in other suitable ways, to be electronically obtained and then stored in the computer memory.

[0126] It should be understood that portions of the application can be implemented in hardware, software, firmware, or combinations thereof. In the embodiments described above, the various steps or methods can be implemented, in part, or in whole, by software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following techniques, which are well known in the art, can be used to implement the application: a hybrid of the techniques mentioned above; a combination of one or more of the techniques mentioned above; a combination of one or more of the techniques mentioned above with one or more other techniques not mentioned above; and / or any other techniques for implementing the application.

[0127] Example 3, with reference to Figures 5-7 For an embodiment of the application, a haptic signal reconstruction method of signal-to-noise ratio adaptive joint optimization of source and channel is provided. In order to verify the beneficial effects of the application, a simulation experiment is carried out for scientific demonstration.

[0128] The present embodiment adopts the AUdataset data set for experiment, which is collected by Lasse et al. of the Department of Electrical and Computer Engineering, Aarhus University. The data set contains multimodal signals of 63 objects: vision, motion and haptics (audio vibration signals). The objects selected in the data set are mainly toys and household items. Each object contains 3 pictures of different angles and 5 recorded haptic vibration signals. Among them, there are similar shape and color objects (visual ambiguity) and similar material and texture but different color objects (haptic ambiguity). These objects containing visual ambiguity and haptic ambiguity will reflect the different effects of vision and haptics on perception. The present example selects 62 available objects, and augments 3 images of each object to 5 (for example, randomly adjusts the brightness and contrast of the picture, and flips the image), and 5 recorded haptic vibration signals to form 310 groups of samples. In the data set, 80% is selected for training, and the remaining 20% is used for testing and performance evaluation.

[0129] Firstly, the method of the application is compared with the pure haptic reconstruction method, that is, the method of the application needs to consider channel interference, while the pure haptic reconstruction method only considers end-side signal reconstruction. The following three haptic reconstruction methods are selected for experimental comparison:

[0130] Existing method one:

[0131] Document: "Cross-Modal Data Generation Using Residue-Fusion GAN With Feature-Matching and Perceptual Losses"

[0132] In the paper "Tactile Image-to-Image Translation on Haptic Data Using Conditional Generative Adversarial Networks" (Authors: Cai S, Zhu K, Ban Y, et al), the cGAN (conditional-GAN) structure and residual-fusion (RF) module are used, and the model is trained by using additional feature matching (FM) and perception loss, to realize the mutual generation between visual pictures and haptic spectrum.

[0133] Existing method two:

[0134] In the paper "Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired People" (Authors: Liu H, Guo D, Zhang X, et al), DiscoGAN is used to learn the conversion between images and spectrum graphs. By using the trained DX and DY as discriminators to double-constrain the generated results, and by adding an auxiliary classification layer to the original discriminator, the model can learn the mutual conversion between the visual domain and the tactile domain.

[0135] Existing method three: literature

[0136] In the paper "Learning cross-modal visual-tactile representation using ensembled generative adversarial networks" (Authors: Li X, Liu H, Zhou J, et al), based on deep learning, the texture attributes of images are taken as input, and after obtaining the category information of the images, they are taken as the selection condition of GAN, so as to realize the cross-modal conversion of texture images to haptic signals through corresponding GAN training.

[0137] The present invention: the method of the present embodiment.

[0138] In the experiment, similarity (SIM) is used as an evaluation index to evaluate the effect of cross-modal generation. SIM is used to measure the similarity between the real haptic signal and the reconstructed haptic signal, and is defined as:

[0139]

[0140] Wherein, the exponent in the denominator h i ,h i ′) is the Chebyshev distance between h i and h i ′ waveform. That is, the maximum distance between two waveform points. The range of SIM is [0, 1]. When the difference between the two is small, SIM is closer to 1, and the two waveforms are more similar, and vice versa.

[0141] Table 1 is an experimental result of the present application

[0142]

[0143] From Table 1 and Figure 5 It can be seen that compared with the above-mentioned state-of-the-art method, the method proposed by us has obvious advantages. All other comparative model inputs are only visual signals, while our model fuses visual and tactile signals, extracts high-level semantic information of the same object in two modalities, better extracts inter-modal common information and intra-modal unique information, and thus better restores the tactile signal. Figure 5 The tactile reconstruction result graph of the comparative method is shown in Figure 6 The tactile spectrum and sequence reconstruction result graph of the method of the present application is shown.

[0144] Secondly, in order to further verify the adaptability of the method of the present application to signal-to-noise ratio, the tactile reconstruction effect of the method of the present application and the comparative method under different signal-to-noise ratios is compared. Here, the comparative method adopts a self-encoder-decoder structure composed of multiple linear layers and convolutional layers instead of the cross-modal joint encoder-decoder structure in the present application. Then, experiments are carried out under different test signal-to-noise ratios, and the structure is as shown in Figure 7 Under different signal-to-noise ratios, the method of the present application can reconstruct the tactile signal in high quality, showing the self-adaptive characteristics to the change of signal-to-noise ratio, while the tactile reconstruction effect of the comparative method under low signal-to-noise ratio is not ideal.

[0145] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A method for haptic signal reconstruction of a signal-to-noise ratio adaptive joint optimization of a source channel, characterized in that, The method comprises the following steps: Preprocessing the image signal and the tactile vibration signal generated by sliding corresponding to each object based on a tactile-image database; Inputting the preprocessed image and tactile signal into a cross-modal joint encoder, performing block processing on the input data, sequentially inputting a hierarchical feature extraction module and a signal-to-noise ratio adaptive module, and obtaining hierarchical fusion features; After the hierarchical fusion features are transmitted through a channel, the hierarchical fusion features enter a cross-modal joint decoder, and the cross-modal joint decoder comprises a hierarchical feature extraction module, a signal-to-noise ratio adaptive module and a signal reconstruction module; The hierarchical feature extraction module decodes tactile low-dimensional features from the hierarchical fusion features, the signal-to-noise ratio adaptive module dynamically adjusts the low-dimensional features according to channel changes, and the signal reconstruction module converts the low-dimensional features into a tactile frequency spectrum, the tactile frequency spectrum is converted into a tactile sequence through inverse short-time Fourier transform ISTFT, and thus the tactile signal reconstruction is completed; Joint training of the cross-modal joint encoder and the cross-modal joint decoder is performed according to training data, and optimal model parameters are obtained; The measured tactile and image signals are input into the optimal cross-modal joint encoder to obtain hierarchical fusion features, and after channel transmission, the received hierarchical fusion features are input into the cross-modal joint decoder to generate a target tactile frequency spectrum, and the target tactile signal is obtained through ISTFT transformation.

2. The SNR adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 1, wherein: The preprocessing comprises cutting the image according to the HxW size, which is represented as, wherein v j represents the jth picture in the data set with size HxW in RGB three channels; The preprocessing of the tactile signal comprises short-time Fourier transform STFT of the tactile sequence to obtain the frequency spectrum of the tactile sequence, which is represented as, passing the spectrum through and obtaining an amplitude spectrum and a phase spectrum Splicing amplitude spectrum and phase spectrum to obtain a haptic signal RGB image V and haptic signal In the channel dimension to five-dimensional data where n = H x W x 5 represents the size of the image-haptic signal pair, and the pairing of image and haptic signal is maintained throughout the processing.

3. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 2, wherein: The block processing comprises dividing the data pair into non-overlapping blocks with a size of 4x4; The feature dimension of each patch is 4 x 4 x 5 = 80 dimensions, and the size of the data tensor is The block-processed data is input into a plurality of hierarchical feature extraction modules and signal-to-noise ratio adaptive modules; The hierarchical feature extraction module is FE e The signal-to-noise ratio adaptive module is SD e (·), respectively connected after the first and last feature extraction components; Let the number of feature extraction components in the encoder be L, the input of the jth feature extraction component be t j j∈{1,2,...,L}, the continuous calculation process is represented as, Touch-image signal pair Divide the image into non-overlapping small blocks in both length and width directions to obtain t1, input the hierarchical feature extraction module to obtain the first layer feature representation Input the first layer feature representation into the signal-to-noise ratio adaptive module to obtain t2, input t2 into multiple feature extraction components, obtain deep layer features after passing through the last layer feature extraction component, input the last layer signal-to-noise ratio adaptive layer, and obtain the remapped features e2.

4. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 3, wherein: The hierarchical feature extraction module FE e (·) includes a merging layer and a Swin Transformer module, learning hierarchical features of the tactile-image signal; The merging layer is responsible for down-sampling and dimension increasing to generate features at different levels, and the SwinTransformer module performs hierarchical feature learning; The first merging layer connects the features of each group of 2x2 adjacent patches in the channel dimension, so that the feature resolution is down-sampled by 2 and the feature dimension is increased by 4 times; The full connection layer is applied on the connected features, and the output dimension is unified to twice of the original dimension, i.e. 2C, and the tensor size is After passing through the Swin Transformer module and the merging layer multiple times, hierarchical feature representation is obtained; The last layer merging layer of the cross-modal joint encoder is responsible for mapping the deep feature information to the same dimension as the channel channel size. After connecting the adjacent patch features, a linear layer is applied to transform the output dimension to the channel dimension b. At this time, the tensor size is 5. The SNR adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 4, wherein: The signal-to-noise ratio adaptive module SD e (·) includes global pooling and a feedforward network; The global pooling is used to capture full-size information of input features, to summarize local features into global features, to connect signal-to-noise ratio information with the pooled features, to input a feedforward network to generate a scaling factor, and to multiply the output features of the hierarchical feature extraction module FE e of the application (·) to obtain weighted features. The output of the feature extraction component is o = {o1, o2,..., o c}, o ipq denotes the element where the i-th channel is located at (p, q). After global average pooling P(·), each feature map is pooled in the channel dimension to extract global information The global pooling operation sums up each two-dimensional feature map o i into a single numerical value s i representing the average activation strength of the entire feature map, denoted as Concatenating the global information with the current SNR value, we get Inputting it into the feedforward network, we have The original input features before pooling are multiplied by the channel condition to obtain the adjusted features The original input features before pooling are multiplied by the channel condition to obtain the adjusted features The original input features before pooling are multiplied by the channel condition to obtain the adjusted features Since the sizes are different, the data of each channel is copied by factor before multiplication, so that the size is expanded to HxWxc.

6. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 5, wherein: The cross-modal joint decoder uses an expansion layer to up-sample the extracted deep features, the expansion layer reshapes the feature maps of adjacent dimensions and reduces the feature dimension to half of the original dimension, thereby completing the dimension reduction of the features; The channel signal is input into a plurality of hierarchical feature extraction modules to obtain features at different levels, and a signal-to-noise ratio adaptive module is connected after the first and last hierarchical feature extraction modules; Let the number of hierarchical feature extraction modules in the cross-modal joint decoder be L, and the input of the jth hierarchical feature extraction module according to the data flow direction be t j j∈{1,2,...,L}, the continuous calculation process is represented as, Signal transmitted from a channel The input hierarchical feature extraction module obtains a first layer reduced dimension feature representation The input signal-to-noise ratio adaptive module obtains t2, t2 is reduced in dimension through a plurality of hierarchical feature extraction modules, and then passes through a signal-to-noise ratio adaptive layer again, performs feature scaling on the shallow layer feature after the last dimension reduction, obtains the remapped feature e2, and performs 4 times scaling through an expansion layer to convert the feature size into The dimension-reduced features are input into a signal reconstruction module, which maps the features from the feature domain to a time-frequency matrix containing amplitude and phase information of the tactile sequence through multiple convolution layers, which is represented as, The multi-level features are input into multiple convolutions, the size is kept unchanged, the channel number is scaled, the channel number is scaled to 2, and the time-frequency matrix containing amplitude and phase information of the tactile sequence is mapped, and the recovered tactile sequence is obtained through ISTFT transformation.

7. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 6, wherein: The joint training comprises selecting a Charbonnier loss as a loss function of the optimization process, which is represented as, wherein, h represents the haptic signal recovered by the decoder, j h represents the real haptic signal, ε = 10 -3 is a constant; The cross-modal joint encoder and the cross-modal joint decoder are jointly trained as a whole network, gradients are calculated by an Adam optimization algorithm, and network parameters are updated by back propagation, and the expression is as follows: where m t and v t are the unbiased estimates of the first and second moments, respectively, g t is the gradient of the parameters, beta1 and beta2 are the decay rates of the first and second moment estimates, respectively, μ is the learning rate, ε is a very small constant to ensure the denominator is not zero, θ t and θ t+1 are the network parameters before and after updating, respectively, where the network parameters refer to the network parameters of the cross-modal joint encoder and the cross-modal joint decoder, the learning rate is selected as 1x10-4, the number of iterations is 1000, and the network parameters are determined after the iteration is completed to complete the model training.

8. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 7, wherein: The reconstruction of the target haptic signal includes passing through the cross-modal joint encoder to generate a hierarchical fusion feature, transmitting the feature through a channel, decoding the feature by the cross-modal joint decoder, reconstructing the spectrum of the target haptic signal, and obtaining the expected haptic signal through ISTFT transformation. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor implements the steps of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of a source channel according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of a source channel according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • 6G-oriented tactile modal signal reconstruction method

    CN114842384A

  • Audio-visual assisted fine-grained tactile signal reconstruction method

    CN115905838A