Signal-to-noise ratio adaptive haptic signal reconstruction method with joint source-channel optimization

By using a unified model design for cross-modal joint encoders and decoders and an adaptive signal-to-noise ratio module, the robustness of cross-modal signal reconstruction models under the influence of channel noise was solved, and stable tactile signal reconstruction under variable channel conditions was achieved.

WO2026031416A1PCT designated stage Publication Date: 2026-02-12NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
PCT/CN2024/136301
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-07
Filing Date
2024-12-03
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing cross-modal signal reconstruction models are not robust enough to channel noise and have poor generalization ability to signal-to-noise ratio, resulting in poor stability of tactile signal reconstruction and inability to adapt to changing channel conditions.

Method used

A unified model design of cross-modal joint encoder and decoder is adopted, combined with a signal-to-noise ratio adaptive module. Through hierarchical feature extraction and adaptive adjustment of signal-to-noise ratio, the signal encoding, transmission and decoding process is optimized, and the robustness of the model to channel changes is improved.

Benefits of technology

Stable reconstruction of tactile signals under different signal-to-noise ratios was achieved, significantly improving the robustness and accuracy of signal reconstruction, reducing the amount of data, and maintaining information integrity and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136301_12022026_PF_FP_ABST
    Figure CN2024136301_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of haptic signal generation. Disclosed is a signal-to-noise ratio adaptive haptic signal reconstruction method with joint source-channel optimization, comprising: preprocessing an image corresponding to each object and a haptic vibration signal generated by sliding; constructing a cross-modal joint encoder for generating hierarchical fusion features, constructing a cross-modal joint decoder, and inputting preprocessed training data into the constructed cross-modal joint encoder and cross-modal joint decoder for training; and inputting in pairs haptic signals and image signals under test into the cross-modal joint encoder and the optimal cross-modal joint decoder to reconstruct a target haptic signal. The signal-to-noise ratio adaptive haptic signal reconstruction method with joint source-channel optimization provided by the present invention simplifies encoding training processes, better performs modal feature fusion, improves the robustness of models to channel variations, and allows the models to exhibit higher haptic signal reconstruction stability under a wider range of signal-to-noise ratio variations.
Need to check novelty before this filing date? Find Prior Art

Description

A haptic signal reconstruction method of signal-to-noise ratio adaptive joint optimization of signal source channel TECHNICAL FIELD

[0001] The present application relates to the technical field of haptic signal generation, in particular to a haptic signal reconstruction method of signal-to-noise ratio adaptive joint optimization of signal source channel. BACKGROUND

[0002] With the progress of technology and the continuous evolution of user needs, traditional audio-visual media has been insufficient to meet the growing expectations of users for immersive and interactive experiences, which has driven the evolution of multimedia services to a higher form. However, existing research has focused on the analysis and transmission of single modal signals, failing to fully address the complex transmission needs of heterogeneous multi-modal signals. Therefore, cross-modal services have emerged, which create a new mode of communication and experience for users by integrating visual, auditory and haptic sensory information. In particular, cross-modal services centered on haptics are widely considered a key development direction for future communication technologies, as they can provide direct and realistic interactive feedback. However, current research on cross-modal services centered on haptics still faces several challenges.

[0003] On the one hand, current cross-modal signal reconstruction schemes have not fully considered the impact of channel noise, and need to further consider the robustness of signal reconstruction under different channel conditions. In actual communication scenarios, haptic signals need to be encoded at the signal source, transmitted in the channel, and decoded and reconstructed at the receiving end. This process involves the interaction of multiple factors, including semantic fusion of cross-modal signals, transmission characteristics of the channel, and reconstruction strategies at the receiving end, etc.

[0004] On the other hand, existing signal communication reconstruction models have low generalization to signal-to-noise ratio in actual applications, and need to be trained and deployed according to different channel conditions, occupying more computing and storage resources. Existing models based on deep neural networks rely on data-driven strategies. The optimization process of the entire system is based on an end-to-end loss function that restores accuracy, and the parameters of the model are adjusted through supervised learning. This model training mechanism easily leads to low generalization of the model to signal-to-noise ratio, resulting in good performance of the model at certain specific signal-to-noise ratios during training, but rapid performance decline at other signal-to-noise ratios.

[0005] Existing research has also made some progress in the cross-modal fusion and generation of haptic signals. In the literature "Information recovery technology for cross-modal communication", Xu et al. proposed an information recovery scheme for the problem of loss or noise pollution of wireless channels that multi-modal data (such as audio, video, and haptic signals) may encounter during transmission. The scheme uses existing multi-modal data at the receiving end to achieve information recovery through same-modal one-to-one search, cross-modal one-to-one search, cross-modal one-to-many search, and other methods. Wei et al. proposed an audio-visual assisted haptic signal reconstruction method based on cloud-edge collaboration in the literature "Haptic Signal Reconstruction for Cross-Modal Communications". For the case of insufficient data received by the edge node, a large audio-visual assisted database stored in the central cloud is used to obtain and transmit knowledge and migrate to the edge node. Then, the edge node is used to mine the inherent semantic consistency and correlation between modalities, and finally the required haptic signal reconstruction is realized. By considering the received audio-visual signals and relating them to the potential correlation of the haptic modality, the reconstruction of the haptic signal is achieved.

[0006] However, these models are mostly limited to one-way conversion at the sending end or receiving end, without considering end-to-end design of the entire communication process, and lack simulation of transmission in actual channels. In actual situations, it is easy to be disturbed by noise, resulting in a decline in the quality of the transmitted data. This will have an adverse effect on multi-modal data streams, especially haptic streams, which require low latency and high reliability. In addition, for a variety of channel conditions and varying channel signal-to-noise ratios, how to design effective channel adaptive technology to significantly improve the stability of haptic signal reconstruction is also a difficult problem faced by haptic-based cross-modal communication. SUMMARY

[0007] In view of the above problems, the present application is proposed.

[0008] Therefore, the present application realizes the encoding, transmission and decoding process of signals based on deep neural networks. On the one hand, through the unified model design of cross-modal joint encoder and decoder, the encoding training process is simplified, and at the same time, the unified model also brings unified haptic-image signal feature representation, which better performs modal feature fusion. On the other hand, through the separately designed signal-to-noise ratio adaptive module, the robustness of the model to channel changes is improved, so that the model performs higher haptic signal reconstruction stability under a larger range of signal-to-noise ratio changes.

[0009] To solve the above technical problems, the present application provides the following technical solutions: a signal-to-noise ratio adaptive joint optimization of haptic signal reconstruction method of source channel, comprising:

[0010] The image signal corresponding to each object and the tactile vibration signal generated by sliding are preprocessed based on a tactile-image database.

[0011] The preprocessed image and tactile signal are input into a cross-modal joint encoder, the input data are block processed, sequentially input into a hierarchical feature extraction module and a signal-to-noise ratio adaptive module, and hierarchical fusion features are obtained.

[0012] The hierarchical fusion features are transmitted through a channel and then input into a cross-modal joint decoder, which includes a hierarchical feature extraction module, a signal-to-noise ratio adaptive module and a signal reconstruction module.

[0013] The hierarchical feature extraction module decodes the tactile low-dimensional features from the hierarchical fusion features, the signal-to-noise ratio adaptive module dynamically adjusts the low-dimensional features according to the channel change, the low-dimensional features are converted into a tactile spectrum by the signal reconstruction module, the tactile spectrum is converted into a tactile sequence through inverse short-time Fourier transform ISTFT, and thus the tactile signal reconstruction is completed.

[0014] The cross-modal joint encoder and the cross-modal joint decoder are jointly trained according to the training data, and the optimal model parameters are obtained.

[0015] The to-be-tested tactile and image signals are input into the optimal cross-modal joint encoder to obtain hierarchical fusion features, the received hierarchical fusion features are input into the cross-modal joint decoder after channel transmission, a target tactile spectrum is generated, and a target tactile signal is obtained through ISTFT transformation.

[0016] As a preferred scheme of the tactile signal reconstruction method for jointly optimizing a signal source and a channel according to the signal-to-noise ratio, the preprocessing includes cutting an image according to an HxW size, and the image is represented as:

[0017] Wherein, v j represents the jth picture with an HxW size in the RGB three-channel picture in the data set.

[0018] The tactile signal is preprocessed, the tactile sequence is subjected to short-time Fourier transform STFT to obtain a spectrum of the tactile sequence, and the spectrum is represented as:

[0019] The spectrum is subjected to and

[0020] , to obtain an amplitude spectrum and a phase spectrum

[0021] The amplitude spectrum and the phase spectrum are spliced to obtain a tactile signal

[0022] RGB image V and haptic signal H are spliced in channel dimension into five-dimensional data where n = H x W x 5 represents the size of the image-haptic signal pair, and the one-to-one pairing of the image and the haptic signal is maintained during processing.

[0023] As a preferred scheme of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio and source channel according to the present application, wherein: the block processing includes dividing the data pair into non-overlapping blocks of 4x4 size.

[0024] The feature dimension of each patch is 4x4x5 = 80, and the size of the data tensor is

[0025] The blocked data is input into a plurality of hierarchical feature extraction modules and a signal-to-noise ratio adaptive module.

[0026] The hierarchical feature extraction module is FE e (·), and the signal-to-noise ratio adaptive module is SD e (·), which are connected after the first and last feature extraction components, respectively.

[0027] Let the number of feature extraction components in the encoder be L, and the input of the jth feature extraction component be t j ,j∈{1,2,...,L}, and the continuous calculation process is represented as:

[0028] The haptic-image signal pair d∈D is divided into non-overlapping small blocks in the length and width directions to obtain t1, which is input into the hierarchical feature extraction module to obtain the first layer feature representation The first layer feature representation is input into the signal-to-noise ratio adaptive module to obtain t2, which is input into a plurality of feature extraction components, and the deep layer feature is obtained after the last layer feature extraction component. The last layer signal-to-noise ratio adaptive layer is input to obtain the feature e2 that is remapped.

[0029] As a preferred scheme of the haptic signal reconstruction method of adaptive joint optimization of signal-to-noise ratio and source channel according to the present application, wherein: the hierarchical feature extraction module FE e (·) includes a merging layer and a Swin Transformer module, and learns the hierarchical features of the haptic-image signal.

[0030] The merging layer is responsible for downsampling and dimensionality increase to produce features at different levels, while the Swin Transformer block performs hierarchical feature learning.

[0031] The first merging layer connects the features of each group of 2x2 adjacent patches in the channel dimension, so that the feature resolution is downsampled by 2 and the feature dimension is increased by 4 times.

[0032] The full connection layer is applied on the connected features to unify the output dimension to twice of the original dimension, i.e., 2C, and the tensor size is

[0033] After multiple times of passing through the Swin Transformer block and the merging layer, the hierarchical feature representation is obtained.

[0034] The last merging layer of the cross-modal joint encoder is responsible for mapping the deep feature information to the dimension same as the channel dimension, and after connecting the adjacent patch features, a linear layer is applied to transform the output dimension to the channel dimension b, and at this time the tensor size is

[0035] As a preferred scheme of the signal-to-noise ratio adaptive joint optimization method of the source channel of the haptic signal reconstruction method, wherein: the signal-to-noise ratio adaptive module SD(·) includes global pooling and a feedforward network.

[0036] The global pooling is used to capture the full-size information of the input features, and the local features are summarized as global features, and the signal-to-noise ratio information is connected with the pooled features to input the feedforward network to generate a scaling factor, and the scaling factor is multiplied with the output features of the hierarchical feature extraction module FE e (·) to obtain weighted features.

[0037] The output of the feature extraction component is o=FE(t)∈R H×W×c , o={o1,o2,...,o c}, o ipq represents the element of the i-th channel at (p,q).

[0038] After global average pooling P(·), each feature map is pooled in the channel dimension to extract global information

[0039] The global pooling operation compresses each two-dimensional feature map o i into a single value s i representing the average activation strength of the entire feature map, which is represented as:

[0040] The global information is concatenated with the current SNR value to obtain which is input into the feedforward network, represented as:

[0041] The product of and the original input feature o∈R H×W×c before pooling is multiplied to obtain the feature e∈R H×W×cSince the two sizes are different, the data of each channel of the factor is copied before multiplication, so that the size is expanded to HxWxc.

[0042] As a preferred scheme of the haptic signal reconstruction method of adaptive SNR joint optimization of source channel according to the application, wherein: the cross-modal joint decoder uses an expansion layer to upsample the extracted deep features, the expansion layer reshapes the feature maps of adjacent dimensions and reduces the feature dimension to half of the original dimension, and completes the dimension reduction of the features.

[0043] The channel signal is input into a plurality of hierarchical feature extraction modules to obtain features at different levels, and an SNR adaptive module is connected after the first and last hierarchical feature extraction modules.

[0044] Suppose the number of hierarchical feature extraction modules in the cross-modal joint decoder is L, and the input of the jth hierarchical feature extraction module according to the data flow is t j ,j∈{1,2,...,L}, the continuous calculation process is represented as:

[0045] The signal transmitted from the channel The first layer of dimension-reduced feature representation is obtained by inputting the hierarchical feature extraction module The t2 is input into the SNR adaptive module, and after dimension reduction by a plurality of hierarchical feature extraction modules, the SNR adaptive layer is again used to scale the shallow features after the last dimension reduction, to obtain the remapped features e 2 The 4-fold scaling is performed by an expansion layer, and the feature size is converted to

[0046] The dimension-reduced features are input into the signal reconstruction module, and through a plurality of convolution layers, the features are mapped from the feature domain to the time-frequency matrix containing the amplitude and phase information of the haptic sequence, represented as:

[0047] The multi-level features are input into a plurality of convolutions, the size is kept unchanged, the channel number is scaled, the channel number is scaled to 2, and the time-frequency matrix containing the amplitude and phase information of the haptic sequence is mapped, and the recovered haptic sequence is obtained by ISTFT transformation.

[0048] As a preferred scheme of the haptic signal reconstruction method of adaptive SNR joint optimization of source channel according to the application, wherein: the joint training includes selecting Charbonnier loss as the loss function of the optimization process, represented as:

[0049] Wherein, represents the haptic signal recovered by the decoder, h jrepresent the real haptic signal, and ε = 10 -3 is a constant.

[0050] The cross-modal joint encoder and the cross-modal joint decoder are jointly trained as a whole network, the gradient is calculated through the Adam optimization algorithm, and the network parameters are updated through back propagation. The expression is as follows:

[0051] Where, m t and v t are the unbiased estimates of the first and second moments respectively. g t is the gradient of the parameter, beta1 and beta2 are the decay rates of the first and second moment estimates respectively. μ is the learning rate, and ε is a very small constant to ensure that the denominator is not zero. θ t and θ t+1 are the network parameters before and after updating respectively. The network parameters here refer to the network parameters composed of the cross-modal joint encoder and the cross-modal joint decoder. The learning rate is selected as 1x10-4, the iteration number is 1000 times, and the network parameters are determined after iteration to complete the model training.

[0052] As a preferred scheme of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of the source channel, wherein: the reconstruction target haptic signal includes a hierarchical fusion feature generated by the cross-modal joint encoder, and the target haptic signal spectrum is reconstructed by the cross-modal joint decoder after channel transmission, and the expected haptic signal is obtained by ISTFT transformation.

[0053] A computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of the source channel as described above when executing the computer program.

[0054] A computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of the source channel as described above.

[0055] The haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of the source channel provided by the present application fully utilizes the advantages of multi-modal feature fusion, effectively fuses different modal semantic information, significantly compresses the source signal, reduces the data volume, and maintains the integrity and accuracy of the information. Further, in order to cope with the influence of channel noise and changes on signal transmission, the present application specially designs a set of channel encoder and decoder. The encoder and decoder can improve the robustness of the signal to channel changes and ensure that the signal remains stable and reliable under various channel conditions. BRIEF DESCRIPTION OF DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some of the embodiments of the present application, and all other drawings obtained by those skilled in the art without creative labor based on the embodiments in the present application should also fall within the protection scope of the present application.

[0057] Fig. 1 is a whole flow chart of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application.

[0058] Fig. 2 is a network structure schematic diagram of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application.

[0059] Fig. 3 is a Swin Transformer module schematic diagram of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application.

[0060] Fig. 4 is a signal-to-noise ratio adaptive module schematic diagram of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application.

[0061] Fig. 5 is a haptic signal reconstruction result diagram of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application and other comparative methods.

[0062] Fig. 6 is a haptic spectrum and sequence reconstruction result diagram of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application.

[0063] Fig. 7 is a haptic reconstruction stability comparison diagram of a haptic signal reconstruction method of a signal-to-noise ratio adaptive joint optimization of a source channel provided by an embodiment of the present application and other methods. DETAILED DESCRIPTION

[0064] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings in the specification. Obviously, the described embodiments are only some of the embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the protection scope of the present application.

[0065] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description, that the present application can be practiced with other systems, and that the present application can be practiced using different techniques. Therefore, the present application is not limited to the embodiments set forth in this description but is instead broadly commensurate with the claims.

[0066] Embodiment 1

[0067] Referring to FIGS. 1-4, for one embodiment of the present application, a method for self-adaptive optimization of signal-to-noise ratio of a haptic signal reconstruction method for a joint source channel is provided, including:

[0068] Step 1: Preprocess the image signal corresponding to each object and the haptic vibration signal generated by sliding on the basis of the disclosed haptic-image database. Convert the haptic vibration signal generated by sliding into an amplitude spectrum and a phase spectrum by short-time Fourier transform (STFT) to represent, and cut the image signal to the size to align the haptic frequency spectrum size.

[0069] Step 1-1, first, preprocess the image signal. Cut the image to HxW size, denoted as: wherein v j represents the jth RGB three-channel picture with a size of HxW in the data set.

[0070] Step 1-2, preprocess the haptic signal, and perform short-time Fourier transform (STFT) on the haptic sequence to obtain the frequency spectrum of the haptic sequence When performing STFT transform again, a suitable window function and window size should be selected, and the constraint of "Constant OverLap Add" (COLA) should be met, so as to ensure that there is no information loss when performing STFT transform. In this way, when restoring the haptic, it can also be restored to the haptic sequence through ISTFT transform. Calculate the amplitude spectrum and the phase spectrum by the following formula, and then splice them to obtain the haptic signal

[0071] Finally, splice the RGB image V and the haptic signal H in the channel dimension into five-dimensional data wherein n=HxWx5 represents the size of the image-haptic signal pair. In the processing process, the one-to-one pairing of the image and the haptic signal is maintained.

[0072] Step 2: The preprocessed image and tactile signal are input into the cross-modal joint encoder. First, the input data is blocked, which can reduce the subsequent embedding sequence length and more effectively capture long-distance dependencies. Then, the hierarchical feature extraction module and the signal-to-noise ratio adaptive module are input in turn, and finally the hierarchical fusion features are obtained, as shown in FIG. 2. Among them, the hierarchical feature extraction module extracts multi-level features of image-tactile signals and fuses them, and the signal-to-noise ratio adaptive module dynamically adjusts the multi-level fusion features generated by the hierarchical feature extraction module according to the real-time channel conditions to adapt to changing channel environments.

[0073] Step 2-1, the preprocessed image and tactile signal are input into the cross-modal joint encoder. The input data is blocked, which can reduce the subsequent embedding sequence length and more effectively capture long-distance dependencies. In addition, cutting the image into small blocks can preserve local structure information. Specifically, the data pair is divided into 4x4 non-overlapping blocks (patches), and through this division method, the feature dimension of each patch is 4x4x5=80 dimensions. At this time, the size of the data tensor is

[0074] Step 2-2, then the blocked data is input into multiple hierarchical feature extraction modules and signal-to-noise ratio adaptive modules. The hierarchical feature extraction module is denoted as FE e (·). Among them, the signal-to-noise ratio adaptive module is denoted as SD e (·), which is connected after the first and last feature extraction components respectively, so as to learn the feature representation under different signal-to-noise ratios of shallow and deep levels, as shown in FIG. 2. Specifically, assuming that the number of feature extraction components in the encoder is L, the input of the jth feature extraction component is t j ,j∈{1,2,...,L}. The continuous calculation process is as follows:

[0075] First, the tactile-image signal pair d∈D is divided into non-overlapping small blocks in the length and width directions to obtain t1, which is input into the hierarchical feature extraction module to obtain the first layer feature representation Then, the first layer feature representation is input into the signal-to-noise ratio adaptive module to obtain t2, and then t2 is input into multiple feature extraction components. After passing through the last layer feature extraction component, the deep layer feature is obtained, and then input into the last layer signal-to-noise ratio adaptive layer to obtain the feature e2 that is remapped.

[0076] Step 2-3, among them, the hierarchical feature extraction module FE e(·) includes a merging layer and a Swin Transformer module, which learns the hierarchical features of the tactile-image signal. Among them, the merging layer is mainly responsible for downsampling and dimensionality, to produce features at different levels, while the Swin Transformer block is mainly for hierarchical feature learning. The first merging layer connects the features of each group of 2x2 adjacent patches in the channel dimension. Through such processing, it will make the feature resolution downsampled by 2, and the feature dimension increased by 4 times due to the connection operation. Therefore, a fully connected layer is applied on the connected features to unify the output dimension to twice the original dimension, that is, 2C, at this time, the tensor size is Then, after passing through the Swin Transformer block and the merging layer multiple times, the hierarchical feature representation is obtained. The last layer merging layer of the cross-modal joint encoder is responsible for mapping the deep feature information to the same dimension as the channel channel size. Therefore, unlike the previous merging layers, after connecting the adjacent patch features, a linear layer is applied to transform the output dimension to the channel dimension b, at this time, the tensor size is

[0077] The Swin Transformer module is the key to extracting features, which will be described in detail below. As shown in FIG. 3, two consecutive Swin Transformer modules, each module consists of a normalization (LayerNorm, LN) layer, a multi-head self-attention module, a residual connection, and a multi-layer perceptron (MLP). The two connected Swin Transformer modules use window-based multi-head self-attention module (W-MSA) and displacement window-based multi-head self-attention module (SW-MSA) respectively.

[0078] In the standard ViT (Vision Tansformer), global attention is used, which calculates the relationship between the target token and all other tokens, increasing the computational complexity. The Swin Transformer module is based on a displacement window. Each input tensor of the Swin Transformer is evenly divided into several non-overlapping windows, which only calculate self-attention within the local window. Because the number of patches within the window is much smaller than the number of patches in the picture, the computational complexity is greatly reduced. On the other hand, WMSA only calculates self-attention within the window, and there is a lack of information exchange between non-overlapping windows. In order to improve the model's ability to build global relationships, SW-MSA needs to be built to achieve cross-window information exchange and increase the receptive field by introducing displacement windows.

[0079] The method of shifting windows can cause a problem because the window size does not align with the tensor size, resulting in an increase in the number of windows and different element sizes in each window after shifting. The document "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" proposes using a cyclic shift with masking to maintain the number of windows after shifting and uniform elements.

[0080] Specifically, the method of shifting windows is implemented between two consecutive Swin Transformer modules. The first module performs WMSA operation to calculate attention within the window, and the second module performs SW-MSA operation to move the window to calculate attention. Assuming that the input of the i-th Swin Transformer module is u i , the specific calculation process is as follows:

[0081] where LN represents the normalization layer, WMSA represents the window multi-head self-attention module, MLP represents the multi-layer perceptron, and SW-MSA represents the shifted window multi-head self-attention module. The output of the i-th Swin Transformer module is denoted as u

[0082] The input token sequence generates three matrices, query, key, and value, denoted as Q, K, and V. d represents the dimension of query or key. B represents the relative position bias.

[0083] Steps 2-4, the SNR adaptive module SD(·) is the key to ensure that the model adapts to different channel levels, which includes global pooling and feedforward network, as shown in FIG. 4. Global pooling is used to capture the full-size information of the input features, and local features are summarized into global features. The SNR information is connected with the pooled features, and the feedforward network (FFN) is used to generate scaling factors. These scaling factors are multiplied with the output features of the hierarchical feature extraction module FE e (·), and the weighted features are obtained. These weighted features are then used in the subsequent encoding and transmission process. The purpose of multiplication is to adjust the importance of features according to different SNR conditions, so that the model can more effectively handle channel noise and transmit data.

[0084] First, the output of the hierarchical feature extraction component is o = FE(t) ∈ R H×W×c , o = {o1, o2,..., o c}, o ipqThe element represents the i-th channel located at (p, q). After global average pooling P(·), each feature map is pooled in the channel dimension, which can extract global information The global pooling operation compresses each two-dimensional feature map o i into a single numerical value s i , which represents the average activation strength of the entire feature map.

[0085] Then, these global information and the current SNR value are concatenated to obtain The input of a feedforward network FFN(·) is obtained, and the output of the network is a scaling factor that determines the weight of each element in the subsequent feature map. The feedforward network consists of multiple linear layers, and the activation function of the last layer is Sigmoid, and the activation function of the remaining layers is Relu. In this way, the output scaling factor is ensured to be between 0 and 1.

[0086] Finally, the product of and the original input feature o∈R H×W×c before pooling is obtained, and the feature e∈R H×W×c adjusted according to the channel condition is obtained. Since the sizes of the two are different, before multiplication, the data of each channel of the factor is copied, so that the size is expanded to HxWx c.

[0087] Through the scaling factor predicted by the feedforward network, the SNR adaptive module can adjust the weight of each element in the original feature map. The SNR adaptive module does not change the size of the input feature. Specifically, each element in the feature map is multiplied by the corresponding scaling factor, thereby realizing the recalibration of the feature. If the scaling factor is close to 0, it means that the corresponding feature is not important under the current SNR condition and should be suppressed; if the scaling factor is close to 1, the feature remains unchanged or is enhanced. Unlike traditional binary masks, the scaling factor is equivalent to a kind of soft mask, and the scaling factor is a value between 0 and 1, which is used to adjust the weight of the feature in proportion, rather than simply switching operation. This method provides a more delicate and continuous control method, allowing the model to adjust the importance of the feature in a more fine-grained manner according to the current channel condition and image content.

[0088] Step 3: After the hierarchical fusion features are transmitted through the channel, they enter the cross-modal joint decoder. The cross-modal joint decoder adopts a symmetric design similar to the encoder structure, including a hierarchical feature extraction module, a signal-to-noise ratio adaptive module, and a signal reconstruction module, as shown in FIG. 2. Among them, the hierarchical feature extraction module decodes the haptic low-dimensional features from the hierarchical fusion features, the signal-to-noise ratio adaptive module dynamically adjusts the low-dimensional features according to the channel changes, and finally, the signal reconstruction module converts the low-dimensional features into haptic frequency spectrum. The haptic frequency spectrum is converted into a haptic sequence through inverse short-time Fourier transform (ISTFT), thereby completing the reconstruction of the haptic signal.

[0089] Step 3-1: After the hierarchical fusion features are transmitted through the channel, they are input into the cross-modal joint decoder, which has a structure corresponding to the cross-modal joint encoder. The cross-modal joint decoder is composed of multiple hierarchical feature extraction modules a signal-to-noise ratio adaptive module and a signal reconstruction module. At this time, the hierarchical feature extraction module is composed of an expansion layer and a Swin Transformer module. Unlike the merging layer used in the encoder, the expansion layer is used in the decoder to upsample the extracted deep features. The expansion layer reshapes the feature maps of adjacent dimensions into larger size feature maps (2x upsampling) and correspondingly reduces the feature dimension to half of the original dimension, completing the dimension reduction of the features.

[0090] Step 3-2: The channel signal is input into multiple hierarchical feature extraction modules to obtain features at different levels. Similarly, the signal-to-noise ratio adaptive module is connected after the first and last hierarchical feature extraction modules.

[0091] Specifically, assuming that the number of hierarchical feature extraction modules in the decoder is L, the input of the jth hierarchical feature extraction module is t j ,j∈{1,2,...,L} according to the data flow. The continuous calculation process is as follows:

[0092] From the channel-transmitted signal input the hierarchical feature extraction module to obtain the first layer of reduced dimension feature representation Next, input the signal-to-noise ratio adaptive module to obtain t2. After that, t2 is reduced in dimension by multiple hierarchical feature extraction modules, and then passes through a layer of signal-to-noise ratio adaptive layer again to scale the shallow features after the last dimension reduction, obtaining the remapped features e2. Second, through a layer of expansion layer, the feature size is converted to

[0093] Step 3-3: Input the reduced dimension features into the signal reconstruction module, and map from the feature domain to the time-frequency matrix containing the amplitude and phase information of the haptic sequence through multiple convolution layers In this stage, multiple convolutions are applied to multi-level feature inputs while maintaining the same size, scaling down the number of channels until the number of channels is reduced to two, accurately mapping them to the time-frequency matrix of the amplitude and phase information of the tactile sequence. Finally, the recovered tactile sequence is obtained through ISTFT transformation.

[0094] Step 4: Jointly train the cross-modal joint encoder and cross-modal joint decoder based on the training data to obtain the optimal model parameters.

[0095] Step 4-1: During the training phase, the model receives tactile image signal pairs and channel signal-to-noise ratio (SNR) information for training, learning the mapping relationship under different SNR levels. Charbonnier loss is chosen as the loss function for the optimization process. The calculation formula is shown in the formula below:

[0096] in, h represents the tactile signal recovered by the decoder. j Representing the actual tactile signal, ε = 10 -3 It is a constant. The existence of the constant ε ensures that the gradient is not too small when it approaches zero, and the square root helps to avoid gradient vanishing. The Charbonnier loss function is highly robust to outliers. Compared to mean squared error (L2 loss), it is insensitive to large residuals, which makes the model more stable during training, less susceptible to noise, and easier to converge quickly. Moreover, the calculation of this loss is relatively simple.

[0097] Step 4-2: Jointly train the cross-modal joint encoder and cross-modal joint decoder as a single network. Calculate the gradient using the Adam optimization algorithm and update the network parameters through backpropagation. The expression is as follows:

[0098] Where, m t and v t These are the unbiased estimates of the first and second moments, respectively. t θ is the gradient of the parameters, and beta1 and beta2 are the decay rates of the first and second moment estimates, respectively. μ is the learning rate, and ε is a very small constant used to ensure that the denominator is not zero. t and θ t+1 These are the network parameters before and after the update, respectively. The network parameters here refer to the parameters of the network composed of the cross-modal joint encoder and the cross-modal joint decoder. A learning rate of 1×10⁻⁴ was selected, and the number of iterations was 1000. After the iterations, the network parameters were determined, and the model training was completed.

[0099] Step 5: input the to-be-tested haptic and image signals into the optimal cross-modal joint encoder obtained in the above step to obtain hierarchical fusion features, then after channel transmission, input the received hierarchical fusion features into the cross-modal joint decoder to generate a target haptic spectrum, and obtain a target haptic signal through ISTFT transformation.

[0100] Step 5-1, select the trained optimal cross-modal signal source encoder, channel codec and haptic decoder network model.

[0101] Step 5-2, input the to-be-tested haptic-image signal pair into the model, first, generate a hierarchical fusion feature through the cross-modal joint encoder. Then, after channel transmission, decode by the cross-modal joint decoder and reconstruct the spectrum of the target haptic signal. Finally, obtain the expected haptic signal through ISTFT transformation.

[0102] Embodiment 2, an embodiment of the present application, which is different from the above embodiment is:

[0103] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products, which are stored in a storage medium and include a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program code storage media.

[0104] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer-readable medium for use by an instruction execution system, device or apparatus, such as a computer-based system, a system including a processor or other system that can fetch and execute instructions from an instruction execution system, device or apparatus, or in conjunction with these instructions execution system, device or apparatus. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by an instruction execution system, device or apparatus or in conjunction with these instruction execution system, device or apparatus.

[0105] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for instance via an optical scanner, then compiled, interpreted or otherwise processed in a suitable manner, if necessary, to generate an electronically readable version of the program, which can then be stored in the computer memory.

[0106] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0107] Embodiment 3, referring to FIGS. 5-7, provides a haptic signal reconstruction method of self-adaptive joint optimization of signal-to-noise ratio and source channel, which is an embodiment of the present application. In order to verify the beneficial effects of the present application, scientific demonstration is carried out through simulation experiment.

[0108] This embodiment uses the AUdataset data set for experiments, which is collected by Lasse et al. of the Department of Electrical and Computer Engineering, Aarhus University. The data set contains multimodal signals of 63 objects: vision, motion and haptics (audio vibration signals). The objects selected in the data set are mainly toys and household items. Each object contains 3 pictures at different angles and 5 recorded haptic vibration signals. Among them, there are similar shape and color objects (visual ambiguity) and similar material and texture but different color objects (haptic ambiguity). These objects containing visual ambiguity and haptic ambiguity will reflect the different effects of vision and haptics on perception. This example selects 62 available objects, and augments 3 images of each object to 5 (for example, randomly adjust the brightness and contrast of the picture, flip the image), and 5 recorded haptic vibration signals to form 310 groups of samples. In the data set, 80% is selected for training, and the remaining 20% is used for testing and performance evaluation.

[0109] First, the method of this invention is compared with a simple tactile reconstruction method. Specifically, the method of this invention needs to consider channel interference, while the simple tactile reconstruction method only considers signal reconstruction at the edge. The following three tactile reconstruction methods are selected for experimental comparison:

[0110] Existing Method 1:

[0111] Literature: "Cross-ModalDataGenerationUsingResidue-FusionGANWithFeature-MatchingandPerceptualLosses"

[0112] (Authors Cai S, Zhu K, Ban Y, et al.) used the cGAN (conditional-GAN) structure and residual-fusion (RF) module, and trained the model with additional feature matching (FM) and perceptual loss to achieve mutual generation between visual images and tactile spectra.

[0113] Existing Method Two:

[0114] The paper "Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired People" (authors: Liu H, Guo D, Zhang X, et al.) relies on DiscoGAN to learn the transformation between images and spectrograms. It uses trained DX and DY as discriminators to dual-constrain the generated results, while adding an auxiliary classification layer to the original discriminator, which utilizes category information. This model can learn the mutual transformation between the visual and tactile domains.

[0115] Existing Method 3: Literature

[0116] The paper "Learning cross-modal visual-tactile representation using ensembled generative adversarial networks" (authors: Li X, Liu H, Zhou J, et al.) is based on deep learning. It takes the texture attributes of an image as input, obtains the image's category information, and uses it as the selection condition for a GAN. Thus, cross-modal conversion from texture images to tactile signals is achieved through the training of the corresponding GAN.

[0117] This invention: The method of this embodiment.

[0118] The similarity (SIM) is used as an evaluation index to evaluate the cross-modal generation effect. The SIM is used to measure the similarity between the real haptic signal and the reconstructed haptic signal, and is defined as:

[0119] wherein the index in the denominator h i ,h i ′ is the Chebyshev distance between h i and h i ′ waveforms. That is, the maximum value of the distance between two waveform points. The SIM ranges from [0, 1]. When the difference between the two is small, the SIM is closer to 1, and the two waveforms are more similar, and vice versa.

[0120] Table 1 is the experimental results of the present application

[0121] As can be seen from Table 1 and Figure 5, compared with the above-mentioned state-of-the-art method, the method proposed by us has obvious advantages. All other comparative model inputs are only visual signals, while our model fuses visual and haptic signals to extract high-level semantic information of the same object from two modalities, better extracts common information between modalities and unique information within modalities, and thus better restores the haptic signal. Figure 5 shows the haptic reconstruction results of the comparative methods, and Figure 6 shows the haptic spectrum and sequence reconstruction results of the method of the present application.

[0122] Secondly, in order to further verify the adaptability of the method of the present application to the signal-to-noise ratio, the method of the present application and the comparative method are compared in terms of haptic reconstruction effect under different signal-to-noise ratios. Here, the comparative method uses a self-encoder-decoder structure composed of multiple linear layers and convolutional layers instead of the cross-modal joint encoder-decoder structure in the present application. Then, experiments are carried out under different test signal-to-noise ratios, as shown in Figure 7. Under different signal-to-noise ratios, the method of the present application can reconstruct the haptic signal with high quality, showing the self-adaptive characteristics to the change of the signal-to-noise ratio, while the haptic reconstruction effect of the comparative method under low signal-to-noise ratio is not ideal.

[0123] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A method for haptic signal reconstruction of a signal-to-noise ratio adaptive joint optimization of a source channel, characterized in that, The method comprises the following steps: Preprocessing the image signal corresponding to each object and the tactile vibration signal generated by sliding based on a tactile-image database; Inputting the preprocessed image and tactile signal into a cross-modal joint encoder, performing block processing on the input data, sequentially inputting a hierarchical feature extraction module and a signal-to-noise ratio adaptive module, and obtaining hierarchical fusion features; After the hierarchical fusion features are transmitted through a channel, they enter a cross-modal joint decoder, which comprises a hierarchical feature extraction module, a signal-to-noise ratio adaptive module and a signal reconstruction module; The hierarchical feature extraction module decodes the tactile low-dimensional features from the hierarchical fusion features, the signal-to-noise ratio adaptive module dynamically adjusts the low-dimensional features according to the channel changes, and the signal reconstruction module converts the low-dimensional features into a tactile frequency spectrum, which is converted into a tactile sequence through inverse short-time Fourier transform (ISTFT), thereby completing the reconstruction of the tactile signal; Joint training of the cross-modal joint encoder and the cross-modal joint decoder is performed according to the training data to obtain optimal model parameters; The optimal cross-modal joint encoder is inputted with the to-be-tested tactile and image signals to obtain hierarchical fusion features, which are transmitted through a channel, and then the received hierarchical fusion features are inputted into the cross-modal joint decoder to generate a target tactile frequency spectrum, which is converted into a target tactile signal through ISTFT.

2. The SNR adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 1, wherein: The preprocessing includes cutting the image according to HxW size, denoted as: wherein v j represents the jth picture in the data set with size HxW in RGB three channels; The haptic signal is pre-processed, and a short-time Fourier transform (STFT) is performed on the haptic sequence to obtain a frequency spectrum of the haptic sequence, denoted as: passing the spectrum through and obtaining the amplitude spectrum and phase spectrum Splicing amplitude spectrum and phase spectrum to obtain a haptic signal RGB image V and haptic signal H are concatenated in the channel dimension into five-dimensional data Wherein, n = H x W x 5 represents the size of the image-tactile signal pair, and the pairing of the image and tactile signal is maintained during the processing.

3. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 2, wherein: The block processing comprises segmenting the data pair into non-overlapping blocks of 4 x 4 size; The feature dimension of each patch is 4 x 4 x 5 = 80 dimensions, and the size of the data tensor is The segmented data is inputted into a plurality of hierarchical feature extraction modules and signal-to-noise ratio adaptive modules; The hierarchical feature extraction module is FE e The signal-to-noise ratio adaptive module is SD e (·), respectively connected after the first and last feature extraction components; Let the number of feature extraction components in the encoder be L, the input of the jth feature extraction component be t j j e {1, 2,..., L}, and the continuous calculation process is represented as: The haptic-image signal pair d e D is divided into non-overlapping patches in both the long and the wide direction to obtain t1, which is input to the hierarchical feature extraction module to obtain the first layer feature representation The first layer feature representation is inputted into the signal-to-noise ratio adaptive module to obtain t2, which is inputted into a plurality of feature extraction components, and the deep layer feature is obtained through the last layer feature extraction component and inputted into the last layer signal-to-noise ratio adaptive layer to obtain the remapped feature e2.

4. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 3, wherein: The hierarchical feature extraction module FE e (·) includes a merging layer and a Swin Transformer module, learning the hierarchical features of the tactile-image signal; The merging layer is responsible for downsampling and dimensionality increase to generate features at different levels, and the Swin Transformer block performs hierarchical feature learning; The first merging layer connects the features of each group of 2 x 2 adjacent patches in the channel dimension, makes the feature resolution downsampled by 2 times, and increases the feature dimension by 4 times; The full connection layer is applied on the connected features, and the output dimension is unified to twice of the original dimension, i.e. 2C, and the tensor size is After passing through the Swin Transformer block and the merging layer multiple times, the hierarchical feature representation is obtained; The last layer merging layer of the cross-modal joint encoder is responsible for mapping the deep feature information to the same dimension as the channel channel size. After connecting the adjacent patch features, a linear layer is applied to transform the output dimension to the channel dimension b. At this time, the tensor size is 5. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 4, wherein: The signal-to-noise ratio adaptive module SD(·) comprises global pooling and a feedforward network; The global pooling is used for capturing full-size information of input features, integrating local features into global features, connecting signal-to-noise ratio information with the pooled features, inputting a feedforward network to generate a scaling factor, and multiplying the output features of the hierarchical feature extraction module FE e of the application (·) to obtain weighted features. The output of the feature extraction component is o = FE(t) e R H×W×c , o = {o1, o2,..., o c} ipq denotes the element at position (p, q) of the i-th channel. After global average pooling P(·), each feature map is pooled in the channel dimension to extract global information The global pooling operation sums up each two-dimensional feature map o i into a single numerical value s i representing the average activation strength of the entire feature map, denoted as: The global information is concatenated with the current SNR value to obtain It is input into a feedforward network, denoted as: Will The original input feature o∈R H×W×c The multiplication obtains the feature e∈R adjusted according to the channel condition H×W×c Since the sizes of the two are different, before multiplication, the data of each channel of the factor is copied, so that the size is expanded to H×W×c.

6. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 5, wherein: The cross-modal joint decoder uses an expansion layer to upsample the extracted deep layer feature, the expansion layer reshapes the feature maps of adjacent dimensions and reduces the feature dimension to half of the original dimension, thereby completing the dimensionality reduction of the feature; The channel signal is inputted into a plurality of hierarchical feature extraction modules to obtain features at different levels, and the signal-to-noise ratio adaptive module is connected after the first and last hierarchical feature extraction modules; Let the number of hierarchical feature extraction modules in the cross-modal joint decoder be L, and the input of the jth hierarchical feature extraction module according to the data flow direction be t j j∈{1,2,...,L}, and the continuous calculation process is represented as: Signal transmitted from a channel The input hierarchical feature extraction module obtains a first layer reduced dimension feature representation The input signal-to-noise ratio adaptive module obtains t2, which is reduced in dimension by multiple hierarchical feature extraction modules, and then again passes through a signal-to-noise ratio adaptive layer to perform feature scaling on the shallow features after the last dimension reduction, to obtain the features e2 that are remapped, and through an expansion layer, the features are scaled by 4 times to convert the feature size to The reduced dimensionality features are input to a signal reconstruction module that maps through multiple convolutional layers from the feature domain to a time-frequency matrix containing amplitude and phase information of the haptic sequence, denoted as: The multi-level features are inputted into a plurality of convolutions, the size is kept unchanged, the channel number is scaled, the channel number is scaled to 2, and the channel number is mapped to the amplitude and phase information of the time-frequency matrix of the tactile sequence, and the recovered tactile sequence is obtained through ISTFT transformation.

7. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 6, wherein: The performing the joint training includes selecting a Charbonnier loss as a loss function of an optimization process, denoted as: wherein h represents the haptic signal recovered by the decoder j ε represents the real haptic signal; ε = 10 -3 is a constant; The cross-modal joint encoder and the cross-modal joint decoder are jointly trained as a whole network, gradients are calculated through an Adam optimization algorithm, and network parameters are updated through back propagation; and the expression is as follows: where m t and v t are the unbiased estimates of the first and second moments, respectively, g t is the gradient of the parameters, beta1 and beta2 are the decay rates of the first and second moment estimates, respectively, μ is the learning rate, ε is a very small constant to ensure the denominator is not zero, θ t and θ t+1 are the network parameters before and after updating, respectively, where the network parameters refer to the network parameters of the cross-modal joint encoder and the cross-modal joint decoder, the learning rate is selected as 1×10-4, the number of iterations is 1000, and the network parameters are determined after the iteration is completed to complete the model training.

8. The SNR-adaptive joint optimization of source-channel for haptic signal reconstruction method of claim 7, wherein: The reconstructed target haptic signal comprises a hierarchical fusion feature generated by a cross-modal joint encoder, and after channel transmission, the cross-modal joint decoder is used for decoding to reconstruct a spectrum of the target haptic signal, and the expected haptic signal is obtained through ISTFT transformation. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor implements the steps of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of the signal source and the channel according to any one of claims 1 to 8 when the computer program is executed.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the haptic signal reconstruction method of the signal-to-noise ratio adaptive joint optimization of the signal source and the channel according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Mechanical arm autonomous operation strategy learning method based on vision-touch fusion

    CN114660934A

  • 6G-oriented tactile modal signal reconstruction method

    CN114842384A

  • Cross-modal communication-oriented image super-resolution reconstruction method

    CN115936997A

  • Audio-visual assisted fine-grained tactile signal reconstruction method

    WO2024104376A1

Cited By

  • Industrial Internet of Things semantic data transmission method and system based on feature enhancement

    CN122179063A