A method for reconstructing fine-grained tactile signals for audiovisual assistance

By employing cluster analysis and cross-modal transfer learning with shared semantic constraints, the method addresses the limitations of existing haptic signal reconstruction, achieving high-quality, fine-grained tactile signal generation for enhanced cross-modal communication.

JP2025515925AActive Publication Date: 2025-05-20NANJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024568338
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-18
Filing Date
2023-11-15
Publication Date
2025-05-20
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

Existing methods for reconstructing haptic signals in cross-modal communication are inadequate, particularly in noisy environments, and fail to generate high-quality, fine-grained tactile sensations due to reliance on single-modal information, uneven data distribution, and lack of direct correspondence between modalities, leading to inaccurate and incomplete tactile signal generation.

Method used

A method involving cluster analysis, cross-modal transfer learning, and shared semantic learning to constrain and optimize haptic, audio, and image features, using clustering, centrality, and sorting constraints to achieve fine-grained classification and reconstruction of haptic signals.

Benefits of technology

The method effectively generates high-quality, fine-grained haptic signals by leveraging multi-modal feature fusion, addressing the limitations of existing methods and enhancing the quality of tactile perception in cross-modal services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515925000001_ABST
    Figure 2025515925000001_ABST
Patent Text Reader

Abstract

The present invention discloses a method for reconstructing fine-grained haptic signals for audiovisual assistance, which first passes the training haptic signal through a haptic autoencoder, extracts haptic features based on a clustering task, and transfers the features to an audio-visual feature extraction network to realize feature extraction of the audio-visual signal; then optimizes a multi-modal fusion mapping function of haptic, audio, and image using triplet constraints, thereby obtaining fusion features from the extracted audio-visual features; and finally, inputs the fusion features into a haptic generation network to realize the reconstruction of fine-grained haptic signals. The present invention better solves the problems of weak supervision and weak matching existing between multi-modal signals, realizes cross-modal shared semantic learning, ensures the structure and semantic completeness of the generated haptic signal, and also enhances its clustering properties, thereby significantly improving the reconstruction quality of the haptic signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the technical field of generating haptic signals, in particular to a method for reconstructing fine-grained haptic signals for audiovisual aids. [Background technology]

[0002] As the related technologies of traditional multimedia applications mature, people's audiovisual needs are largely satisfied, and they begin to pursue more dimensions and higher levels of perception experience. It is expected that haptic information will gradually be integrated into existing audio / video multimedia services to form multimodal services, bringing richer interactive experiences. Cross-modal communication technology has been proposed to support cross-modal services, and has a certain effectiveness in ensuring the quality of multimodal streams. However, when cross-modal communication is applied to multimodal services that mainly use haptics, some technical challenges still remain. First, the haptic stream is very sensitive to interference and noise in the wireless link, which results in the degradation and even loss of the haptic signal at the receiving end. This problem is serious and unavoidable, especially in the application scenarios of remote operation, such as remote industrial control and remote surgery. Secondly, service providers do not have tactile collection devices, but users need to sense touch, and especially in virtual interactive application scenes such as online immersive shopping, holographic museum guides, virtual interactive movies, etc., users have a very high need for tactile perception, and therefore require the ability to generate "virtual" touch sensations or tactile signals based on video and audio signals.

[0003] Currently, haptic signals that are corrupted or partially missing due to unreliable wireless communication and communication noise interference can be self-recovered in two ways. The first is based on traditional signal processing techniques, which use sparse representation to search for a specific signal with the most similar structure, and then use the specific signal to estimate the missing part of the corrupted signal. The second is to mine and utilize the spatio-temporal correlation of the signal itself to realize intra-modal self-recovery and reconstruction. However, when the haptic signal is severely corrupted or even non-existent, intra-modal based reconstruction schemes fail.

[0004] In recent years, some studies have focused on the correlation between different modalities and realized cross-modal reconstruction. Li et al. in "Learning cross-modal visual-tactile representation using ensemble generative adversarial networks" proposed to use image features to obtain the required category information, and then use this information together with noise as the input of an adversarial network to generate a tactile spectrum map of the corresponding category. In this method, the information obtained by mining the semantic correlation and categories between each modality is limited, so the results generated are often inaccurate. Kuniyuki Takahashi et al. in "Deep Visuo-Tactile Learning: Estimation of Tactile Properties from Images" extend a single encoder-decoder network to embed both visual and tactile attributes into a latent space, focusing on the degree of tactile properties of materials represented by latent variables. Furthermore, in the paper "Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces From Images," Matthew Purr et al. proposed a cross-modal learning framework with adversarial learning and cross-domain associative classification to estimate tactile physical properties from a single image. Although such methods exploit the semantic information of the modalities, they do not generate a complete tactile signal and therefore have no practical meaning for cross-modal services.

[0005] The above existing cross-modal generation methods also have the following defects: The training of the models all relies on large-scale training data to ensure the effectiveness of the models, and they all only use single-modal information, but in reality, the advantage of single modality cannot bring a sufficient amount of information, and when different modalities jointly describe the same meaning, they may contain an uneven amount of information, and the complementation and enhancement of information between modalities contributes to improving the generation effect. In actual application scenarios, the annotation cost of large-scale data sets is huge, and it is often easier to obtain coarse-grained large classification categories, and fine-grained categories are not clear. In addition, there is no direct correspondence between samples of different modalities, and there are challenges in weak supervision and weak matching. Summary of the Invention [Problem to be solved by the invention]

[0006] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a method for reconstructing fine-grained haptic signals for audiovisual assistance. First, a cluster analysis is performed within the coarse-grained categories to obtain a fine-grained classification of samples. Second, a modal co-semantic constraint is performed in the fine-grained categories, the purpose of which is to minimize the intra-category dissimilarity and maximize the inter-category dissimilarity. Finally, a correlation constraint is performed by searching for audiovisual samples that positively match the haptic signal in the fine-grained subcategories, thereby realizing a high-quality, fine-grained reconstruction of the haptic signal. [Means for solving the problem]

[0007] In order to solve the above technical problems, the present invention provides a method for reconstructing fine-grained haptic signals for audiovisual assistance, A step of inputting a haptic signal to a haptic autoencoder and extracting features of the haptic signal through a clustering task; Transferring and optimizing the feature extraction ability of the haptic autoencoder to the audio feature extraction network and the image feature extraction network by a cross-modal transfer learning method; constraining the extracted haptic, audio and image modal features by jointly considering clustering constraints, centrality constraints and sorting constraints to bring modal features belonging to the same meaning close to each other, but separate modal features not belonging to the same meaning, to obtain haptic, audio and image modal features with fine-grained classification; Constructing a triplet set based on haptic, audio and image modal features, performing shared semantic learning of triplet constraints, optimizing a multi-modal fusion mapping function, and obtaining a fusion feature containing shared semantic information; and preconfiguring a haptic generation network and inputting the fused features into the haptic generation network to reconstruct a haptic signal.

[0008] Further, the step of preconfiguring a haptic generative network and inputting the fused features to the haptic generative network to reconstruct a haptic signal includes: Presetting parameters of a haptic autoencoder, an audio feature extraction network, and an image feature extraction network; Presetting parameters of a multimodal fusion mapping function and parameters of a haptic generation network; training a haptic autoencoder, an audio feature extraction network, an image feature extraction network, a multimodal fusion mapping function, and a haptic generation network; The method includes inputting the just-received image signal and audio signal into a trained image feature extraction network and an audio feature network, respectively, to obtain image features and audio features, respectively, then inputting the features into a multi-modal fusion mapping function to obtain fusion features, and finally inputting the fusion features into a trained haptic generation network to obtain a reconstructed haptic signal.

[0009] Furthermore, inputting the haptic signal into a haptic autoencoder and extracting features from the haptic signal through a clustering task is The haptic signal is input to a haptic autoencoder for learning, and the corresponding haptic features are extracted. Then, the haptic signal is clustered based on the K-means algorithm according to the haptic features, i.e. h is the input haptic signal, h = {h i} i=1,…,N where i denotes the sort subscript of the input haptic signal, N is the total amount of input haptic signals, and the encoding module E h After passing through (·), i h =E h (h i ;θ he ) is the haptic signal h i is a feature representation of f h ={f i h} i=1,…,N and θ he are the parameters of the encoding module, and f i h Decryption module D h (·) input and output haptic signal

number

number

number

number

number

[0010] Furthermore, the feature extraction ability of the haptic autoencoder can be transferred to an audio feature extraction network and an image feature extraction network by a cross-modal transfer learning method. The feature self-adaptation method realizes the transfer by minimizing the maximum average difference criterion between the tactile and audiovisual domains, i.e. Let P, Q, and R be the distributions of the haptic, audio, and image signal sets, respectively, and let MMD be the MMD between the haptic and audio signals. k (P,Q) and the MMD between the tactile and visual signals is MMD k (P,R) and the reproducing kernel Hilbert space H k In H kcontains a set of functions f defined on a non-empty set, and the square of the MMD is

number

number

number

number

number

number

[0011] Further optimizing the audio feature extraction network and the image feature extraction network further comprises: The classification loss function is

number

[0012] Furthermore, by constraining the extracted tactile, audio, and image modal features by jointly considering the clustering constraint, the centrality constraint, and the sorting constraint, each modal feature belonging to the same meaning is brought close to each other, but each modal feature not belonging to the same meaning is separated, and tactile, audio, and image modal features with fine-grained classification are obtained. In order to ensure compactness between features of the same fine-grained subcategory, clustering learning under center constraint is performed on three types of modal signals. In order to achieve better fine-grained classification performance, features of the same subcategory should be adjacent in a common space, with the objective of minimizing the intra-category variance. Clustering learning is driven by minimizing the distance from a feature to its subcategory center. The subcategory center of the tactile signal is set as the common subcategory center of the cross-modal signal to ensure compactness between semantic features of the same category of cross-modal signals. The loss function of clustering learning under center constraint is:

number

[0013] Furthermore, to ensure that the features of different fine-grained subcategories have a certain sparsity, we perform clustering learning under sorting constraints on the three types of modal signals. The goal of the centering constraint is to minimize the within-category variance, while the goal of the sorting constraint is to maximize the between-category variance, so that the feature outputs of different subcategories are less similar than the feature outputs of the same subcategory. The sorting constraint is:

number

[0014] Furthermore, constructing a triplet set based on the modal features of tactile, audio and image, performing shared semantic learning of triplet constraints, optimizing a multi-modal fusion mapping function, and obtaining a fusion feature containing shared semantic information is One haptic signal sample from a fine-grained subcategory, h i randomly selecting a sample as an anchor; Image dataset from h i belongs to the same category as f and has semantic features i h The closest sample to v i + Select h as the positive match sample, i does not belong to the same category as f and has semantic features i h The closest sample to v j - Selecting as a negative match sample; This allows us to create a triplet set {(h i ,v i + ,v j - )}; and Similarly, the triplet set {(h i ,a i + ,a j - )}, Anchor point h i Semantic features of f i hand positive match semantic features within the corresponding subclasses in the audio-visual modalities.

number

number

number

number

[0015] Furthermore, we construct a haptic generation network, i.e., a haptic generation network G(·), whose structure is D h (·) and its network parameters θ hd The parameter θ of G(·) G Let be the initial value of The fused features containing the necessary semantic information are input to the haptic generation network G(·) to obtain the desired haptic signal h′, and the generated haptic signal h′ is expressed as E h (·) gives the tactile feature f h′ , and select the category center to represent the tactile feature f h′ We apply semantic constraints to the,final loss function,

number

number

number

number

[0016] Furthermore, we pre-set the parameters of the haptic autoencoder, the audio feature extraction network, and the image feature extraction network. Presetting parameters of the multimodal fusion mapping function and parameters of the haptic generation network; Training the haptic autoencoder, the audio feature extraction network, the image feature extraction network, the multimodal fusion mapping function, and the haptic generation network includes step 1 and step 2; In step 1, θ v , θ a , θ he , θ hd , and M are preset, and {s i h}, {s i a}, {s i v}, and the step 1 is Network parameter θ v , θ a , θ he , θ hd , node tag {s i h}, {s i a}, {s i v} and the category center matrix M are initialized, and the number of clusters C and the learning rate μ 1 11, setting the number of iterations T; {s i h}, {s i a}, {s i v} and M are fixed, and θ is calculated based on the stochastic gradient descent method. v , θ a , θ he , θ hd Optimize, i.e.,

Number

Number

Number

number

number

[0017] Compared with the prior art, the present invention has the following technical advantages by using the above technical solutions:

[0018] The present invention uses the deep clustering algorithm of cross-modal transfer to learn fine-grained classification of three kinds of modal samples, and then performs shared semantic learning, fully exerting the advantages of multi-modal feature fusion, and finally realizing the generation of fine-grained haptic signals based on clustering constraints, making full use of the existing dataset with weak supervision and weak matching problems, thereby generating high-quality fine-grained haptic signals, which better meets the requirements of cross-modal services. [Brief description of the drawings]

[0019] [Figure 1] 1 is a flowchart of a method for reconstructing fine-grained haptic signals for audiovisual assistance according to the present invention. [Diagram 2] FIG. 2 is a structural schematic diagram of a complete network according to the present invention; [Diagram 3] FIG. 1 is a schematic diagram of the architecture of a deep clustering model based on cross-modal transfer according to the present invention. [Figure 4] 13A-13C show tactile signal reconstruction results of the present invention and other comparison methods. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0021] The present invention provides a method for reconstructing fine-grained haptic signals for audiovisual assistance, the flow chart of which is shown in FIG. 1, and the method includes the following steps 1 to 5.

[0022] In step 1, first, the haptic signal is input into the haptic autoencoder, and feature extraction of the haptic signal is realized by clustering. Then, the feature extraction ability of the haptic autoencoder is transferred and optimized to the audio feature extraction network and the image feature extraction network by using cross-modal transfer learning technology. Then, the extracted haptic, audio, and image modal features are further constrained by jointly considering the clustering constraint, the centrality constraint, and the sorting constraint, so that the modal features belonging to the same meaning are close to each other, but the modal features that do not belong to the same meaning are separated, and the haptic, audio, and image modal features with fine-grained classification are obtained.

[0023] (1-1) First, feature learning with clustering constraints is performed on three types of modal signals, namely, haptic, image, and audio, to obtain segment features with fine-grained subcategories, which may be divided into the following three steps, namely, step (1-1-1) to step (1-1-3).

[0024] (1-1-1) In the first step, the haptic signal is input to the autoencoder for learning, the corresponding haptic features are extracted, and the haptic signal is clustered based on the K-means algorithm according to the haptic features, i.e., Specifically, i is the input haptic signal, and h={h i} i=1,···,N where i denotes the sort subscript of the input haptic signal, N is the total amount of input haptic signals, and the encoding module E h After passing through (·), i h =E h (h i ;θ he ) is the haptic signal h i is a feature representation of f h ={f i h} i=1,···,N and θ he Let f be the parameters of the encoding module. i h Decryption module D h (·) input and output haptic signal

number

number

number

number

number

[0025] (1-1-2) In the second step, the acquired haptic features are transferred to the feature extraction process of the image and audio signals using cross-modal transfer learning technology. That is, the transfer is realized by minimizing the maximum mean difference (MMD) criterion between the haptic domain and the audiovisual domain using a feature self-adaptation method.

[0026] Specifically, the distributions of the haptic, audio, and image signal sets are P, Q, and R, respectively. The MMD between the haptic signal and the audio signal is MMD k (P,Q) and the MMD between the tactile and visual signals is MMD k (P,R). The reproducing kernel Hilbert space H k In H k contains a set of functions f defined on a non-empty set, and the square of the MMD is

number

number

number

number

[0027] handle

number

number

[0028] L CT By optimizing θ, we can guide the information flow between the haptic feature extraction autoencoder model and the audio-visual feature extraction network, thereby effectively transferring the feature extraction ability of the autoencoder for the haptic modal to the audio-visual feature extraction network, i.e., θ a and θ v can be estimated.

[0029] (1-1-3), the third step is to further optimize the audio-visual feature extraction network, where the video feature extraction network is selected from the design style of the VGG network, i.e., it has a 3×3 convolution filter and a 2×2 max pooling layer without padding with a step width of 2, the network is divided into 4 blocks, each block contains 2 convolution layers and 1 pooling layer, with the number of filters doubling between successive blocks, and finally, max pooling is performed on all spatial positions to generate a single 512-dimensional semantic feature vector. Then, the semantic feature vector is input into one 3-layer fully connected neural network (256-128-32) and one fully connected layer with K nodes and softmax function, the 32-dimensional vector is the visual signal feature vector, which is further trained for subsequent triplet constraint-based shared semantic learning, where K is the number of fine-grained subcategories in each coarse-grained category, and K is 3 in this experimental dataset. The associated network structure for audio signals is similar in setup to visual signals and shares classifiers with visual signals.

[0030] The designed classification loss function is

number

[0031] It is particularly noted that the tag acquisition adopts the following incremental policy: first, pseudo tags are attached to the data, and the data are used to optimize the model, so that the optimized model's ability to classify audiovisual signals is enhanced; then, the trained model is used to perform pseudo tag operations, so that the pseudo tags are updated. Through such an incremental optimization method, the fine-grained classification ability of the network is gradually improved.

[0032] In short, the total objective function in step (1-1) is a combination of the three loss functions above, L clu =L clu h +L CT +L clu av may be shown as L clu By minimizing h , θ a , and θ v After determining the parameters, the features f corresponding to the haptic signal, the audio signal, and the image signal can be obtained. h , f a , f v You can get (1-2) To ensure compactness between features of the same fine-grained subcategory, clustering learning under center constraint is performed on the three types of modal signals. To achieve better fine-grained classification performance, features of the same subcategory should be adjacent in a common space, and the objective is to minimize intra-category variance. Specifically, clustering learning is driven by minimizing the distance from a feature to its subcategory center. The subcategory center of the tactile signal is set as the common subcategory center of the cross-modal signal to ensure compactness between cross-modal signal semantic features of the same category. The loss function of clustering learning under center constraint is

number

[0033] (1-3) In order to ensure that the features of different fine-grained subcategories have a certain sparsity, we perform clustering learning under sorting constraints on the three types of modal data. The goal of the centrality constraint is to minimize the within-category variance, while the goal of the sorting constraint is to maximize the between-category variance, so that the feature outputs of different subcategories are less similar than the feature outputs of the same subcategory. The sorting constraint is:

number

[0034] In step 2, after obtaining tactile, audio, and image features with fine-grained classification, the three kinds of modal features are constructed into a triplet set, triplet constraint shared semantic learning is carried out, and the multimodal fusion mapping function is optimized to obtain the fusion features containing shared semantic information, which lays the foundation for the generation of tactile signals.

[0035] Specifically, step 2 is as follows:

[0036] (2-1), one haptic signal sample from a fine-grained subcategory, h i Randomly select a sample as an anchor, and then select h from the image dataset. i belongs to the same category as f and has semantic features i h The closest sample to v i + Select h as a positive match sample, and then i does not belong to the same category as f and has semantic features i h The closest sample to v j - As a negative match sample, we select the triplet set {(h i ,v i+ ,v j - Similarly, a triplet set {(h i ,a i + ,a j - )} can be obtained. Anchor point h i Semantic features of f i h and positive match semantic features within the corresponding subclasses in the audio-visual modalities.

number

number

number

[0037] (2-2) Based on the above, we introduce an integration paradigm to highly fuse multimodal features. Specifically, audiovisual data is first passed through an audio feature extraction network and an image feature extraction network, respectively, to obtain features f a and f v Then, f a and f v The process is as follows: f m =F m (f a ,f v ;θ m ) (However, f m is the output of multimodal fusion in the shared semantic subspace, i.e., the fusion feature, and F m (·) is the parameter θ mis a mapping function for F m (·) is f a and f v (linearly weighted by ).

[0038] f + a and f + v f + m Fusing to f - a and f ― v f ― m Similarly, triplet loss is used to constrain the fused features, i.e.

number

[0039] (2-3), the objective function of shared semantic learning may be modeled by combining three loss functions: L Syn =L syn a +L syn v +L syn m It may be shown as:

[0040] L Syn By minimizing m thereby laying the foundation for the next stage of tactile signal generation.

[0041] In step 3, the fused features are input into a haptic generation network to generate the desired haptic signal h′.

[0042] Specifically, step 3 is as follows:

[0043] First, we constructed a haptic generation network and a haptic decoder D hSince (·) is trained as part of the complete autoencoder in step (1), we now construct a separate haptic generation network G(·) whose structure is D h (·) (32-128-256-Z) and its network parameters θ hd The parameter θ of G(·) G The fusion feature including the necessary semantic information is input to the haptic generation network G(·) to obtain the desired haptic signal h′, and the generated haptic signal h′ is expressed as E h (·) gives a 32-dimensional tactile feature f h′ , and select the category center to map the tactile feature f h′ Obviously, we apply semantic constraints to the tactile features f h′ The distance between and the category center of the corresponding category is as small as possible, but the distance between and the center of other categories is as large as possible. The final loss function is

number

number

number

number

[0044] In step 4, the model is trained, that is, the model training is divided into two steps: the first step is to estimate the parameters in the haptic autoencoder, the audio feature extraction network, and the image feature extraction network, and the second step is to estimate the parameters of the multimodal fusion mapping function and the haptic generation network, through which the haptic autoencoder, the audio feature extraction network, the image feature extraction network, the multimodal fusion mapping function, and the haptic generation network are trained.

[0045] Specifically, step 4 is as follows:

[0046] θ v , θ a , θ he , θ hd , and M are estimated, and {s i h}, {s i a}, {s i v} to optimize.

[0047] Step 411, network parameter θ v , θ a , θ he , θ hd , node tag {s i h}, {s i a}, {s i v} and initialize the category center matrix M, set the number of clusters C, and set the learning rate μ 1 = 0.0001, and the number of iterations T = 600.

[0048] Step 412, {s i h}, {s i a}, {s i v} and M are fixed, and θ v , θ a , θ he , θ hdOptimize, that is,

Number

[0049] Step 413, Fix θ v , θ a , θ he , θ hd , and M, and optimize {s i h}, {s i a}, {s i v}, that is,

Number

[0050] Step 414, Fix θ v , θ a , θ he , θ hd , and {s i h}, {s i a}, {s i v}, and optimize M, that is,

Number

[0051] Step 415, If t < T, jump to Step 412, if t = t + 1, continue with the next iteration, otherwise, end the iteration.

[0052] Step 416, After T iterations, the parameters θ of the optimal audio feature extraction network a , the parameters θ of the image feature extraction network v , the parameters θ of the tactile autoencoder he , θ hd , the node - based tags {s i h}, {si a}, {s i v}, and obtain the cluster center vector matrix M of the tactile data.

[0053] In particular, m k When updating

number

[0054] (4-2), based on SGD, θ m and θ G Complete the estimation of.

[0055] Step 421, θ m Initialize with batch size bactch=64 and learning rate μ 2 , μ 3 =0.0001, iteration count n 1 =600.

[0056] Step 422, L Syn Based on this, we use stochastic gradient descent to find θ m Fine-tune the

number

[0057] Step 423, L Gen Based on this, we use stochastic gradient descent to find θ G Update, i.e.,

number

[0058] Step 424, n <n 1 If so, jump to step 422; if n=n+1, continue to next iteration; otherwise, end iteration.

[0059] Step 425, after 600 iterations, obtain the optimal n=n+1.

[0060] In step 5, the just-received image signal and audio signal are input into the trained image feature extraction network and audio feature network respectively to obtain image features and audio features respectively, then the above features are input into the multi-modal fusion mapping function to obtain fusion features, and finally, the fusion features are input into the trained haptic generation network to obtain the reconstructed haptic signal.

[0061] Specifically, step 5 is as follows:

[0062] The just received image signal v and audio signal a are input to the trained image feature extraction network and audio feature extraction network, respectively, to extract image features.

number

number

number

number

number

[0063] As can be seen from the following experimental results, compared with the conventional methods, the present invention realizes the synthesis of tactile signals through the complementary fusion of multi-modal meanings, and achieves a higher generation effect.

[0064] This embodiment uses the LMT cross-modal dataset for experiments, which is proposed in the paper "Multimodal feature-based surface material classification" and includes samples of nine semantic categories, namely, grid, stone, metal, wood, rubber, fiber, foam, foil, and paper, textile products, and fabrics. This embodiment selects five major categories (each of which contains three sub-categories) for experiments. The LMT dataset is reconstructed, and first, the training set and test set of each material example are referred to, and 20 image samples, 20 audio signal samples, and 20 tactile signal samples are obtained for each example, respectively. Then, the neural network is trained by expanding the data, specifically, each image is flipped horizontally and vertically, rotated at any angle, and techniques such as random scaling, cutting, and offset are used in addition to the conventional method. With this, the data of each category is expanded to 100, so that there are a total of 1500 images, with a dimension of 224×1024×1024. * In the dataset, 80% is selected to be used for training, while the remaining 20% ​​is used for testing and performance evaluation. The experiment is initially set up so that the fine-grained categories are unknown.

[0065] (1) Clustering results To verify the effectiveness of the clustering method of the present invention, the clustering method is compared with several baseline methods, including:

[0066] K-means (KM) The K-Means algorithm is used to cluster the samples of image, audio, and haptic modalities, respectively.

[0067] Autoencoder + K-means (AE+KM, Autoencoder followed by K-means) This is a two-step method: first, we obtain feature representations for each modality by reconstructing and learning signal samples of different modalities, and then cluster them using K-means.

[0068] Triple deep clustering model (3-DCN, 3-Deep Clustering Network) Clustering is performed using DCN models for signals of different modalities.

[0069] The present invention uses the method of this embodiment.

[0070] We select the average values ​​in multiple subcategories to present the results, and mainly adopt three indices: normalized mutual information (NMI), adjusted Rand index (ARI), and clustering accuracy (ACC). The experimental results are shown in Table 1.

[0071] [Table 1]

[0072] Table 1 shows the results of applying the present invention, 3-DCN, AE+KM, and KM to the LMT dataset. As can be seen, the method of the present invention is extremely competitive in this dataset, and the results are obviously better than traditional clustering algorithms and common deep clustering algorithms. Analysis shows that this may be because the application scenes of other algorithms are all unimodal, which is likely to lead to imbalance in clustering results. Theoretically, the number of samples in a category of coexisting cross-modal data should be equal. In addition, the features of subcategories learned by the present clustering method are obviously more compact, and the discrimination between different categories is also higher, which is beneficial for the subsequent reconstruction of tactile signals.

[0073] (2) Results of tactile reconstruction After determining the fine-grained category, we compare the proposed fine-grained tactile reconstruction method with several other methods as follows.

[0074] Existing method 1 The ensemble generative adversarial networks (E-GANs, Ensembled GANs) in the paper "Learning cross-modal visual-tactile representation using ensembled generative adversarial networks" (authors X. Li, H. Liu, J. Zhou, and F. Sun) uses image features to obtain the required category information, and then uses the category information together with noise as input to the generative adversarial network to generate a tactile spectrum map of the corresponding category, and finally converts it into a tactile signal.

[0075] Existing method 2 The deep vision-tactile learning method (DVTL, deep vision-tactile learning) in the paper "Deep Visuo-Tactile Learning: Estimation of Tactile Properties from Images" (authors: Kuniyuki Takahashi and Jethro Tan) extends the traditional encoder-decoder network with latent variables to embed visual and tactile attributes into the latent space.

[0076] Existing method 3 The paper "Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces From Images" (authors: Matthew Purri and Kristin Dana) proposes a joint-encoding-classification GAN (JEC-GAN), which uses different encoding networks to encode the examples of each modality into a shared internal space, and then approximates the embedded visual and haptic samples in the latent space through paired constraints. Finally, the visual information is used as input to reconstruct the corresponding haptic signal through a generative network.

[0077] This experiment is analyzed from two perspectives: quantitative and qualitative. First, Table 2 shows the tactile signal reconstruction performance of each method from multiple perspectives, including root mean square error (RMSE), structural similarity (SIM), and classification accuracy (ACC). Table 2 shows the experimental results of the present invention.

[0078] [Table 2]

[0079] As can be seen from Table 2 and Figure 4, compared with the most advanced methods mentioned above, the method according to the present invention has obvious advantages. The reasons are as follows: (1) the haptic signal reconstruction method according to the present invention clarifies fine-grained subcategories by cross-modal clustering algorithm, and effectively improves the compactness and distinctiveness of semantic features by using center constraint and sorting constraint; (2) when the correspondence between visual-auditory-haptic samples is specified by humans, the accuracy is not sufficient due to the relatively strong subjective consciousness; on the contrary, in the fine-grained haptic signal reconstruction method for audiovisual assistance according to the present invention, the model self-selects the audiovisual semantic features that are closest to the haptic semantic features as the input of its generation network during training.

[0080] In another embodiment, the haptic encoder in step 1 of the present invention may be replaced by one-dimensional convolutional neural networks (1D-CNN) using a feedforward neural network.

[0081] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto, and any changes or replacements that a person skilled in the art can easily think of within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention.

Claims

1. 1. A method for reconstructing fine-grained haptic signals for audiovisual aids, comprising: A step of inputting a haptic signal to a haptic autoencoder and extracting features of the haptic signal through a clustering task; Transferring and optimizing the feature extraction ability of the haptic autoencoder to the audio feature extraction network and the image feature extraction network by a cross-modal transfer learning method; constraining the extracted haptic, audio and image modal features by jointly considering clustering constraints, centrality constraints and sorting constraints to bring modal features belonging to the same meaning close to each other, but separate modal features not belonging to the same meaning, to obtain haptic, audio and image modal features with fine-grained classification; Constructing a triplet set based on haptic, audio and image modal features, performing shared semantic learning of triplet constraints, optimizing a multi-modal fusion mapping function, and obtaining a fusion feature containing shared semantic information; A method for reconstructing a fine-grained haptic signal for audiovisual assistance, comprising the steps of presetting a haptic generative network and inputting the fusion features into the haptic generative network to reconstruct the haptic signal.

2. The step of preconfiguring a haptic generative network and inputting the fused features into the haptic generative network to reconstruct a haptic signal includes: Presetting parameters of a haptic autoencoder, an audio feature extraction network, and an image feature extraction network; Presetting parameters of a multimodal fusion mapping function and parameters of a haptic generation network; training a haptic autoencoder, an audio feature extraction network, an image feature extraction network, a multimodal fusion mapping function, and a haptic generation network; The method for reconstructing a fine-grained haptic signal for audiovisual assistance as described in claim 1, characterized in that it includes: inputting the just received image signal and audio signal into a trained image feature extraction network and an audio feature network, respectively, to obtain image features and audio features, respectively; then inputting the features into a multimodal fusion mapping function to obtain fusion features; and finally inputting the fusion features into a trained haptic generation network to obtain a reconstructed haptic signal.

3. Inputting haptic signals into a haptic autoencoder and extracting features from the haptic signals through a clustering task is The haptic signal is input to a haptic autoencoder for learning, and corresponding haptic features are extracted. Then, clustering based on the K-means algorithm is performed on the haptic signal based on the haptic features, i.e. h is the input haptic signal, and h = {h i } i=1,・・・,N where i denotes the sort subscript of the input haptic signal, N is the total amount of input haptic signals, and the encoding module E h After passing through (・), i h = E h (h i ;θ he ) is the haptic signal h i is a feature representation of h = {f i h } i=1,・・・,N and θ he are the parameters of the encoding module, and f i h Decoding module D h Input to (・) and output haptic signal [0061] where θ hd are the parameters of the decoding module, and the feature f i h Clustering is performed based on the K-means algorithm on the corresponding category tag s i h We output the loss function [0062] (however, [0063] is the reconstruction error of the encoder, N is the number of haptic signals, [0,64] is the clustering error of K-means, M is the cluster center vector matrix for obtaining tactile data by the K-means algorithm, and m c denotes the center of mass of the cth cluster, and θ h =[θ he ,θ hd ] are the parameters of the encoder and decoder modules, and s j,i h is i h is the j-th element of j,i h If the value is 1 and all other elements are 0, then s i h The original haptic signal h corresponding to i indicates that it belongs to the jth category, and l is the least squares loss [0.65] where λ is a regularization parameter, λ≧0; L clu h By minimizing θ h Estimate f i h and s i h and acquiring a fine-grained haptic signal based on the fine-grained haptic signal.

4. Transferring the feature extraction ability of a haptic autoencoder to an audio feature extraction network and an image feature extraction network through a cross-modal transfer learning method. The feature self-adaptation method realizes the transfer by minimizing the maximum average difference criterion between the tactile and audiovisual domains, i.e. The distributions of the haptic, audio, and image signal sets are P, Q, and R, respectively, and the MMD between the haptic and audio signals is MMD k (P,Q), and the MMD between the tactile signal and the visual signal is MMD k (P, R), and the reproducing kernel Hilbert space H k In the k contains a set of functions f defined on a non-empty set, and the square of the MMD is [0.66] and Here, the haptic, audio, and image signals are passed through respective feature extraction networks φ to obtain extracted feature vectors, and the haptic feature vector is φ h (h;θ he ), i.e. the output of the encoding module of the autoencoder, and the audio and image feature vectors are denoted as φ a (a;θ a ) = f a , φ v (v;θ v ) = f v and the three modal feature sets are (f h ,f a ,f v ) and θ a and θ v are the parameters of the audio and image feature extraction networks, respectively, and θ he are the parameters of the encoding module, and for any function f ∈ H k and for any X∈P, [0.67] and μ k (P) is P's H k is the mean embedding in distribution P, i.e., H k f(X) is an element representation in the space, and X is expressed as a function f in H k It indicates mapping to space, and <・,・> Hk is the dot product operation, and similarly, [0068] and μ k (Q) is Q's H k is the average embedding in [0,69] and μ k (R) is R's H k It is the average embedding in handle [Number 70] The value of is calculated as the loss function of cross-modal transfer, and the specific formula is [Equation 71] And, L CT By optimizing θ, we guide the information flow between the haptic feature extraction autoencoder model and the audio-image feature extraction network, and effectively transfer the feature extraction ability of the autoencoder for the haptic modality to the audio-image feature extraction network, i.e., θ a and θ v 4. The method of claim 3, further comprising: estimating a fine-grained haptic signal for audiovisual assistance.

5. Further optimizing the audio feature extraction network and the image feature extraction network includes: The classification loss function is [Equation 72] (However, s i a and s i v are the category tags of the audio signal and the image signal, respectively, and L clu av By minimizing θ a and θ v The optimal value of p(f i a ;θ a ) means that the audio feature extraction network input is f i a , the network parameters are θ a If the audio signal category s i a is the probability of obtaining i v ;θ v ) means that the image feature extraction network input is f i v , the network parameters are θ v If the image signal category s i v is the probability of obtaining a = {f i a } i=1,・・・,N , f v = {f i v } i=1,・・・,N (which is) L clu h , L CT、 and L clu av Combining Total objective function L clu = L clu h +L CT +L clu av and L clu By minimizing h , θ a、 and θ v After determining the parameters, the features f corresponding to the haptic signal, the audio signal, and the image signal can be obtained. h , f a , f v and obtaining a fine-grained haptic signal based on the fine-grained haptic signal.

6. Constraining the extracted tactile, audio, and image modal features by jointly considering the clustering constraint, the centrality constraint, and the sorting constraint, so as to bring the modal features belonging to the same meaning close to each other, but separate the modal features not belonging to the same meaning, and obtaining the tactile, audio, and image modal features with fine-grained classification; In order to ensure compactness between features of the same fine-grained subcategory, clustering learning under center constraint is performed on three types of modal signals. In order to achieve better fine-grained classification performance, features of the same subcategory should be adjacent in a common space, with the objective of minimizing the intra-category variance. Clustering learning is driven by minimizing the distance from a feature to its subcategory center. The subcategory center of the tactile signal is set as the common subcategory center of the cross-modal signal to ensure compactness between cross-modal signal semantic features of the same category. The loss function of clustering learning under center constraint is: [73] is defined as where N is the number of haptic, audio, and image signals, and s i h , s i a , and s i v denote the categories of haptic signals, audio signals, and image signals, respectively, audio signals and image signals share the cluster center vector matrix M of haptic signals, and each modal feature f i a , f i v , and f i h 2. The method for reconstructing a fine-grained haptic signal for audiovisual assistance according to claim 1, further comprising: bringing the fine-grained haptic signals closer to each other.

7. In order to ensure that the features of different fine-grained subcategories have a certain sparsity, we perform clustering learning under sorting constraints on the three types of modal signals. The goal of the centering constraint is to minimize the within-category variance, while the goal of the sorting constraint is to maximize the between-category variance, so that the feature outputs of different subcategories are less similar than the feature outputs of the same subcategory. The sorting constraint is: [74] is defined as where C is the total number of categories after the tactile signals are clustered by the K-means algorithm, and each modal feature f i a , f i v , and f i h 7. The method for reconstructing fine-grained haptic signals for audiovisual assistance according to claim 6, further comprising: separating modal features having different meanings as much as possible while bringing them closer together.

8. Constructing a triplet set based on modal features of tactile, audio and image, performing shared semantic learning of triplet constraints, optimizing a multi-modal fusion mapping function, and obtaining a fusion feature containing shared semantic information. One haptic signal sample from a fine-grained subcategory i randomly selecting a sample as an anchor; From the image dataset i belongs to the same category as and has semantic features f i h The closest sample to i + Select h as a positive match sample, i does not belong to the same category as f and has semantic features i h The closest sample to j - Selecting as a negative match sample; This allows for a triplet set {(h i , v i + , v j - )) and Similarly, the triplet set {(h i , a i + , a j - )) and Anchor point h i Semantic feature f of i h and positive match semantic features within the corresponding subclasses in the audio-image modal [Number 75] and minimize the distance between i h and Semantic features of negative matches [76] By maximizing the semantic feature between and and having one minimum interval δ, the two triplet loss functions obtained are [77] and Introducing an integrated paradigm, we achieve high-level fusion of multimodal features, i.e. f a and v The process is as follows: f m =F m (f a ,f v ;θ m ) (However, f m is the output of multimodal fusion in the shared semantic subspace, i.e., the fusion feature, and F m (・) is the parameter θ m is a multimodal fusion mapping function of F m (・) is f a and f v (linear weighting of f + a and + v f + m Fusing to f - a and ― v f ― m To merge into The triplet loss is used to constrain the fusion features, i.e. [78] and The objective function of shared semantic learning is modeled by combining three loss functions: L Syn = L syn a +L syn v +L syn m and L Syn By minimizing m and acquiring a fine-grained haptic signal based on the fine-grained haptic signal.

9. Inputting the fused features into a haptic generation network to reconstruct a haptic signal; A haptic generation network is constructed, that is, a haptic generation network G(·) is constructed, and its structure is D h (·) and its network parameters θ hd The parameter θ of G(·) G The initial value of The fused features containing the necessary semantic information are input to a haptic generation network G(·) to obtain a desired haptic signal h′, and the generated haptic signal h′ is expressed as E h (・) represents the tactile feature f h′ , and select the category center to map the tactile feature f h′ We apply semantic constraints to the,final loss function, [79] is shown as however, [Number 80] and [Number 81] The feature is h and h′ It shows the similarity to [0082] GA h′ The clustering loss of θ is calculated by optimizing the loss function. G and obtaining an optimal value of G(·), i.e., determining G(·).

10. Pre-configure the parameters of the haptic autoencoder, audio feature extraction network, and image feature extraction network. Presetting parameters of the multimodal fusion mapping function and parameters of the haptic generation network; Training the haptic autoencoder, the audio feature extraction network, the image feature extraction network, the multimodal fusion mapping function, and the haptic generation network includes step 1 and step 2; In step 1, θ v , θ a , θ he , θ hd , and M are preset, and {s i h }, {s i a }, {s i v }, and the step 1 is Network parameter θ v , θ a , θ he , θ hd , node tag {s i h }, {s i a }, {s i v }, and the category center matrix M is initialized, the number of clusters C, the learning rate μ 1、 and a step 11 of setting the number of iterations T; {s i h }, {s i a }, {s i v } and M are fixed, and θ v , θ a , θ he , θ hd Optimize, i.e., [Number 83] and where step 12 is where ∇ is the partial derivative of each loss function; θ v , θ a , θ he , θ hd , and M are fixed, and {s i h }, {s i a }, {s i v }, i.e. [Number 84] Step 13, θ v , θ a , θ he , θ hd , and {s i h }, {s i a }, {s i v } is fixed and M is optimized, i.e. [Number 85] Step 14, If t<T, jump to step 412; if t=t+1, continue with the next iteration; otherwise, end the iteration; After T iterations, the optimal audio feature extraction network parameters θ a , the parameters of the image feature extraction network θ v , the parameters of the haptic autoencoder θ he , θ hd , node tag {s i h }, {s i a }, {s i v }, and a cluster center vector matrix M of the haptic data; In step 2, θ m and θ G The step 2 is as follows: θ m , learning rate μ 2 , μ 3 , iteration number n 1 Step 21 of initializing L Syn Based on this, we use stochastic gradient descent to find θ m Estimate: [Number 86] Step 22, L Gen Based on this, we use stochastic gradient descent to find θ G Update, i.e., [Number 87] Step 23, n<n 1 if n=n+1, jump to step 22; if n=n+1, continue with the next iteration; otherwise, end the iteration; n 1 After iterations, the optimal θ m and θ G 10. The method of claim 9, further comprising the step of:

Citation Information

Patent Citations

  • Cross-modal image generation method and device based on audio-tactile signal fusion

    CN113627482A

  • Audio and video auxiliary tactile signal reconstruction method based on cloud edge collaboration

    CN113642604A

  • Learning data generation program, learning data generation apparatus and learning data generation method

    JP2020079984A