A method for reconstructing fine-grained tactile signals for audiovisual aids

By employing clustering and cross-modal transfer learning with shared semantic constraints, the method addresses noise interference and limited information in haptic signal reconstruction, achieving high-quality, fine-grained tactile signal generation for cross-modal services.

JP7742677B2Active Publication Date: 2025-09-22NANJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024568338
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-18
Filing Date
2023-11-15
Publication Date
2025-09-22
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

Existing methods for reconstructing haptic signals in cross-modal services face challenges due to noise interference, limited information utilization, and the need for fine-grained tactile signal generation, especially in applications like remote control and virtual interactive experiences, where traditional signal processing and cross-modal reconstruction methods fail to provide accurate tactile sensations.

Method used

A method involving clustering analysis, cross-modal transfer learning, and shared semantic learning to reconstruct fine-grained haptic signals by constraining and optimizing features from haptic, audio, and image modalities using a haptic autoencoder, audio feature extraction network, and image feature extraction network, with clustering, centrality, and sorting constraints to enhance feature classification and fusion.

Benefits of technology

The method achieves high-quality, fine-grained reconstruction of haptic signals, effectively utilizing existing datasets with weak supervision and matching, meeting the demands of cross-modal services by generating accurate tactile sensations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742677000090
    Figure 0007742677000090
  • Figure 0007742677000091
    Figure 0007742677000091
  • Figure 0007742677000092
    Figure 0007742677000092
Patent Text Reader

Abstract

The present invention discloses a method for reconstructing fine-grained haptic signals for audiovisual assistance, which first passes the training haptic signal through a haptic autoencoder, extracts haptic features based on a clustering task, and transfers the features to an audio-visual feature extraction network to realize feature extraction of the audio-visual signal; then optimizes a multi-modal fusion mapping function of haptic, audio, and image using triplet constraints, thereby obtaining fusion features from the extracted audio-visual features; and finally, inputs the fusion features into a haptic generation network to realize the reconstruction of fine-grained haptic signals. The present invention better solves the problems of weak supervision and weak matching existing between multi-modal signals, realizes cross-modal shared semantic learning, ensures the structure and semantic completeness of the generated haptic signal, and also enhances its clustering properties, thereby significantly improving the reconstruction quality of the haptic signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of generating haptic signals, and in particular to a method for reconstructing fine-grained haptic signals for audiovisual aids. [Background technology]

[0002] As traditional multimedia application technologies mature, people's audiovisual needs are largely met, and they are beginning to pursue more dimensional and higher-level sensory experiences. Haptic information is gradually being integrated into existing audio / video multimedia services to form multimodal services, providing a richer and more interactive experience. Cross-modal communication technologies have been proposed to support cross-modal services, and while they have some effectiveness in ensuring the quality of multimodal streams, applying cross-modal communication to haptic-based multimodal services still faces several technical challenges. First, haptic streams are highly sensitive to interference and noise in wireless links, resulting in degradation or even loss of haptic signals at the receiving end. This problem is particularly serious and unavoidable in remote control applications, such as remote industrial control and remote surgery. Secondly, service providers do not have tactile collection devices, but users need to sense tactile sensations. Particularly in virtual interactive application scenes such as online immersive shopping, holographic museum guides, and virtual interactive movies, users have a very high need for tactile perception, and therefore require the ability to generate "virtual" touch sensations or tactile signals based on video and audio signals.

[0003] Currently, haptic signals corrupted or partially lost due to unreliable wireless communication and communication noise interference can be self-recovered in two ways. The first is based on traditional signal processing techniques. It uses sparse representation to search for a specific signal with the most similar structure, and then uses this specific signal to estimate the missing parts of the corrupted signal. The second is to mine and utilize the spatiotemporal correlation of the signal itself to achieve intra-modal self-recovery and reconstruction. However, when the haptic signal is severely corrupted or even non-existent, intra-modal reconstruction schemes fail.

[0004] In recent years, several studies have focused on the correlation between different modalities and achieved cross-modal reconstruction. In their paper, "Learning Cross-Modal Visual-Tactile Representation Using Ensembled Generative Adversarial Networks," Li et al. proposed using image features to obtain the necessary category information, and then using this information along with noise as input to a generative adversarial network to generate a tactile spectral map of the corresponding category. However, this method often produces inaccurate results due to the limited information obtained by mining semantic correlations and categories between modalities. In their paper, "Deep Visuo-Tactile Learning: Estimation of Tactile Properties from Images," Kuniyuki Takahashi et al. extended a single encoder-decoder network to embed both visual and tactile attributes into a latent space, focusing on the degree of tactile properties of materials represented by the latent variables. Furthermore, in the paper "Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces From Images," Matthew Purr et al. proposed a cross-modal learning framework with adversarial learning and cross-domain associative classification to estimate tactile physical properties from a single image. Although such methods utilize modal semantic information, they do not generate a complete tactile signal, making them of no practical value for cross-modal services.

[0005] The above-mentioned existing cross-modal generation methods also suffer from the following shortcomings: The training of these models all relies on large-scale training data to ensure the effectiveness of the models, and they all only utilize single-modal information. However, in reality, the advantages of a single modality cannot provide a sufficient amount of information. When different modalities jointly describe the same meaning, they may contain unequal amounts of information. Complementary and enhanced information between modalities contributes to improved generation effectiveness. In practical applications, the annotation costs for large datasets are enormous, and coarse-grained, broad classification categories are often more easily obtained, while fine-grained categories are unclear. Furthermore, there is no direct correspondence between samples of different modalities, posing challenges in weak supervision and weak matching. Summary of the Invention [Problem to be solved by the invention]

[0006] The technical problem to be solved by the present invention is to provide a method for reconstructing fine-grained haptic signals for audiovisual assistance that overcomes the drawbacks of the prior art. First, cluster analysis is performed within coarse-grained categories to obtain fine-grained classifications of samples. Next, modal co-semantic constraints are performed in the fine-grained categories, with the goal of minimizing intra-category differences and maximizing inter-category differences. Finally, correlation constraints are performed by searching for audiovisual samples that positively match the haptic signals in the fine-grained subcategories, thereby achieving high-quality, fine-grained reconstruction of the haptic signals. [Means for solving the problem]

[0007] In order to solve the above technical problems, the present invention provides a method for reconstructing fine-grained tactile signals for audiovisual assistance, inputting a haptic signal into a haptic autoencoder and extracting features from the haptic signal through a clustering task; transferring and optimizing the feature extraction ability of the haptic autoencoder to the audio feature extraction network and the image feature extraction network by a cross-modal transfer learning method; Constraining the extracted tactile, audio, and image modal features by jointly considering clustering constraints, centrality constraints, and sorting constraints to bring modal features belonging to the same meaning closer together but separate modal features that do not belong to the same meaning, thereby obtaining tactile, audio, and image modal features with fine-grained classification; Constructing a triplet set based on tactile, audio and image modal features, performing shared semantic learning of triplet constraints, optimizing a multimodal fusion mapping function, and obtaining fusion features containing shared semantic information; and preconfiguring a haptic generation network and inputting the fused features into the haptic generation network to reconstruct a haptic signal.

[0008] Furthermore, the step of preconfiguring a haptic generation network and inputting the fused features into the haptic generation network to reconstruct a haptic signal includes: Presetting parameters of a haptic autoencoder, an audio feature extraction network, and an image feature extraction network; Presetting parameters of a multimodal fusion mapping function and parameters of a haptic generation network; training a haptic autoencoder, an audio feature extraction network, an image feature extraction network, a multimodal fusion mapping function, and a haptic generation network; The method includes inputting the just-received image signal and audio signal into a trained image feature extraction network and an audio feature network, respectively, to obtain image features and audio features, then inputting the features into a multimodal fusion mapping function to obtain fusion features, and finally inputting the fusion features into a trained haptic generation network to obtain a reconstructed haptic signal.

[0009] Furthermore, inputting tactile signals into a tactile autoencoder and extracting features from the tactile signals through a clustering task is The tactile signals are input to a tactile autoencoder for learning, and the corresponding tactile features are extracted. Then, the tactile signals are clustered based on the K-means algorithm according to the tactile features, i.e., h is the input tactile signal, and h = {h i} i=1,…,N where i denotes the sort subscript of the input haptic signal, N is the total amount of input haptic signals, and the encoding module E of the autoencoder h After passing through (·), f i h =E h (h i ;θ he ) is the tactile signal h i is a feature representation of f h ={f i h} i=1,…,N and θ he are the parameters of the encoding module, and f i h Decryption module D h (·) input and output haptic signals

number

number

number

number

number

[0010] Furthermore, transferring the feature extraction ability of a haptic autoencoder to an audio feature extraction network and an image feature extraction network through a cross-modal transfer learning method is The feature self-adaptation method minimizes the maximum average difference criterion between the tactile and audiovisual areas to achieve the transition, i.e., The distributions of the haptic, audio, and image signal sets are P, Q, and R, respectively, and the MMD between the haptic and audio signals is MMD. k (P,Q), and the MMD between the tactile signal and the visual signal is MMD k (P,R) and the reproducing kernel Hilbert space H k In H kcontains a set of functions f defined on a non-empty set, and the square of the MMD is

number

number

number

number

number

number

[0011] Further optimizing the audio feature extraction network and the image feature extraction network may include: The classification loss function is

number

[0012] Furthermore, by constraining the extracted tactile, audio, and image modal features by jointly considering the clustering constraint, the centrality constraint, and the sorting constraint, the modal features belonging to the same meaning are brought close to each other, but the modal features not belonging to the same meaning are separated, thereby obtaining tactile, audio, and image modal features with fine-grained classification. To ensure compactness between features of the same fine-grained subcategory, clustering learning under center constraints is performed on three types of modal signals. To achieve better fine-grained classification performance, features of the same subcategory should be adjacent in a common space, with the goal of minimizing intra-category variance. Clustering learning is driven by minimizing the distance from a feature to its subcategory center. The subcategory center of the tactile signal is set as the common subcategory center of the cross-modal signal to ensure compactness between semantic features of the same category of cross-modal signals. The loss function of clustering learning under center constraints is:

number

[0013] Furthermore, to ensure that the features of different fine-grained subcategories have a certain sparsity, we perform clustering learning under sorting constraints on the three types of modal signals. The goal of the centrality constraint is to minimize the intra-category variance, while the goal of the sorting constraint is to maximize the inter-category variance, so that the feature outputs of different subcategories are less similar than the feature outputs of the same subcategory. The sorting constraint is:

number

[0014] Furthermore, constructing a triplet set based on the modal features of tactile, audio, and image, conducting shared semantic learning of triplet constraints, optimizing the multimodal fusion mapping function, and obtaining fusion features containing shared semantic information is One haptic signal sample h from a fine-grained subcategory i randomly selecting a sample and setting the sample as an anchor; Image dataset from h i belongs to the same category as f and has semantic features i h The closest sample to v i + Select as a positive match sample, and h i does not belong to the same category as f and has semantic features i h The closest sample to v j - Selecting as a negative match sample; This allows us to create a triplet set {(h i ,v i + ,v j - )}; and Similarly, the triplet set {(h i ,a i + ,a j - )}, and Anchor point h i Semantic features of f i hand positive match semantic features within the corresponding subclassifications in the audio-visual modal.

number

number

number

number

[0015] Furthermore, we construct a tactile generation network, i.e., a tactile generation network G(·), whose structure is D h (·) and its network parameter θ hd The parameter θ of G(·) G The initial value of The fused features containing the necessary semantic information are input to the haptic generation network G(·) to obtain the desired haptic signal h′, and the generated haptic signal h′ is expressed as E h (·) gives the tactile feature f h′ and select the category center to map the tactile feature f h′ We apply semantic constraints to the

number

number

number

number

[0016] Furthermore, we pre-set the parameters of the haptic autoencoder, audio feature extraction network, and image feature extraction network. Presetting the parameters of the multimodal fusion mapping function and the parameters of the haptic generation network; Training the haptic autoencoder, audio feature extraction network, image feature extraction network, multimodal fusion mapping function, and haptic generation network includes step 1 and step 2; In step 1, θ v , θ a , θ he , θ hd , and M are preset, and {s i h}, {s i a}, {s i v}, and step 1 is Network parameter θ v , θ a , θ he , θ hd , node tag {s i h}, {s i a}, {s i v}, and a category center matrix M, and step 11 of setting the number of clusters C, the learning rate μ1, and the number of iterations T; {s i h}, {s i a}, {s i v}, and M is fixed, and θ is calculated based on stochastic gradient descent. v , θ a , θ he , θ hd Optimize, i.e.,

number

[0017] Compared with the prior art, the present invention has the following technical effects by using the above technical solution.

[0018] [[ID= fifty]] After learning the fine-grained classification of three types of modal samples by the cross-modal transfer deep clustering algorithm, the present invention performs shared semantic learning, fully exploits the advantages of multi-modal feature fusion, and finally realizes the generation of fine-grained tactile signals based on clustering constraints, maximally utilizes the existing dataset with problems of weak supervision and weak matching, thereby generating high-quality and fine-grained tactile signals, which better meet the requirements of cross-modal services. [Brief Description of the Drawings] <H

[0019] [Figure 1]1 is a flowchart of a method for reconstructing fine-grained haptic signals for audiovisual aids according to the present invention; [Figure 2] 1 is a structural schematic diagram of a complete network according to the present invention; [Figure 3] FIG. 1 is a schematic diagram of the architecture of a deep clustering model based on cross-modal transfer according to the present invention. [Figure 4] 10A and 10B show the reconstruction results of tactile signals of the present invention and other comparative methods. DETAILED DESCRIPTION OF THE INVENTION

[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0021] The present invention provides a method for reconstructing fine-grained haptic signals for audiovisual assistance, the flow chart of which is shown in FIG. 1, and the method includes the following steps 1 to 5.

[0022] In step 1, the haptic signal is first input into a haptic autoencoder, and feature extraction of the haptic signal is achieved through clustering. Next, cross-modal transfer learning technology is used to transfer and optimize the feature extraction ability of the haptic autoencoder to an audio feature extraction network and an image feature extraction network, respectively. After that, the extracted haptic, audio, and image modal features are further constrained by jointly considering clustering constraints, centrality constraints, and sorting constraints, so that modal features belonging to the same meaning are close together, but modal features that do not belong to the same meaning are separated, and haptic, audio, and image modal features with fine-grained classification are obtained.

[0023] (1-1) First, clustering-constrained feature learning is performed on three types of modal signals, namely, haptic, image, and audio, to obtain segment features with fine-grained subcategories, which may be divided into the following three steps, namely, step (1-1-1) to step (1-1-3).

[0024] (1-1-1) In the first step, the tactile signal is input to the autoencoder for learning, and the corresponding tactile features are extracted. Then, the tactile signal is clustered based on the K-means algorithm according to the tactile features. That is, Specifically, h i is the input tactile signal, and h={h i} i=1,···,N where i denotes the sort subscript of the input haptic signal, N is the total amount of input haptic signals, and the encoding module E of the autoencoder h After passing through (·), f i h =E h (h i ;θ he ) is the tactile signal h i is a feature representation of f h ={f i h} i=1,···,N and θ he Assume that are the parameters of the encoding module, and f i h Decryption module D h (·) input and output haptic signals

number

number

number

number

number

[0025] (1-1-2) In the second step, the acquired tactile features are transferred to the feature extraction process of the image and audio signals using cross-modal transfer learning technology. That is, the transfer is achieved by minimizing the maximum mean difference (MMD) criterion between the tactile domain and the audiovisual domain using a feature self-adaptation method.

[0026] Specifically, let the distributions of the haptic, audio, and image signal sets be P, Q, and R, respectively. The MMD between the haptic signal and the audio signal is defined as MMD k (P,Q), and the MMD between the tactile signal and the visual signal is MMD k (P,R). The reproducing kernel Hilbert space Hk In H k contains a set of functions f defined on a non-empty set, and the square of the MMD is

number

number

number

number

[0027] handle

number

number

[0028] L CT By optimizing θ, we can guide the information flow between the haptic feature extraction autoencoder model and the audio-visual feature extraction network, thereby effectively transferring the feature extraction ability of the autoencoder for the haptic modality to the audio-visual feature extraction network, i.e., θ a and θ v can be estimated.

[0029] (1-1-3) In the third step, the audio-visual feature extraction network was further optimized. The video feature extraction network was designed in the VGG network style, with a 3x3 convolutional filter and a 2x2 max-pooling layer with a step width of 2 and no padding. The network was divided into four blocks, each containing two convolutional layers and one pooling layer, with the number of filters doubling between successive blocks. Finally, max-pooling was performed across all spatial locations to generate a single 512-dimensional semantic feature vector. This semantic feature vector was then input to a three-layer fully connected neural network (256-128-32) with K nodes and one fully connected layer with a softmax function. The 32-dimensional vector, which is the visual signal feature vector, was further trained using subsequent triplet-constrained shared semantic learning, where K is the number of fine-grained subcategories in each coarse-grained category, and K is 3 in this experimental dataset. The related network structure for audio signals is similar in setting to visual signals and shares classifiers with visual signals.

[0030] The designed classification loss function is

number

[0031] It is particularly noteworthy that the tags are obtained using the following incremental policy: first, pseudo-tags are attached to the data, and then these data are used to optimize the model, so that the optimized model's ability to classify audiovisual signals is enhanced; then, the trained model is used to perform pseudo-tag operations, thereby updating the pseudo-tags. This incremental optimization method gradually improves the network's fine-grained classification ability.

[0032] In short, the total objective function in step (1-1) is a combination of the three loss functions above, L clu =L clu h +L CT +L clu av may be shown as L clu By minimizing h , θ a , and θ v After determining the parameters, the features f corresponding to the haptic signal, audio signal, and image signal can be obtained. h , f a , f v You can get (1-2) To ensure compactness between features of the same fine-grained subcategory, clustering learning under center constraint is performed on three types of modal signals. To achieve better fine-grained classification performance, features of the same subcategory should be adjacent in a common space, with the goal of minimizing intra-category variance. Specifically, clustering learning is driven by minimizing the distance from a feature to its subcategory center. By setting the subcategory center of the tactile signal as the common subcategory center of the cross-modal signal, compactness between semantic features of cross-modal signals of the same category is ensured. The loss function of clustering learning under center constraint is:

number

[0033] (1-3) To ensure that the features of different fine-grained subcategories have a certain sparsity, we perform clustering learning under sorting constraints on three types of modal data. The goal of the centrality constraint is to minimize the within-category variance, while the goal of the sorting constraint is to maximize the between-category variance, so that the feature outputs of different subcategories are less similar than the feature outputs of the same subcategory. The sorting constraint is:

number

[0034] In step 2, after obtaining tactile, audio, and image features with fine-grained classification, the three modal features are constructed into a triplet set, triplet-constrained shared semantic learning is performed, and the multimodal fusion mapping function is optimized to obtain fusion features containing shared semantic information, which lays the foundation for generating tactile signals.

[0035] Specifically, Step 2 is as follows:

[0036] (2-1), one tactile signal sample from a fine-grained subcategory h i Randomly select a sample as an anchor, and then select h from the image dataset. i belongs to the same category as f and has semantic features i h The closest sample to v i + Select as a positive match sample, and then h i does not belong to the same category as f and has semantic features i h The closest sample to v j - As a negative match sample, we select the triplet set {(h i ,v i+ ,v j - Similarly, a triplet set {(h i ,a i + ,a j - )} can be obtained. Anchor point h i Semantic features of f i h and positive match semantic features within the corresponding subclassifications in the audio-visual modal.

number

number

number

[0037] (2-2) Based on the above, we introduce an integration paradigm to achieve advanced multimodal feature fusion. Specifically, audiovisual data is first passed through an audio feature extraction network and an image feature extraction network, respectively, to obtain features f a and f v and then f a and f v The process is as follows: f m =F m (f a ,f v ;θ m ) (However, f m is the output of multimodal fusion in the shared semantic subspace, i.e., the fusion feature, and F m (·) is the parameter θ mis a mapping function of F m (·) is f a and f v (linear weighting of the

[0038] f + a and f + v f + m fused to f - a and f ― v f ― m Similarly, triplet loss is used to constrain the fused features, i.e.,

number

[0039] (2-3), the objective function of shared semantic learning may be modeled by combining three loss functions: L Syn =L syn a +L syn v +L syn m It may be shown as:

[0040] L Syn By minimizing m , which lays the foundation for the next stage of tactile signal generation.

[0041] In step 3, the fused features are input into a haptic generation network to generate the desired haptic signal h′.

[0042] Specifically, Step 3 is as follows:

[0043] First, we constructed a haptic generation network and a haptic decoder D hSince (·) is trained as part of the complete autoencoder in step (1), we now construct a separate haptic generation network G(·), whose structure is D h (·) is (32-128-256-Z), and its network parameters θ hd The parameter θ of G(·) G The fused features containing the necessary semantic information are input to the haptic generation network G(·) to obtain the desired haptic signal h′, and the generated haptic signal h′ is expressed as E h (·) gives the 32-dimensional tactile feature f h′ and select the category center to map the tactile feature f h′ Obviously, the tactile feature f h′ The distance between and the category center of the corresponding category is made as small as possible, but the distance between and the center of other categories is made as large as possible. The final loss function is

number

number

number

number

[0044] In step 4, the model is trained. That is, the model training is divided into two steps. In the first step, the parameters of the haptic autoencoder, audio feature extraction network, and image feature extraction network are estimated. In the second step, the parameters of the multimodal fusion mapping function and the haptic generation network are estimated. Through these steps, the haptic autoencoder, audio feature extraction network, image feature extraction network, multimodal fusion mapping function, and haptic generation network are trained.

[0045] Specifically, Step 4 is as follows:

[0046] θ v , θ a , θ he , θ hd , and M are estimated, and {s i h}, {s i a}, {s i v} to optimize.

[0047] Step 411, network parameter θ v , θ a , θ he , θ hd , node tag {s i h}, {s i a}, {s i v}, and the category center matrix M is initialized, the number of clusters C is set, the learning rate μ1=0.0001, and the number of iterations T=600.

[0048] Step 412, i h}, {s i a}, {s i v}, and M are fixed, and θ v , θ a , θ he , θ hd Optimize, i.e.,

Number

[0049] Step 413. Fix θ v θ a θ he θ hd and M, and optimize {s i h}, {s i }, {s i v [[ID=<>]] v θ a θ he θ hd and {s i h}, {s i a [[ID=<>]] i v}, {s i v} and optimize M, that is,

Number

[0051]

[0052] [[ID=<>]]

[0053] a v he hd i h i It should be noted that there seems to be some incomplete or unclear parts in the original text, especially in the middle part where some tags are not properly closed or there are some "u" and "<>" notations that might be errors. This translation is done based on the best understanding of the existing text.​​​​​​​a}, {s i v}, and obtain the cluster center vector matrix M of the tactile data.

[0053] In particular, m k When updating

number

[0054] (4-2), based on SGD, θ m and θ G Complete the estimation of.

[0055] Step 421, θ m is initialized with a batch size of bactch = 64, learning rates μ2, μ3 = 0.0001, and number of iterations n1 = 600.

[0056] Step 422, L Syn Based on this, we use stochastic gradient descent to find the m Fine-tune the

number

[0057] Step 423, L Gen Utilize stochastic gradient descent based on it to update θ G That is,

Number

[0058] Step 424, if n < n1, jump to Step 422, if n = n + 1, continue the next iteration, otherwise, end the iteration.

[0059] Step 425, after 600 iterations, obtain the optimal n = n + 1.

[0060] In Step 5, input the just received image signal and audio signal into the trained image feature extraction network and audio feature network respectively, obtain the image feature and audio feature respectively, then input the above features into the multimodal fusion mapping function to obtain the fusion feature, and finally input the fusion feature into the trained tactile generation network to obtain the reconstructed tactile signal.

[0061] Step 5 is specifically as follows.

[0062] Input the just received image signal v and audio signal a into the trained image feature extraction network and audio feature extraction network respectively, and obtain the image feature

Number

Number

Number

Number

number

[0063] As can be seen from the following experimental results, compared with the conventional method, the present invention realizes the synthesis of tactile signals through the complementary fusion of multimodal meanings, and achieves a higher generation effect.

[0064] This example uses the LMT cross-modal dataset proposed in the paper "Multimodal feature-based surface material classification," which includes samples from nine semantic categories: grid, stone, metal, wood, rubber, fiber, foam, foil, paper, textiles, and fabrics. This example selects five major categories (each with three subcategories) for the experiment. To reconstruct the LMT dataset, we first refer to the training and test sets for each material example, and obtain 20 image samples, 20 audio signal samples, and 20 tactile signal samples for each example. Next, we train the neural network by augmenting the data. Specifically, we flip each image horizontally and vertically, rotate them at any angle, and use techniques such as random scaling, cutting, and offsetting in addition to traditional methods. Thus, we augmented the data for each category by 100, resulting in a total of 1500 images, each with a dimension of 224. * 224. In the dataset, 80% is selected to be used for training, while the remaining 20% ​​is used for testing and performance evaluation. The experiment is initially set up so that the fine-grained categories are unknown.

[0065] (1) Clustering results To verify the effectiveness of the clustering method of the present invention, the clustering method is compared with several baseline methods, including:

[0066] K-means (KM) The K-Means algorithm clusters the samples of the image, audio, and haptic modalities, respectively.

[0067] Autoencoder + K-means (AE+KM, Autoencoder followed by K-means) This is a two-step method: first, we obtain feature representations for each modality by reconstructing and learning signal samples of different modalities, and then cluster them using K-means.

[0068] Triple Deep Clustering Model (3-DCN, 3-Deep Clustering Network) Clustering is performed using DCN models for signals of different modalities.

[0069] The present invention uses the method of this example.

[0070] When presenting the results, we select the average values ​​across multiple subcategories and mainly adopt three indices: normalized mutual information (NMI), adjusted Rand index (ARI), and clustering accuracy (ACC). The experimental results are shown in Table 1.

[0071] [Table 1]

[0072] Table 1 shows the results of applying the present invention, 3-DCN, AE+KM, and KM to the LMT dataset. As can be seen, the method of the present invention demonstrates extremely high competitiveness on this dataset, achieving significantly better results than traditional clustering algorithms and common deep clustering algorithms. Analysis reveals that the application scenarios of other algorithms are all unimodal, which may lead to unbalanced clustering results. Theoretically, the number of samples in a given category of coexisting cross-modal data should be equal. Furthermore, the subcategory features learned by this clustering method are significantly more compact and more discriminative between different categories, which contributes to the subsequent reconstruction of tactile signals.

[0073] (2) Tactile reconstruction results After determining the fine-grained category, we compare the proposed fine-grained tactile reconstruction method with several methods below.

[0074] Existing Method 1 The paper "Learning cross-modal visual-tactile representation using ensemble generative adversarial networks" (authors X. Li, H. Liu, J. Zhou, and F. Sun) uses ensemble generative adversarial networks (E-GANs, Ensembled GANs) to obtain the necessary category information using image features, then uses the category information along with noise as input to the generative adversarial network to generate a tactile spectral map of the corresponding category, which is then converted into a tactile signal.

[0075] Existing Method 2 The deep vision-tactile learning method (DVTL, Deep Visio-Tactile Learning) in the paper "Deep Visuo-Tactile Learning: Estimation of Tactile Properties from Images" (authors: Kuniyuki Takahashi and Jethro Tan) extends the traditional encoder-decoder network with latent variables to embed visual and tactile attributes into a latent space.

[0076] Existing Method 3 The paper "Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces From Images" (authors: Matthew Purri and Kristin Dana) proposes a joint-encoding-classification GAN (JEC-GAN), which uses different encoding networks to encode each modal instance into a shared internal space, and then couples the embedded visual and tactile samples together in the latent space using paired constraints. Finally, using visual information as input, the corresponding tactile signal is reconstructed by a generative network.

[0077] This experiment is analyzed from two perspectives: quantitative and qualitative. First, Table 2 shows the tactile signal reconstruction performance of each method from multiple perspectives: root mean square error (RMSE), structural similarity (SIM), and classification accuracy (ACC). Table 2 shows the experimental results of the present invention.

[0078] [Table 2]

[0079] As can be seen from Table 2 and Figure 4, the method of the present invention has clear advantages over the most advanced methods mentioned above. The reasons are as follows: (1) The haptic signal reconstruction method of the present invention uses a cross-modal clustering algorithm to clarify fine-grained subcategories and utilizes centrality and sorting constraints to effectively improve the compactness and distinctiveness of semantic features; (2) when the correspondence between visual, auditory, and haptic samples is specified by humans, the accuracy is insufficient due to a relatively strong subjective awareness. In contrast, in the fine-grained haptic signal reconstruction method for audiovisual assistance of the present invention, during training, the model self-selects the audiovisual semantic features that are closest to the haptic semantic features as the input of its generative network.

[0080] In another embodiment, the haptic encoder in step 1 of the present invention may be replaced by one-dimensional convolutional neural networks (1D-CNN) using a feedforward neural network.

[0081] The above description is merely a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto, and any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present invention should be included in the protection scope of the present invention.

Claims

1. 1. A method for reconstructing fine-grained haptic signals for audiovisual aids, comprising: inputting a haptic signal into a haptic autoencoder and extracting features from the haptic signal through a clustering task; Transferring and optimizing the feature extraction ability of the haptic autoencoder to the audio feature extraction network and the image feature extraction network by a cross-modal transfer learning method, i.e., optimizing the parameters of the audio feature extraction network and the image feature extraction network by minimizing the cross-modal transfer loss function and the classification loss function derived from the maximum mean difference (MMD) criterion between the haptic domain and the audiovisual domain; A step of performing clustering constraint feature learning and clustering learning under center constraint and sorting constraint for three types of modal signals of tactile, audio, and image, thereby approximating tactile, audio, and image modal features that belong to the same meaning but separating modal features that do not belong to the same meaning, and obtaining tactile, audio, and image modal features with fine-grained classification; Constructing a triplet set based on tactile, audio, and image modal features, and performing triplet-constrained shared semantic learning, i.e., optimizing parameters of a multimodal fusion mapping function by modeling and minimizing the objective function of shared semantic learning, and inputting the audio features and image features into the optimized multimodal fusion mapping function to obtain fusion features containing shared semantic information; a step of presetting a haptic generation network and inputting the fused features into the haptic generation network to reconstruct a haptic signal.

2. The step of preconfiguring a haptic generative network and inputting the fused features into the haptic generative network to reconstruct a haptic signal includes: Presetting parameters of a haptic autoencoder, an audio feature extraction network, and an image feature extraction network; Presetting parameters of a multimodal fusion mapping function and parameters of a haptic generation network; training a haptic autoencoder, an audio feature extraction network, an image feature extraction network, a multimodal fusion mapping function, and a haptic generation network; 2. The method for reconstructing fine-grained haptic signals for audiovisual assistance according to claim 1, comprising: inputting the just-received image signal and audio signal into a trained image feature extraction network and an audio feature network, respectively, to obtain image features and audio features, respectively; then inputting the features into a multimodal fusion mapping function to obtain fusion features; and finally inputting the fusion features into a trained haptic generation network to obtain a reconstructed haptic signal.

3. Inputting tactile signals into a tactile autoencoder and extracting features from the tactile signals through a clustering task is The haptic signal is input to a haptic autoencoder for learning, and the corresponding haptic features are extracted. Then, the haptic signal is clustered based on the K-means algorithm according to the haptic features, i.e., h is the input haptic signal, and h = {h i } i=1,・・・,N where i denotes the sort subscript of the input haptic signal, N is the total amount of input haptic signals, and the encoding module E h After passing through (・), f i h = E h (h i ;θ he ) is the tactile signal h i is a feature representation of f h = {f i h } i=1,・・・,N and θ he are the parameters of the encoding module, and f i h Decryption module D h Input to (・) and output tactile signal [Number 61] where θ hd are the parameters of the decoding module, and the feature f i h Clustering is performed based on the K-means algorithm on the corresponding category tag s i h The parameters in the above process are jointly estimated, and the loss function [Number 62] (however, [Number 63] is the reconstruction error of the encoder, N is the number of haptic signals, [Number 64] is the clustering error of K-means, M is the cluster center vector matrix for obtaining tactile data using the K-means algorithm, and m c denotes the center of mass of the cth cluster, and θ h =[θ he ,θ hd ] are parameters of the encoder and decoder modules, and s j,i h is i h is the j-th element of j,i h If the value is 1 and all other elements are 0, then s i h The original haptic signal h corresponding to i indicates that belongs to the jth category, and l is the least squares loss [Number 65] where λ is a regularization parameter, λ≧0; L clu h By minimizing θ h Estimate and f i h and s i h and acquiring a fine-grained haptic signal based on the fine-grained haptic signal.

4. Transferring the feature extraction ability of a haptic autoencoder to an audio feature extraction network and an image feature extraction network using a cross-modal transfer learning method is The feature self-adaptation method minimizes the maximum average difference criterion between the tactile and audiovisual areas to achieve the transition, i.e., The distributions of the haptic, audio, and image signal sets are P, Q, and R, respectively, and the MMD between the haptic and audio signals is MMD k (P, Q) and the tactile signal and the visual The MMD between the signal is MMD k (P, R) and the reproducing kernel Hilbert space H k In this case, H k contains a function set f defined on a non-empty set, and the square of the MMD is [Number 66] and However, the haptic, audio, and image signals are passed through respective feature extraction networks φ to obtain extracted feature vectors, and the haptic feature vector is φ h (h;θ he ), i.e., the output of the encoding module of the autoencoder, and the audio and image feature vectors are denoted as φ a (a;θ a ) = f a , φ v (v;θ v ) = f v and the three modal feature sets are (f h ,f a ,f v ) and θ a and θ v are the parameters of the audio and image feature extraction networks, respectively, and θ he are the parameters of the encoding module, and any function f∈H k and for any X∈P, [Number 67] and μ k (P) is P's H k is the mean embedding in the distribution P, i.e., H k f(X) is a representation of one element in the space, and X is expressed as H by the function f. k It indicates mapping to space, and <・,・> Hk is the dot product operation, and similarly, [Number 68] and μ k (Q) is Q's H k is the average embedding in [Number 69] and μ k (R) is R's H k is the mean embedding in handle [Number 70] The value of is calculated as the loss function of cross-modal transfer, and the specific formula is [Number 71] And, L CT By optimizing θ, we guide the information flow between the haptic feature extraction autoencoder model and the audio-visual feature extraction network, and effectively transfer the feature extraction ability of the autoencoder for the haptic modality to the audio-visual feature extraction network. a and θ v 4. The method for reconstructing a fine-grained haptic signal for audiovisual assistance according to claim 3, further comprising:

5. Further optimizing the audio and image feature extraction networks teeth, The classification loss function is [Number 72] (However, s i a and s i v are the category tags of the audio signal and the image signal, respectively, and L clu av By minimizing θ a and θ v Further obtain the optimal value of p(f i a ;θ a ) means that the audio feature extraction network input is f i a , the network parameters are θ a If the audio signal category s i a is the probability of obtaining i v ;θ v ) means that the image feature extraction network input is f i v , the network parameters are θ v If the image signal category s i v is the probability of obtaining a = {f i a } i=1,・・・,N , f v = {f i v } i=1,・・・,N ) and L clu h , L CT、 and L clu av Combined, Total objective function L clu =L clu h +L CT +L clu av and L clu By minimizing h , θ a、 and θ v After determining the parameters, the features f corresponding to the haptic signal, the audio signal, and the image signal can be obtained. h , f a , f v 5. The method for reconstructing a fine-grained haptic signal for audiovisual assistance according to claim 4, further comprising:

6. By constraining the extracted tactile, audio, and image modal features by jointly considering the clustering constraint, the centrality constraint, and the sorting constraint, each modal feature belonging to the same meaning is brought close to each other, but each modal feature not belonging to the same meaning is separated, and tactile, audio, and image modal features with fine-grained classification are obtained. To ensure compactness between features of the same fine-grained subcategory, clustering learning under center constraints is performed on three types of modal signals. To achieve better fine-grained classification performance, features of the same subcategory should be adjacent in a common space, with the goal of minimizing intra-category variance. Clustering learning is driven by minimizing the distance from a feature to its subcategory center. The subcategory center of the tactile signal is set as the common subcategory center of the cross-modal signal to ensure compactness between semantic features of the same category of cross-modal signals. The loss function of clustering learning under center constraints is: [Number 73] is defined as where N is the number of haptic, audio, and image signals, and s i h , s i a , and s i v indicate the categories of haptic signals, audio signals, and image signals, respectively, and audio signals and image signals share the cluster center vector matrix M of haptic signals. i a , f i v , and f i h 2. The method for reconstructing a fine-grained haptic signal for audiovisual assistance according to claim 1, further comprising: bringing the fine-grained haptic signals closer to each other.

7. To ensure that the features of different fine-grained subcategories have a certain sparsity, we perform clustering learning under sorting constraints on three types of modal signals. The goal of the centrality constraint is to minimize the intra-category variance, while the goal of the sorting constraint is to maximize the inter-category variance, so that the feature outputs of different subcategories are less similar than the feature outputs of the same subcategory. The sorting constraint is: [Number 74] is defined as where C is the total number of categories after the tactile signals are clustered using the K-means algorithm, and each modal feature f i a , f i v , and f i h 7. The method for reconstructing fine-grained haptic signals for audiovisual assistance according to claim 6, wherein the modal features having different meanings are separated as much as possible while the modal features having different meanings are brought closer together.

8. Constructing a triplet set based on tactile, audio, and image modal features, performing shared semantic learning of triplet constraints, optimizing a multimodal fusion mapping function, and obtaining fusion features containing shared semantic information is One haptic signal sample from a fine-grained subcategory i randomly selecting a sample and setting the sample as an anchor; From the image dataset i belongs to the same category as i h The sample v closest to i + Select as a positive match sample, and h i does not belong to the same category as i h The sample v closest to j - Selecting as a negative match sample; This allows us to create a triplet set {(h i , v i + , v j - )}; and Similarly, the triplet set {(h i , a i + , a j - )}, and Anchor point h i Semantic feature f of i h and positive match semantic features within the corresponding subcategories in the audio-visual modal. [Number 75] and minimize the distance between i h and semantic features of negative matches [Number 76] By maximizing the semantic feature between and there is one minimum interval δ, the two triplet loss functions obtained are [Number 77] and We introduce an integration paradigm to achieve high-level fusion of multimodal features, i.e., f a and f v The process is as follows: f m =F m (f a ,f v ;θ m ) (However, f m is the output of multimodal fusion in the shared semantic subspace, i.e., the fusion feature, and F m (・) is the parameter θ m is the multimodal fusion mapping function of F m (・) is f a and f v (linear weighting of f + a and f + v f + m fused to f - a and f ― v f ― m and The triplet loss is used to constrain the fused features, i.e., [Number 78] and The objective function of shared semantic learning is modeled by combining three loss functions. L Syn =L syn a +L syn v +L syn m and L Syn By minimizing m and obtaining a fine-grained haptic signal based on the fine-grained haptic signal.

9. Inputting the fused features into a haptic generation network to reconstruct the haptic signal is A haptic generation network G(·) is constructed, and its structure is D h (·) and its network parameters θ hd The parameter θ of G(·) G and the initial value of The fused features containing the necessary semantic information are input to the haptic generation network G(·) to obtain the desired haptic signal h′, and the generated haptic signal h′ is input to E h (・) indicates the tactile feature f h′ and select the category center to map the tactile feature f h′ We apply semantic constraints to the [Number 79] is shown as however, [Number 80] and [Number 81] The feature is h and f h′ indicates similarity with [Number 82] ga f h′ The clustering loss of θ is calculated by optimizing the loss function. G and obtaining an optimal value of G(·), i.e., determining G(·).

10. Pre-configure the parameters of the haptic autoencoder, audio feature extraction network, and image feature extraction network. Presetting the parameters of the multimodal fusion mapping function and the parameters of the haptic generation network; Training the haptic autoencoder, the audio feature extraction network, the image feature extraction network, the multimodal fusion mapping function, and the haptic generation network includes step 1 and step 2; In step 1, θ v , θ a , θ he , θ hd , and M are preset, and {s i h }, {s i a }, {s i v }, and step 1 is Network parameter θ v , θ a , θ he , θ hd , node system tag {s i h }, {s i a }, {s i v } and category center matrix M are initialized, the number of clusters C, the learning rate μ 1、 and a step 11 of setting the number of iterations T; {s i h }, {s i a }, {s i v } and M are fixed, and θ v , θ a , θ he , θ hd Optimize, i.e., [Number 83] and where step 12 is where ∇ is the partial derivative of each loss function; θ v , θ a , θ he , θ hd , and M are fixed, and {s i h }, {s i a }, {s i v }, i.e., [Number 84] Step 13, where θ v , θ a , θ he , θ hd , and {s i h }, {s i a }, {s i v } and optimize M, i.e., [Number 85] Step 14, where If t<T, jump to step 412; if t=t+1, continue with the next iteration; otherwise, end the iteration; After T iterations, the optimal audio feature extraction network parameters θ a , the parameters θ of the image feature extraction network v , the parameters θ of the haptic autoencoder he , θ hd , node system tag {s i h }, {s i a }, {s i v }, and a cluster center vector matrix M of the haptic data; In step 2, θ is calculated based on the stochastic gradient descent method. m and θ G Step 2 is estimated as follows: θ m , learning rate μ 2 , μ 3 , iteration number n 1 Step 21 of initializing L Syn Based on this, we use stochastic gradient descent to find the m is estimated, i.e., [Number 86] Step 22, where L Gen Based on this, we use stochastic gradient descent to find the G Update, i.e., [Number 87] Step 23, where n<n 1 if n=n+1, jump to step 22, if n=n+1, continue with the next iteration, otherwise, end the iteration, step 24; n 1 After iterations, the optimal θ m and θ G 10. The method for reconstructing a fine-grained haptic signal for audiovisual aids according to claim 9, further comprising the step of:

Citation Information

Patent Citations

  • Cross-modal image generation method and device based on audio-tactile signal fusion

    CN113627482A

  • Audio and video auxiliary tactile signal reconstruction method based on cloud edge collaboration

    CN113642604A

  • Learning data generation program, learning data generation apparatus and learning data generation method

    JP2020079984A