A Fine-Grained Tactile Signal Reconstruction Method with Audio-Visual Aids
Through audio-visual-assisted cross-modal transfer learning and deep clustering algorithm, high-quality, fine-grained tactile signal reconstruction under weak supervision conditions is achieved, and the problems of signal degradation and loss in the existing technology are solved, meeting the needs of cross-modal services.
Patent Information
- Application Number
- CN202211446581.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-11-18
AI Technical Summary
The prior art has shortcomings in generating high-quality, fine-grained haptic signals, especially in the case of wireless communication unreliability and noise interference, making it difficult to achieve effective haptic signal reconstruction.
Through audio-visual assisted methods, cross-modal transfer learning and deep clustering algorithms are used to perform fine-grained classification and shared semantic learning to achieve high-quality reconstruction of tactile signals. The specific steps include: inputting the haptic signal to the haptic autoencoder for feature extraction, transfer learning transfers the feature extraction capability to the audio and image feature extraction network, joint constraint optimization of modal features, constructing a triple set for shared semantic learning, and finally input the haptic generation network for reconstruction.
High-quality, fine-grained tactile signal reconstruction under weakly supervised and weakly paired data sets is achieved, and the problems of signal degradation and loss in the prior art are overcome, and the needs of cross-modal services are met.
Smart Images

Figure CN115905838B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tactile signal generation, in particular to a fine-grained tactile signal reconstruction method assisted by audiovisual means. Background Art
[0002] With the maturity of the technologies related to traditional multimedia applications, while people's audiovisual needs have been greatly satisfied, they have begun to pursue more-dimensional and higher-level sensory experiences. And tactile information has gradually been integrated into existing audio-visual multimedia services to form multi-modal services, which are expected to bring a more extreme and rich interactive experience. Cross-modal communication technology has been proposed to support cross-modal services. Although it is effective to a certain extent in ensuring the quality of multi-modal streams, there are still some technical challenges when applying cross-modal communication to multi-modal services dominated by touch. First of all, the tactile stream is very sensitive to interference and noise in the wireless link, resulting in the degradation or even loss of tactile signals at the receiving end. Especially in remote operation application scenarios, such as remote industrial control, remote surgery, etc., this problem is serious and inevitable. Secondly, service providers do not have tactile acquisition devices, but users need tactile perception. Especially in virtual interaction application scenarios, such as online immersive shopping, holographic museum guides, virtual interactive movies, etc., users have extremely high requirements for tactile senses, which requires the ability to generate "virtual" touch sensations or tactile signals based on video and audio signals.
[0003] Currently, for tactile signals damaged or partially missing due to the unreliability of wireless communication and communication noise interference, self-recovery can be carried out from two aspects. The first category is based on traditional signal processing techniques. It finds a specific signal with the most similar structure by using sparse representation, and then uses it to estimate the missing part of the damaged signal. The second is to mine and utilize the spatio-temporal correlation of the signal itself to achieve in-modal self-repair and reconstruction. However, when the tactile signal is severely damaged or even does not exist, the in-modal reconstruction scheme will fail.
[0004] In recent years, some studies have focused on the correlations between different modalities and achieved cross-modal reconstruction based on this. Li et al. proposed in the literature "Learning cross-modal visual-tactile representation using ensembled generative adversarial networks" to obtain the required category information using image features and then use it together with noise as the input of the generative adversarial network to generate the corresponding category of tactile spectrogram. This method does not fully explore the semantic correlations between modalities, and the information provided by the obtained categories is limited. Therefore, the generated results are often not precise enough. Kuniyuki Takahashi et al. extended an encoder-decoder network in the literature "Deep Visuo-Tactile Learning: Estimation of Tactile Properties from Images", embedding both visual and tactile properties into the latent space and focusing on the degree of material tactile properties represented by the latent variables. Further, Matthew Purr et al. proposed a cross-modal learning framework with adversarial learning and cross-domain joint classification to estimate tactile physical properties from a single image in the literature "Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces From Images". Although these methods utilize the semantic information of modalities, they do not generate complete tactile signals, which has no practical significance for cross-modal services.
[0005] The above existing cross-modal generation methods also have the following defects: The training of their models all depends on large-scale training data to ensure the effectiveness of the models. In addition, they all only utilize the information of a single modality. In fact, the advantages of a single modality cannot provide us with enough information. When different modalities jointly describe the same semantics, they may contain unequal amounts of information. The complementarity and enhancement of information between modalities will help improve the generation effect. In actual application scenarios, the annotation cost of large-scale datasets is huge. What is relatively easy to obtain is often the coarse-grained large categories, while the fine-grained categories are not clear. In addition, there is no explicit correspondence between samples of different modalities, presenting difficult problems of weak supervision and weak pairing. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a method for fine-grained tactile signal reconstruction assisted by audio-visual. First, clustering analysis is performed within the coarse-grained categories to obtain the fine-grained classification of the samples. Then, modal common semantic constraints are imposed under the fine-grained categories, aiming to minimize the intra-class differences and maximize the inter-class differences. Finally, by searching for audio-visual samples that are positively matched with the tactile signals within the fine-grained sub-categories for correlation constraints, high-quality and fine-grained reconstruction of the tactile signals is achieved.
[0007] To solve the above technical problem, the present invention proposes a method for fine-grained tactile signal reconstruction assisted by audio-visual, comprising the following steps:
[0008] Input the tactile signals into a tactile autoencoder, and perform feature extraction on the tactile signals through a clustering task;
[0009] Through a cross-modal transfer learning method, transfer the feature extraction ability of the tactile autoencoder to an audio feature extraction network and an image feature extraction network, and optimize them;
[0010] By jointly considering clustering constraints, center constraints, and ranking constraints, impose constraints on the extracted tactile, audio, and image modal features, so that the modal features belonging to the same semantics are close to each other, while the modal features not belonging to the same semantics are separated, and obtain tactile, audio, and image modal features with fine-grained classification;
[0011] Based on the tactile, audio, and image modal features, construct a triple set, perform shared semantic learning with triple constraints, and optimize the multi-modal fusion mapping function to obtain fusion features containing shared semantic information;
[0012] Pre-set a tactile generation network, and input the fusion features into the tactile generation network to reconstruct the tactile signals.
[0013] Further, the pre-set tactile generation network inputs the fusion features into the tactile generation network to reconstruct the tactile signals, including:
[0014] Pre-set the parameters of the tactile autoencoder, audio feature extraction network, and image feature extraction network;
[0015] Pre-set the parameters of the multi-modal fusion mapping function and the parameters of the tactile generation network;
[0016] Train the tactile autoencoder, audio feature extraction network, image feature extraction network, multi-modal fusion mapping function, and tactile generation network;
[0017] The newly received image signal and audio signal are respectively input into the trained image feature extraction network and audio feature network to obtain image features and audio features; then the above features are input into the multimodal fusion mapping function to obtain fusion features; finally, the fusion features are input into the trained tactile generation network to obtain the reconstructed tactile signal.
[0018] Furthermore, the tactile signal is input into the tactile autoencoder, and through a clustering task, feature extraction of the tactile signal is performed, including:
[0019] The tactile signal is input into the tactile autoencoder for learning to extract the corresponding tactile features, and clustering based on the K-means algorithm is performed on the tactile signal based on the tactile features:
[0020] h is the input tactile signal, h = {h i} i=1,…,N , i represents the sorting subscript of the input tactile signal, N is the total amount of input tactile signals. After passing through the encoding module E h (·) of the autoencoder, f i h = E h (h i ; θ he ) is the feature representation of the tactile signal h i , f h = {f i h} i=1,…,N , θ he are the parameters of the encoding module; f i h is input into the decoding module D h (·) to obtain the output tactile signal where θ hd are the parameters of the decoding module; in addition, clustering based on the K-means algorithm is performed on the feature f i h , and the corresponding class label output is The parameters in the above process are jointly measured, and the loss function is designed as follows:
[0021]
[0022]
[0023] Among them, is the reconstruction error of the encoder, and N is the number of tactile signals; is the clustering error of K-means. M is the matrix of cluster center vectors of the tactile data obtained by the K-means algorithm. The c-th column m c in the M matrix represents the centroid of the c-th cluster; θ h= [θ he , θ hd are the parameters of the encoder module and the decoder module; is the j-th element of, if the element value is 1 and other elements are 0, it indicates that the corresponding original tactile signal h i belongs to the j-th class; l is the least squares loss λ is the regularization parameter, λ ≥ 0;
[0024] By minimizing θ is measured h , and f i h and
[0025] Further, through the cross-modal transfer learning method, the feature extraction ability of the tactile autoencoder is transferred to the audio feature extraction network and the image feature extraction network, including:
[0026] Adopt the feature adaptation method to minimize the maximum mean discrepancy criterion between the tactile domain and the audiovisual domain to achieve the transfer:
[0027] The distributions of the tactile, audio, and image signal sets are P, Q, and R respectively; among them, the MMD between the tactile signal and the audio signal is denoted as MMD k (P, Q), and the MMD between the tactile signal and the visual signal is MMD k (P, R); in the reproducing kernel Hilbert space H k in, H k includes the function set f defined on the non-empty set; the square of MMD is:
[0028]
[0029]
[0030] Among them, the tactile, audio, and image signals pass through their respective feature extraction networks φ to obtain the extracted feature vectors; among them, the tactile feature vector is denoted as φ h (h; θ he ) is the output of the encoding module of the autoencoder, and the audio and image feature vectors φ a (a; θ a ) = f a , φ v (v; θ v ) = f v , and the feature sets of the three modalities are represented as (f h , f a , f v); θ a and θ v are the parameters of the audio and image feature extraction networks respectively, and θ he is the parameter of the encoding module; for any function f ∈ H k , and any X ∈ P, μ k (P) is the average embedding of P in H k , that is, an element representation of the distribution P in the H k space, and f(X)_ represents that X is mapped to H k space through the function f, is the inner product operation; similarly μ k (Q) is the average embedding of Q in H k , μ k (R) is the average embedding of R in H k .
[0031] Calculate the corresponding and values as the loss function for cross-modal transfer. The specific formula is as follows:
[0032]
[0033] By optimizing L CT , guide the information flow between the tactile feature extraction autoencoder model and the audio and image feature extraction networks, and effectively transfer the feature extraction ability of the autoencoder for the tactile modality to the audio and image feature extraction networks, that is, measure θ a and θ v .
[0034] Furthermore, further optimize the audio feature extraction network and the image feature extraction network, including:
[0035] The classification loss function is as follows:
[0036]
[0037] Among them, and are the class labels of the audio signal and the image signal respectively. By minimizing , further obtain the optimal values of θ a and θ v , where p(f i a ; θ a ) means that when the input of the audio feature extraction network is f i a , the network parameter is θ aThe probability of obtaining the audio signal category ; p(f i v ; θ v ) means that when the input of the image feature extraction network is f i v , and the network parameter is θ v , the probability of obtaining the image signal category ; f a = {f i a} i=1,…,N , f v = {f i v} i=1,…,N ;
[0038] Combine L CT and to obtain the total objective function:
[0039]
[0040] By minimizing L clu , the optimal θ h , θ a and θ v can be obtained; after the parameters are determined, the features f h , f a , f v corresponding to the tactile signal, audio signal and image signal can be obtained.
[0041] Furthermore, by jointly considering the clustering constraint, center constraint, and sorting constraint to constrain the extracted tactile, audio, and image modality features, the modality features belonging to the same semantics are made close, while the modality features not belonging to the same semantics are separated, and the tactile, audio, and image modality features with fine-grained classification are obtained, including:
[0042] Perform clustering learning under the center constraint on the three modality signals to ensure the compactness between the features of the same fine-grained subcategory; to achieve better fine-grained classification performance, the features of the same subclass should be adjacent in the common space, which is to minimize the within-class variance; the clustering learning is driven by minimizing the distance from the feature to the center of its subclass. Among them, the center of the subclass of the tactile signal is used as the common subclass center of the cross-modal signals, so as to ensure the compactness between the semantic features of the cross-modal signals of the same category; the loss function of the clustering learning under the center constraint is defined as follows:
[0043]
[0044] where N is the number of tactile, audio, and image signals; and represent the categories of tactile signals, audio signals, and image signals respectively; the audio signals and image signals share the clustering center vector matrix M of the tactile signals; through the above process, the modal features f i a , f i a , and f i h are close to each other.
[0045] Furthermore, clustering learning under sorting constraints is performed on the three modal signals to ensure that the features of different fine-grained sub-categories have a certain degree of sparsity; the goal of the center constraint is to minimize the within-class variance, while the goal of the sorting constraint is to maximize the between-class variance, so that the feature outputs of different sub-categories are more dissimilar than those of the same sub-category; the definition of the sorting constraint is as follows:
[0046]
[0047] where C is the total number of categories after the tactile signals are clustered by the K-means algorithm; through the above process, the modal features f i a , f i a , and f i h are closer, while the modal features with different semantics are separated as much as possible.
[0048] Furthermore, a triple set is constructed based on the tactile, audio, and image modal features, and shared semantic learning with triple constraints is performed to optimize the multi-modal fusion mapping function, obtaining fusion features containing shared semantic information, including:
[0049] Randomly select a tactile signal sample h i in a certain fine-grained sub-category, and use this sample as the anchor;
[0050] Select a sample in the image dataset that i belongs to the same category as h i h and has the closest semantic feature to f as the positive matching sample, and select a sample that i does not belong to the same category as h i h and has the closest semantic feature to f as the negative matching sample;
[0051] Thus, a triple set is formed for the samples in the dataset
[0052] Similarly, a set of triples composed of tactile signal samples and audio signal samples is obtained.
[0053] By minimizing the semantic feature f of the anchor h i and the semantic feature of the positive match within the corresponding subclass in the audio and image modalities i h The distance between them is maximized, and the semantic feature between f and the semantic feature of the negative match is maximized, and there is a minimum interval δ. The two triple loss functions obtained are as follows: i h and the semantic feature of the negative match The semantic feature between them, and there is a minimum interval δ. The two triple loss functions obtained are as follows:
[0054]
[0055]
[0056] Introduce the joint normal form to perform deep fusion on multi-modal features:
[0057] Fuse f a and f v The process is as follows:
[0058] f m = F m (f a , f v ; θ m ),
[0059] where f m is the output of multi-modal fusion in the shared semantic subspace, that is, the fusion feature; F m (·) is the multi-modal fusion mapping function with parameters θ m , and F m (·) takes the linear weighted sum of f a and f v ;
[0060] Fuse and into Fuse and into
[0061] Use the triple loss to constrain the fusion feature:
[0062]
[0063] The objective function of shared semantic learning is modeled by combining three loss functions, expressed as:
[0064]
[0065] By minimizing L Syn , the optimal θ is obtained m .
[0066] Furthermore, construct a tactile generation network: build a tactile generation network G(·), whose structure is the same as that of D h (·), and use its network parameter θ hd as the initial value of the parameter θ G of G(·);
[0067] Input the fusion feature containing the required semantic information into the tactile generation network G(·) to obtain the desired tactile signal h′, and remap the generated tactile signal h′ through E h (·) to the tactile feature f h′ , and select the class center to perform semantic constraint on it. The final loss function is expressed as:
[0068]
[0069] where represents the similarity between the feature f h and f h′ , is the clustering loss of f h′ , and they jointly serve as the regularization term of the loss function; by optimizing this loss function, the optimal value of θ G is obtained, that is, G(·) is determined.
[0070] Furthermore, preset the parameters of the tactile autoencoder, audio feature extraction network, and image feature extraction network;
[0071] Preset the parameters of the multimodal fusion mapping function and the parameters of the tactile generation network;
[0072] Train the tactile autoencoder, audio feature extraction network, image feature extraction network, multimodal fusion mapping function, and tactile generation network, including:
[0073] Step 1: Preset θ v , θ a , θ he , θ hd and M, and optimize
[0074] Step 11. Initialize the network parameters θ v , θ a , θ he , θ hd ; node class label category center matrix M; set the number of clustering clusters C, learning rate μ1, and number of iterations T;
[0075] Step 12, fix and M, and optimize θ based on the stochastic gradient descent method v , θ a , θ he , θ hd :
[0076]
[0077]
[0078]
[0079]
[0080] where is to take the partial derivative of each loss function;
[0081] Step 13, fix θ v , θ a , θ he , θ hd , and M, and optimize
[0082]
[0083]
[0084]
[0085] Step 14, fix θ v , θ a , θ he , θ hd , and optimize M.
[0086]
[0087] Step 15, if t < T, then jump to Step 412, t = t + 1, and continue the next iteration; otherwise, terminate the iteration;
[0088] Step 16, after T rounds of iteration, obtain the parameters θ of the optimal audio feature extraction network a , the parameters θ of the image feature extraction network v , the parameters θ of the tactile autoencoder he , θ hd , the node class labels and the clustering center vector matrix M of the tactile data;
[0089] Step 2: Measure θ based on the stochastic gradient descent method mand θ G :
[0090] Step 21: Initialize θ m ; learning rates μ2, μ3, number of iterations n1;
[0091] Step 22: Estimate θ according to L Syn using the stochastic gradient descent method m :
[0092]
[0093] Step 23: Update θ according to L Gen using the stochastic gradient descent method G :
[0094]
[0095] Step 24: If n < n1, jump to Step 22, n = n + 1, and continue the next iteration; otherwise, terminate the iteration;
[0096] Step 25: After n1 rounds of iteration, obtain the optimal θ m and θ G .
[0097] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0098] With the help of the deep clustering algorithm for cross-modal transfer, after learning the fine-grained classification of three-modal samples and performing shared semantic learning, the present invention gives full play to the advantages of multi-modal feature fusion, and finally realizes the generation of fine-grained tactile signals based on clustering constraints; makes the best use of the existing dataset with weak supervision and weak matching problems; and generates high-quality and fine-grained tactile signals, which better meet the requirements of cross-modal services. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] Figure 1 is a flowchart of a method for reconstructing fine-grained tactile signals assisted by audio-visual of the present invention.
[0100] Figure 2 is a schematic diagram of the complete network structure of the present invention.
[0101] Figure 3 is a schematic diagram of the architecture of a deep clustering model based on cross-modal transfer of the present invention.
[0102] Figure 4 is a diagram of the tactile signal reconstruction results of the present invention and other comparison methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0103] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0104] The present invention provides a method for reconstructing fine-grained tactile signals with audio-visual assistance. The flowchart of the method is as Figure 1 shown, and the method includes the following steps:
[0105] Step 1: First, input the tactile signal into a tactile autoencoder, and through clustering, achieve feature extraction of the tactile signal; then, use cross-modal transfer learning technology to transfer the feature extraction ability of the tactile autoencoder into the audio feature extraction network and the image feature extraction network respectively, and optimize them. Then, by jointly considering clustering constraints, center constraints, and ranking constraints, further constrain the extracted tactile, audio, and image modality features, so that the modality features belonging to the same semantics are close, while the modality features not belonging to the same semantics are separated, obtaining tactile, audio, and image modality features with fine-grained classification;
[0106] (1-1) First, perform feature learning with clustering constraints on the tactile, image, and audio modality signals to obtain distinguishing features with fine-grained subcategories. This can be further divided into three steps:
[0107] (1-1-1) The first step: First, input the tactile signal into the autoencoder for learning, extract the corresponding tactile features, and perform clustering on the tactile signal based on the K-means algorithm based on the tactile features:
[0108] Specifically, assume h i is the input tactile signal, h = {h i} i=1,…,N , i represents the sorting subscript of the input tactile signal, N is the total amount of input tactile signals. After passing through the encoding module E h (·) of the autoencoder, f i h = E h (h i ; θ he ) is the feature representation of the tactile signal h i , f h = {f i h} i=1,…,N , θ he is the parameter of the encoding module; input the feature f i h into the decoding module D h (·), and obtain the output tactile signal where θ hd is the parameter of the decoding module; in addition, for the feature f i hPerform clustering based on the K-means algorithm and output the corresponding class labels as Jointly measure the parameters in the above process and design the loss function as follows:
[0109]
[0110]
[0111] Among them, is the reconstruction error of the encoder, N is the number of tactile signals, is the clustering error of K-means, M is the matrix of cluster center vectors of tactile data obtained by the K-means algorithm, and the c-th column m in the M matrix c represents the centroid of the c-th cluster. θ h =[θ he ,θ hd are the parameters of the encoder module and the decoder module; is the j-th element of, if the element value is 1 and other elements are 0, it indicates that the corresponding original tactile signal h i belongs to the j-th class; l is the least squares loss λ≥0 is the regularization parameter;
[0112] By minimizing θ h can be measured, and f i h and
[0113] (1-1-2) Step 2: Through cross-modal transfer learning technology, transfer the obtained tactile features to the feature extraction processes of image signals and audio signals. That is, adopt the feature adaptation method to minimize the maximum mean discrepancy (MMD) criterion between the tactile domain and the audiovisual domain to achieve the transfer.
[0114] Specifically, let the distributions of the tactile, audio, and image signal sets be P, Q, and R respectively. Among them, the MMD between the tactile signal and the audio signal is denoted as MMD k (P,Q), and the MMD between the tactile signal and the visual signal is MMD k (P,R). In the reproducing kernel Hilbert space H k in, H k includes the function set f defined on the non-empty set, and the square of the MMD is:
[0115]
[0116]
[0117] Among them, the tactile, audio, and image signals pass through their respective feature extraction networks φ to obtain the extracted feature vectors. Among them, the tactile feature vector can be expressed as φ h (h; θ he ), which is the output of the encoding module of the autoencoder in the previous step. The audio and image feature vectors φ a (a; θ a ) = f a , φ v (v; θ v ) = f v . The feature sets of the three modalities can be expressed as (f h , f a , f v ). θ a and θ v are the parameters of the audio and image feature extraction networks respectively, and θ he is the parameter of the encoding module. For any function f ∈ H k , and any X ∈ P, μ k (P) is the average embedding of P in H k , that is, an element representation of the distribution P in the H k space. f(X) represents that X is mapped to the H k space through the function f. is the inner product operation; similarly μ k (Q) is the average embedding of Q in H k , μ k (R) is the average embedding of R in H k .
[0118] Calculate the corresponding and values as the loss function for cross-modal transfer. The specific formula is as follows:
[0119]
[0120] By optimizing L CT , the information flow between the tactile feature extraction autoencoder model and the audio and image feature extraction networks can be guided, so that the feature extraction ability of the autoencoder for the tactile modality can be effectively transferred to the audio and image feature extraction networks, that is, θ a and θ v can be estimated.
[0121] (1-1-3) In the third step, the feature extraction networks for audio and images are further optimized. Here, the video feature extraction network adopts the design style of the VGG network, that is, it has a 3×3 convolutional filter and a 2×2 max pooling layer with a stride of 2 and no padding; the network is divided into four blocks, each block contains two convolutional layers and one pooling layer, and the number of filters between consecutive blocks doubles; finally, max pooling is performed at all spatial positions to generate a single 512-dimensional semantic feature vector. Then it is input into a three-layer fully connected neural network (256-128-32) and a fully connected layer with K nodes and a softmax function, where the 32-dimensional vector is the feature vector of the visual signal, and it will also be trained by subsequent shared semantic learning based on triplet constraints. K is the number of fine-grained subcategories under each coarse-grained category. In the dataset of this experiment, K is taken as 3. The relevant network structure settings for the audio signal are the same as those for the visual signal, and they share the classifier with the visual signal.
[0122] The designed classification loss function is as follows:
[0123]
[0124] Among them, and are the class labels of the audio signal and the image signal respectively. By minimizing , the optimal values of θ a and θ v can be further obtained, where p(f i a ; θ a ) means the probability of obtaining the audio signal class i a when the input of the audio feature extraction network is f a and the network parameters are θ ; p(f i v ; θ v ) means the probability of obtaining the image signal class i v when the input of the image feature extraction network is f v and the network parameters are θ .
[0125] It should be particularly noted that: the acquisition of this label adopts a progressive strategy: first, pseudo-labels are assigned to the data, and these data are used to optimize the model. The classification ability of the optimized model for audio-visual signals will be enhanced. Further, the trained model is used to perform the pseudo-label operation, thereby updating the pseudo-labels. Through such a progressive optimization method, the fine-grained classification ability of the network will be gradually improved.
[0126] In summary, the overall objective function of step (1-1) is a combination of the above three loss functions and can be expressed as:
[0127]
[0128] By minimizing L clu , the optimal θ h , θ a and θ v can be obtained; after the parameters are determined, the features f h , f a , f v corresponding to the tactile signal, audio signal, and image signal can be obtained;
[0129] (1-2) Perform clustering learning under central constraint on the three-modal signals to ensure the compactness between the features of the same fine-grained subcategory. To achieve better fine-grained classification performance, the features of the same subclass should be adjacent in the common space, which is to minimize the within-class variance. Specifically, the clustering learning is driven by minimizing the distance from the feature to the center of its subclass. Among them, the center of the subclass of the tactile signal is used as the common subclass center of the cross-modal signals, so as to ensure the compactness between the semantic features of the cross-modal signals of the same category. The loss function of the clustering learning under central constraint is defined as follows:
[0130]
[0131] where N is the number of tactile, audio, and image signals, and represent the categories of the tactile signal, audio signal, and image signal respectively; it should be noted that the clustering center vector matrix M of the tactile data is shared by the audio data and the image data. By minimizing L cen , the problem of large within-class differences in fine-grained classification can be effectively solved. Through the above process, the modal features f i a , f i a , and f i h with similar semantics are close to each other.
[0132] (1-3) Perform clustering learning under sorting constraint on the three-modal data to ensure a certain sparsity of the features of different fine-grained subcategories. The goal of the central constraint is to minimize the within-class variance, while the goal of the sorting constraint is to maximize the between-class variance, so that the feature outputs of different subclasses are more dissimilar than those of the same subclass. The sorting constraint is defined as follows:
[0133]
[0134] Among them, C is the total number of categories after the tactile signals are clustered by the K-means algorithm, that is, the number of clustering clusters. By minimizing L rank , the problem of small inter-class differences in fine-grained classification can be effectively solved. Through the above process, each modal feature f i a , f i a , and f i h that have similar semantics are further approximated, while making the modal features with different semantics as separated as possible.
[0135] Step 2: After obtaining the tactile, audio, and image features with fine-grained classification, construct a triple set from the three modal features, perform shared semantic learning with triple constraints, optimize the multi-modal fusion mapping function, and obtain the fusion features containing shared semantic information, laying a foundation for generating tactile signals.
[0136] The specific steps of Step 2 are as follows:
[0137] (2-1) Randomly select a tactile signal sample h i within a certain fine-grained subclass, use this sample as the anchor, and then select a sample in the image dataset that i belongs to the same class as h i h and has the closest semantic feature to f as the positive matching sample. Then select another sample that i does not belong to the same class as h i h and has the closest semantic feature to f as the negative matching sample. Thus, a triple set is formed for all samples in the dataset Similarly, a triple set composed of tactile signal samples and audio signal samples can be obtained By minimizing the distance between the semantic feature f i of the anchor point h i h and the semantic features of the positive matches within the corresponding subclasses in the audio and image modalities , and maximizing the semantic features between f i h and the negative matches , and there should be a minimum interval δ. Here, δ is taken as 1. The two resulting triple loss functions are as follows:
[0138]
[0139]
[0140] (2-2) On this basis, the joint paradigm is introduced to deeply fuse multi-modal features. Specifically, first, the audio-visual data respectively pass through the audio feature extraction network and the image feature extraction network to obtain features f a and f v . After that, f a and f v are fused, and the process is as follows:
[0141] f m = F m (f a , f v ; θ m )
[0142] where f m is the output of multi-modal fusion in the shared semantic subspace, that is, the fused feature; F m (·) is a mapping function with parameters θ m . Generally speaking, F m (·) takes the linear weighting of f a and f v .
[0143] Fuse and into Fuse and into Similarly, the fused feature is constrained by the triplet loss:
[0144]
[0145] (2-3) The objective function of shared semantic learning can be modeled by combining three loss functions, expressed as:
[0146]
[0147] By minimizing L Syn , the optimal θ m is obtained, thus laying a foundation for the generation of tactile signals in the next stage.
[0148] Step 3: Input the fused feature into the tactile generation network to generate the desired tactile signal h'.
[0149] Step 3 is specifically as follows:
[0150] First, construct the tactile generation network. Since the tactile decoder D h (·) was trained as part of a complete autoencoder in step (1), a separate tactile generation network G(·) is built here, and its structure is the same as that of D h(·) are the same (32 - 128 - 256 - Z), and its network parameter θ hd as the parameter θ of G(·) G as the initial value. Input the fused feature containing the required semantic information into the haptic generation network G(·) to obtain the desired haptic signal h′, and remap the generated haptic signal h′ through E h (·) to the 32 - dimensional haptic feature f h′ , and select the class center to perform semantic constraint on it. Obviously, the distance between it and the class center of the corresponding class should be as small as possible, while the distance between it and the class centers of other classes should be as large as possible. The final loss function can be expressed as:
[0151]
[0152] where represents the similarity between the feature f h and f h′ , is the clustering loss of f h′ , and they jointly serve as the regularization terms of the loss function. By optimizing this loss function, the optimal value of θ G can be obtained, that is, G(·) is determined.
[0153] Step (4), training of the above - mentioned model: Divide the model training into two steps. In the first step, estimate the parameters in the haptic auto - encoder, audio feature extraction network, and image feature extraction network; in the second step, estimate the parameters of the multi - modal fusion mapping function and the parameters of the haptic generation network. Through this step, train the haptic auto - encoder, audio feature extraction network, image feature extraction network, multi - modal fusion mapping function, and haptic generation network.
[0154] Step 4 is specifically as follows: Complete the estimation of θ v , θ a , θ he , θ hd and M, and optimize
[0155] Step 411. Initialize the network parameters θ v , θ a , θ he , θ hd ; the node class label the category center matrix M; set the number of clustering clusters C, learning rate μ1 = 0.0001, and the number of iterations T = 600;
[0156] Step 412. Fix and M, and optimize θ v , θ a , θ he , θ hd :
[0157]
[0158]
[0159]
[0160]
[0161] Among them, is to take the partial derivative of each loss function;
[0162] Step 413, fix θ v , θ a , θ he , θ hd , and M, and optimize
[0163]
[0164]
[0165]
[0166] Step 414, fix θ v , θ a , θ he , θ hd , and Optimize M
[0167]
[0168] Step 415, if t < T, then jump to Step 412, t = t + 1, and continue the next iteration; otherwise, terminate the iteration;
[0169] Step 416, after T rounds of iteration, after T rounds of iteration, obtain the optimal audio feature extraction network parameters θ a , image feature extraction network θ v , tactile autoencoder parameters θ he , θ hd , node class labels and the clustering center vector matrix M of tactile data.
[0170] Specifically, when updating m k it does not simply use where is the index set assigned to cluster k from the first sample to the current sample, but the historical data that has appeared is not sufficient to represent the global cluster structure situation, and It may not be correct. Therefore, this algorithm assumes a premise that the number of data samples contained in each cluster is roughly balanced. Based on this, the above gradient update steps are designed to update m k , using to control the learning rate, where is the number of times the algorithm assigns the sample to cluster k before processing the i-th sample. In this way, the update of M can also be regarded as an SGD step.
[0171] (4-2), Based on SGD, complete the estimation of θ m and θ G .
[0172] Step 421, Initialize θ m ; The batch size batch = 64, the learning rates μ2, μ3 = 0.0001, and the number of iterations n1 = 600.
[0173] Step 422, According to L Syn Use the stochastic gradient descent method to fine-tune θ m ,:
[0174]
[0175] Step 423, According to L Gen Use the stochastic gradient descent method to update θ G :
[0176]
[0177] Step 424, If n < n1, then jump to Step 422, n = n + 1, and continue the next iteration; otherwise, terminate the iteration;
[0178] Step 425, After 600 rounds of iteration, obtain the optimal n = n + 1
[0179] Step (5), Input the newly received image signal and audio signal into the trained image feature extraction network and audio feature network respectively, and obtain the image feature and audio feature respectively; then input the above features into the multi-modal fusion mapping function to obtain the fusion feature; finally, input the fusion feature into the trained tactile generation network to obtain the reconstructed tactile signal.
[0180] Step (5) is specifically as follows:
[0181] Input the newly received image signal v and audio signal a into the trained image feature extraction network and audio feature extraction network respectively, and obtain the image feature and audio feature and input them into the trained multi-modal fusion mapping function to obtain the fusion feature Input Input the trained G(.), and finally obtain the reconstructed tactile signal
[0182] The following experimental results show that, compared with the existing methods, the present invention achieves better generation effects in tactile signal synthesis by using the complementary fusion of multimodal semantics.
[0183] This embodiment uses the LMT cross-modal dataset for experiments. This dataset was proposed in the literature "Multimodal feature-based surface material classification" and includes samples of nine semantic categories: mesh, stone, metal, wood, rubber, fiber, foam, foil and paper, textiles and fabrics. This embodiment selects five major categories (each major category contains three minor categories) for experiments. The LMT dataset is reorganized. First, the training set and test set of each material instance are combined to obtain 20 image samples, 20 audio signal samples and 20 tactile signal samples for each instance respectively. Then the data is augmented to train the neural network. Specifically, each image is flipped horizontally and vertically, rotated at any angle, and techniques such as random scaling, shearing and offset are used in addition to the traditional methods. Thus, the data of each category is extended to 100, so there are 1500 images in total, with a size of 224*224. In the dataset, 80% is selected for training, while the remaining 20% is used for testing and performance evaluation. In this experiment, it is default that the fine-grained categories are unknown.
[0184] (1) Clustering results: To verify the effectiveness of the clustering method proposed by the present invention, it is compared with a variety of baseline methods, including:
[0185] K-means (KM): Use the K-Means algorithm to cluster the samples of image, audio, and tactile modalities respectively.
[0186] Autoencoder + K-means (Autoencoder followed by K-means, AE+KM): This is a two-stage method. First, the signal samples of different modalities are reconstructed and learned to obtain the feature representations of each modality, and then K-means is used for clustering.
[0187] Triple Deep Clustering Network (3-DCN): Use the DCN model to cluster the signals of different modalities respectively.
[0188] The present invention: The method of this embodiment.
[0189] When presenting the results, the average values on multiple subclasses are selected, and three main metrics are used: Normalized Mutual Information (NMI), Adjusted Rand Index (ARI), and Clustering Accuracy (ACC). The experimental results are shown in Table 1 as follows:
[0190] Table 1 Performance comparison of the proposed clustering method with other competing methods
[0191]
[0192] Table 1 shows the results of applying the present invention, 3-DCN, AE+KM, and KM to the LMT dataset. It can be seen that the method of the present invention performs very competitively in this dataset, and the results are significantly better than those of traditional clustering algorithms and general deep clustering algorithms. The possible reason for the analysis is that the application scenarios of other algorithms are all single-modal, which is likely to cause unbalanced clustering results. Theoretically, the number of samples within a certain category of co-existing cross-modal data should be equal. In addition, the subclass features learned by this clustering method are significantly more compact and the differences between different classes are also greater, which will contribute to the subsequent tactile signal reconstruction.
[0193] (2) Tactile reconstruction results: Based on the determination of fine-grained categories, the proposed fine-grained tactile reconstruction method is compared with the following methods:
[0194] Existing method one: The Ensemble Generative Adversarial Networks (Ensembled GANs, abbreviated as E-GANs) in the literature "Learning cross-modal visual-tactile representation using ensembled generative adversarial networks" (authors X. Li, H. Liu, J. Zhou, and F. Sun) uses image features to obtain the required category information, and then takes it together with noise as the input of the generative adversarial network to generate the corresponding category of tactile spectrogram, and finally converts it into tactile signals.
[0195] Existing method two: The Deep Visio-Tactile Learning (abbreviated as DVTL) in the literature "Deep Visuo-Tactile Learning: Estimation of Tactile Properties from Images" (authors: Kuniyuki Takahashi and Jethro Tan) extends the traditional encoder-decoder network with latent variables and embeds visual and tactile attributes in the latent space.
[0196] Existing Method 3: In the literature "Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces From Images" (authors: Matthew Purri and Kristin Dana), a Joint-encoding-classification GAN (JEC-GAN for short) is proposed. It encodes instances of each modality into a shared intrinsic space through different encoding networks, and uses pairwise constraints to make the embedded visual samples and tactile samples close in the latent space. Finally, taking visual information as input, the corresponding tactile signal is reconstructed through the generation network.
[0197] This experiment is analyzed from both quantitative and qualitative perspectives. First, Table 2 shows the performance of tactile signal reconstruction under each method from multiple perspectives such as root mean square error (RMSE), structural similarity (SIM), and classification accuracy (ACC). Table 2 shows the experimental results of the present invention:
[0198] Table 2 Performance comparison of the proposed audiovisual-assisted fine-grained tactile reconstruction method with other competing methods
[0199]
[0200] From Table 2 and Figure 4 It can be seen that compared with the above-mentioned state-of-the-art methods, the method we proposed has obvious advantages for the following reasons: The tactile reconstruction signal method proposed in the present invention clarifies the fine-grained subclasses through a cross-modal clustering algorithm, and effectively improves the compactness and distinguishability of semantic features by using central constraints and ranking constraints; (2) When the correspondence of audiovisual-tactile samples is specified manually, the subjective awareness is strong and thus not precise enough. On the contrary, in the audiovisual-assisted fine-grained tactile signal reconstruction method proposed in the present invention, the audiovisual semantic features that are closest to the tactile semantic features are automatically selected by the model as the input of its generation network during training.
[0201] In other embodiments, the tactile encoder in step (1) of the present invention uses a feedforward neural network, which can be replaced by one-dimensional convolutional neural networks (1D-CNN for short).
[0202] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for reconstructing fine-grained tactile signals with audio-visual assistance, characterized in that It includes the following steps: Input the tactile signal into the tactile autoencoder, and through the clustering task, extract the features of the tactile signal; Through the cross-modal transfer learning method, transfer the feature extraction ability of the tactile autoencoder to the audio feature extraction network and the image feature extraction network, and optimize them; By jointly considering the clustering constraint, the center constraint, and the ranking constraint, constrain the extracted tactile, audio, and image modality features, so that the modality features belonging to the same semantics are close, while the modality features not belonging to the same semantics are separated, and obtain tactile, audio, and image modality features with fine-grained classification, including: performing clustering learning under the center constraint on the three modality signals to ensure the compactness between the features of the same fine-grained subcategory; in order to achieve better fine-grained classification performance, the features of the same subclass should be adjacent in the common space, which is to minimize the within-class variance; drive the clustering learning by minimizing the distance from the feature to the center of its subclass; where the center of the subclass of the tactile signal is used as the common subclass center of the cross-modal signals, so as to ensure the compactness between the semantic features of the cross-modal signals of the same category; performing clustering learning under the ranking constraint on the three modality signals to ensure a certain sparsity of the features of different fine-grained subcategories; the goal of the center constraint is to minimize the within-class variance, while the goal of the ranking constraint is to maximize the between-class variance, so that the feature outputs of different subclasses are more dissimilar than the feature outputs of the same subclass; The loss function of the clustering learning under the center constraint is defined as follows: ; Among them, N is the number of tactile, audio, and image signals; , and respectively represent the categories of tactile signals, audio signals, and image signals; the audio signals and image signals share the clustering center vector matrix of the tactile signals ; through the above process, the modal features , and with similar semantics are close to each other; The definition of the ranking constraint is as follows: ; Among them, is the total number of categories after the tactile signals are clustered by the K-means algorithm; through the above process, each modal feature with similar semantics , and are further close, and each modal feature with different semantics is separated as much as possible; Construct a triplet set based on the tactile, audio, and image modality features, perform shared semantic learning with triplet constraints, optimize the multi-modal fusion mapping function, and obtain the fusion features containing shared semantic information; Pre-set a tactile generation network, and input the fusion features into the tactile generation network to reconstruct the tactile signal.
2. The method for reconstructing a fine-grained tactile signal assisted by vision and hearing according to claim 1, wherein: The pre-set tactile generation network inputs the fusion features into the tactile generation network to reconstruct the tactile signal, including: Pre-set the parameters of the tactile autoencoder, the audio feature extraction network, and the image feature extraction network; Pre-set the parameters of the multi-modal fusion mapping function and the parameters of the tactile generation network; Train the tactile autoencoder, the audio feature extraction network, the image feature extraction network, the multi-modal fusion mapping function, and the tactile generation network; Input the newly received image signal and audio signal into the trained image feature extraction network and audio feature network respectively, and obtain the image feature and audio feature respectively; then input the above features into the multi-modal fusion mapping function to obtain the fusion features; finally, input the fusion features into the trained tactile generation network to obtain the reconstructed tactile signal.
3. The method for reconstructing a fine-grained tactile signal assisted by vision and hearing according to claim 2, wherein: Input the tactile signal into the tactile autoencoder, and through the clustering task, extract the features of the tactile signal, including: Input the tactile signal into the tactile autoencoder for learning, extract the corresponding tactile features, and perform clustering on the tactile signal based on the tactile features using the K-means algorithm: is the input tactile signal, , where i represents the sorting subscript of the input tactile signal, N is the total amount of the input tactile signals. After passing through the encoding module of the autoencoder , is the feature representation of the tactile signal ; , are the parameters of the encoding module; Input into the decoding module , and the output tactile signal is obtained, where are the parameters of the decoding module; In addition, perform clustering on the feature based on the K-means algorithm, and the output corresponding class label is ; Jointly measure the parameters in the above process, and design the loss function as follows: ; ; Among them, is the reconstruction error of the encoder, N is the number of tactile signals; is the clustering error of K-means, is the clustering center vector matrix of tactile data obtained by the K-means algorithm, In the matrix, the th column represents the centroid of the th cluster; are the parameters of the encoder module and the decoder module; is the th element of . If the value of the element is 1 and the other elements are 0, it indicates that the original tactile signal corresponding to belongs to the th class; is the least squares loss ; is the regularization parameter, ; By minimizing , it is measured that , and and are obtained.
4. A fine-grained tactile signal reconstruction method assisted by audio-visual, according to claim 3, characterized in that: Through the cross-modal transfer learning method, transfer the feature extraction ability of the tactile autoencoder to the audio feature extraction network and the image feature extraction network, including: Adopt the feature adaptation method to minimize the maximum mean difference criterion between the tactile domain and the audio-visual domain to achieve the transfer: The distributions of the tactile, audio, and image signal sets are respectively , and ; among them, the MMD between the tactile signal and the audio signal is denoted as , and the MMD between the tactile signal and the visual signal is ; in the reproducing kernel Hilbert space , includes a set of functions defined on a non-empty set ; the square of the MMD is: ; ; Among them, the tactile, audio, and image signals pass through their respective feature extraction networks to obtain the extracted feature vectors; among them, the tactile feature vector is represented as which is the output of the encoding module of the autoencoder, and the audio and image feature vectors , , and the feature sets of the three modalities are represented as ; and are the parameters of the audio and image feature extraction networks respectively, is the parameter of the encoding module; for any function , and any , is 's average embedding in , that is, an element representation of the distribution P in the space, represents mapped to the space through the function , is the inner product operation; similarly , is 's average embedding in , , is 's average embedding in ; Calculate the corresponding and values as the loss function for cross-modal transfer. The specific formula is as follows: ; By optimizing , guiding the information flow between the tactile feature extraction autoencoder model and the audio and image feature extraction networks, effectively transferring the feature extraction ability of the autoencoder for the tactile modality to the audio and image feature extraction networks, that is,[ and are measured.[ 5. A fine-grained tactile signal reconstruction method assisted by audio-visual, according to claim 4, characterized in that: Further optimize the audio feature extraction network and the image feature extraction network, including: The classification loss function is as follows: ; Among them, and are the class labels of the audio signal and the image signal respectively. By minimizing , the optimal values of and are further obtained, where the meaning of is the probability of obtaining the audio signal class when the input of the audio feature extraction network is and the network parameters are ; the meaning of is the probability of obtaining the image signal class when the input of the image feature extraction network is and the network parameters are ; , ; Combine 、 and to obtain the overall objective function: ; By minimizing , the optimal , and can be obtained; after the parameters are determined, the features corresponding to the tactile signal, audio signal, and image signal can be obtained .
6. A fine-grained tactile signal reconstruction method assisted by audio-visual, according to claim 1, characterized in that: Construct a triple set based on tactile, audio, and image modal features, perform triple-constrained shared semantic learning, and optimize the multi-modal fusion mapping function to obtain fusion features containing shared semantic information, including: Randomly select a haptic signal sample within a certain fine-grained subclass , and use this sample as an anchor; Select one in the image dataset that is in the same class as and has the closest semantic feature distance to as the positive matching sample. Select one that is not in the same class as and has the closest semantic feature distance to as the negative matching sample; Select the sample with the closest semantic feature distance as the negative matching sample. as the negative matching sample; Thus, a triple set is formed by the samples in the data set ; Similarly, a set of triples composed of tactile signal samples and audio signal samples is obtained ; By minimizing the semantic features of the anchor and the semantic features of the positive matches within the corresponding subcategories in the audio and image modalities to obtain the distance between them, and maximizing the distance between and the semantic features of the negative matches while having a minimum interval between them, the resulting two triple loss functions are as follows: ; ; Introduce the joint paradigm to deeply fuse the multi-modal features: Fuse and The process is as follows: ; Among them, is the output of multimodal fusion in the shared semantic subspace, that is, the fused feature is the multimodal fusion mapping function with parameters and takes and for linear weighting; Combine and into , combine and into ; Use the triple loss to constrain the fusion features: ; The objective function of the shared semantic learning is modeled by combining three loss functions, expressed as: ; By minimizing to obtain the optimal .
7. A fine-grained tactile signal reconstruction method assisted by audio-visual, according to claim 6, characterized in that: Input the fusion features into the tactile generation network to reconstruct the tactile signal, including: Construct a haptic generation network: Build a haptic generation network , whose structure is the same as , and set its network parameters as the parameters initial values; Input the fused features containing the required semantic information into the haptic generation network Obtain the desired haptic signals and remap the generated haptic signals through to haptic features Select the class center to perform semantic constraints on them, and the final loss function is expressed as: ; Among them, , represents the similarity of features and , is 's clustering loss, and they jointly serve as the regularization term of the loss function; by optimizing this loss function, the optimal value of is obtained, that is, to determine .
8. A fine-grained tactile signal reconstruction method assisted by audio-visual, according to claim 7, characterized in that: Preset the parameters of the tactile autoencoder, the audio feature extraction network, and the image feature extraction network; Preset the parameters of the multi-modal fusion mapping function and the parameters of the tactile generation network; Train the tactile autoencoder, the audio feature extraction network, the image feature extraction network, the multi-modal fusion mapping function, and the tactile generation network, including: Step 1: Preset and and optimize : Step 11: Initialize network parameters ; Node class label ; Class center matrix ; Set the number of clustering clusters C, learning rate , number of iterations T; Step 12, Fix and , optimize based on the stochastic gradient descent method : ; ; ; ; Among them, is to take the partial derivative of each loss function; Step 13, fixation , and , optimization : ; ; ; Step 14, Fix , and , Optimize : ; Step 15. If t , then jump to Step 12 and continue the next iteration; otherwise, terminate the iteration. Step 16: After T rounds of iteration, obtain the parameters of the optimal audio feature extraction network , the parameters of the image feature extraction network , the parameters of the tactile autoencoder , the node class labels and the clustering center vector matrix of the tactile data ; Step 2: Based on the stochastic gradient descent method, measure and : Step 21, Initialization ; Learning rate , Number of iterations ; Step 22. According to Estimate using the stochastic gradient descent method : ; Step 23, according to Update using the stochastic gradient descent method : ; Step 24, if , then jump to Step 22 , continue the next iteration; otherwise terminate the iteration Step 25, after rounds of iteration, the optimal is obtained.
Citation Information
Patent Citations
Cross-modal image generation method and device based on audio-tactile signal fusion
CN113627482A
Image-tactile signal mutual reconstruction method and device
CN114595739A