Multi-channel touch-slip optical fiber sensing recognition method based on multi-modal fusion
Through multi-channel fiber grating sensors and multimodal fusion models, the problems of single-point single-channel measurement and electromagnetic interference are solved, high-precision recognition of complex surface objects is achieved, and the robot's perception ability is improved.
Patent Information
- Application Number
- CN202411120471.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-08-15
AI Technical Summary
Most existing touch-slip fiber optic sensors only perform single-point single-channel data measurements, and the multimodal fusion algorithm is insufficient, resulting in reduced recognition performance when identifying objects on complex surfaces. Traditional electrical signal sensors are also susceptible to electromagnetic interference.
By adopting multi-channel fiber Bragg grating sensors and combining them with a multi-modal mutual fusion model, high-precision measurement and full fusion of vibration and stress signals are achieved through data collection, demodulation and multi-channel multi-modal mutual fusion model of multi-channel touch and slip information, and information such as the hardness, texture and curvature of the object surface can be identified.
It achieves high-precision and high-accuracy recognition in complex surface object identification, avoids electromagnetic interference, and improves the robot's perception performance.
Smart Images

Figure CN118981747B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of flexible electronic technology, and in particular relates to an identification system that obtains tactile and slip information on an object surface through a multi-channel fiber grating sensor and fuses the tactile and slip multimodal information. Background Art
[0002] With the convergence of informatization and industrialization, the intelligent industry continues to rise, and robotic biomimetic perception technology is developing. Intelligent robots are widely used in industries and services. Touch and slip are the most important connection between humans and the world, and are also crucial means for robotic hands to perceive external information. Tactile sensation can detect stress signals when a human hand touches an object, while slip can detect sliding signals when a human hand touches an object. Combining stress and vibration information and applying them to intelligent biomimetic robots can improve the robot's perception performance, enabling it to accomplish more complex tasks such as accurately identifying and detecting objects.
[0003] A multi-channel fiber-optic tactile sensing and recognition system based on multimodal fusion fuses acquired multimodal information to form a holistic perception of the tactile sensation of an object's surface. Traditional electrical signal tactile sensors are susceptible to electromagnetic interference. While they can sense vibration, stress, and deformation, they can only accurately measure a single physical quantity. The fiber-optic sensor employed in the present invention effectively avoids electromagnetic interference and offers advantages such as small size, light weight, excellent stability, and good multiplexing. However, currently used fiber-optic Bragg grating (FBG) sensors for tactile sensation mostly measure data at a single point and in a single channel. The corresponding multimodal fusion algorithms only fuse single-point tactile sensation data. This presents limitations for existing sensing and recognition systems when dealing with objects with complex surfaces, such as those with different shapes or surfaces with varying curvatures. Furthermore, the network architecture used in multimodal fusion algorithms often only incorporates partial features from one modality into another, resulting in insufficient fusion of multimodal features. As the network depth increases, the recognition performance of the sensing system decreases. Summary of the Invention
[0004] This invention proposes a multi-channel fiber-optic touch-slip sensing system based on multimodal fusion. The data acquisition component includes the touch-slip sensor optical path and fiber Bragg grating (FBG) packaging. The dataset construction component demodulates the multi-channel touch-slip information collected by the fiber Bragg grating sensor to obtain vibration and stress signals, and uses the results to construct a multi-channel touch-slip dataset. The multi-channel stress signal modal and vibration signal modal fusion component uses training to generate a multi-channel touch-slip multimodal fusion model to identify object surfaces. The solution provided by this invention is as follows:
[0005] A multi-channel touch-slip optical fiber sensing recognition method based on multi-modal fusion includes the following steps:
[0006] The first step is to collect multimodal information of touch and slip
[0007] A multi-channel fiber optic sensing system for touch and slip information based on fiber Bragg gratings was built. The fiber Bragg grating sensors were encapsulated in different silicone sleeves, which were placed at the end of the robotic arm to collect multimodal information of touch and slip.
[0008] The second step is to build a multi-channel touch and slip dataset
[0009] A robotic arm was used to contact silicone blocks of varying hardness, texture, shape, and curvature, collecting multi-channel interference signals. The center wavelength offset of each interference signal was extracted as stress information, and the instantaneous frequency of the interference signal was demodulated to obtain vibration information. The multi-channel stress information, vibration information, and labels were combined to form a multi-channel tactile-slip dataset. The tactile-slip information of each channel exhibited different manifestations at corresponding locations, and the data from each channel was combined to obtain information on the hardness, texture, shape, and curvature of the object's surface.
[0010] The third step is to build and train a multi-channel touch-slip multimodal mutual fusion model. The network structure of the multi-channel touch-slip multimodal mutual fusion model consists of two parts. The first part is a single-channel multimodal mutual attention fusion module, and the second part is a multi-channel fusion module.
[0011] (1) The single-channel multimodal mutual attention fusion module performs a preliminary fusion of the single-channel touch and slip information to obtain fusion features. The module includes an N-layer multimodal mutual decoder module and N-1 iterative multimodal interaction modules. The multimodal mutual decoder module fuses the vibration information and stress information, and introduces an iterative multimodal interaction module to enhance the stress information. The structure of each layer of the multimodal mutual decoder module is the same, and the output of each layer is the stress information feature fused with the vibration information feature. The output result of each layer is input into the iterative multimodal interaction module to obtain the fusion feature of the vibration information and stress information of the layer. After iteration, the fusion feature of the vibration information and stress information is finally obtained.
[0012] (2) The multi-channel fusion module, based on the GRU attention mechanism, further fuses the fusion features of each channel to form a multi-channel touch-slip multimodal fusion feature, and normalizes it to realize the recognition of object surface information.
[0013] Furthermore, the single-channel multimodal mutual attention fusion module includes:
[0014] Input the stress information and vibration information into the multimodal mutual decoder module; use the Transformer encoder to capture the dependencies between different positions in the vibration information data, namely the deep vibration feature F DThe input of the first layer of the multimodal mutual decoder module is stress information and deep vibration features F D , stress information generates stress information features F through multi-head self-attention S , projecting stress information features onto the bond through a linear layer Sum Similarly, the deep vibration feature F is transformed into D Projection to key Sum Then the interactive attention matrix is generated and normalized to obtain the vibration information features of the fused strain mode. and fusion of the stress information characteristics of the vibration mode
[0015] Will Each element in is considered as a weight of the linear layer and applied to On the top, obtain the initial fusion features
[0016] Will Perform linear mapping and then project it to the key of multi-head cross attention and value F S Projected to the weight of multi-head cross attention d k for The dimension of the multimodal mutual decoder module is obtained by the first layer output result, that is, the stress information feature of the first layer fused with the vibration information feature
[0017] The stress information features of the first layer are fused with the vibration information features and stress information characteristics F S Input iterative multimodal interaction module, through the linear layer mapping F S , generate the stress characteristics of the current layer Then use the linear layer to project the touch-slip multimodal fusion features First, we map F with a linear layer 1 S Generate the output of the current layer of the iterative multimodal interaction module Next, we use the linear layer 3 to project the multimodal fusion features of touch and slip to obtain Calculate the attention matrix of the reorganized stress feature and normalize it to obtain
[0018]
[0019] Then pass and mapped with a linear layer 2 Generated Obtaining new stress characteristics
[0020]
[0021] Will Inject into middle:
[0022]
[0023] in, is a learnable weight. As each layer is iterated, BN stands for batch normalization and the output is is the fusion feature of the first layer of vibration information and stress information, which is combined with the deep vibration feature F D Together they serve as the input to the second layer of the multimodal mutual decoder module; Generate stress information feature F through multi-head self-attention dec1 , through the linear layer F dec1 Projection to key Sum Similarly, project vibration information characteristics to bonds Sum Generate an interactive attention matrix; normalize the interactive attention matrix to obtain the vibration information characteristics of the fused strain mode and multiple fusion features of fused vibration modes
[0024] Will Each element in is considered as a weight of the linear layer and applied to Above, that is, the initial fusion feature Will Perform linear mapping and then project it to the key of multi-head cross attention and value F S Projected to the weight of multi-head cross attention d k for The dimension of the multimodal mutual decoder module is obtained by the second layer output result, that is, the stress feature of the second layer fused with the vibration feature
[0025] The stress characteristics of the second layer are fused with the vibration characteristics and the first floor Input iterative multimodal interaction module; first use linear layer 1 to map Generate the output of the current layer of the iterative multimodal interaction module Next, we use the linear layer 3 to project the multimodal fusion features of touch and slip to obtain Calculate the attention matrix of the reorganized stress feature and normalize it to obtain Then pass and mapped with a linear layer 2 Generated Obtaining new stress characteristics
[0026] Will Inject into In the second layer, the fusion features of vibration information and stress information are obtained
[0027] Repeat the above steps iteratively to obtain the fusion features of vibration information and stress information of the N-1th layer.
[0028] Input it into the last layer of the single-channel multimodal mutual decoder module to obtain the final fusion features of vibration information and stress information
[0029] Furthermore, the multi-channel fusion module includes:
[0030] 1) The fusion features of the vibration information and stress information obtained from each channel are finally concatenated into a multi-channel fusion feature sequence F;
[0031] 2) Input the sequence F into the multi-channel fusion module to obtain the output and hidden state of each dimension through GRU,
[0032] 3) Extract vibration information and stress information fusion feature v through attention mechanism:
[0033] 4) Perform normalization calculation to obtain the attention weight matrix:
[0034] A=softmax(v)
[0035] The label corresponding to the eigenvector is determined according to the attention weight matrix to achieve overall recognition of the hardness, texture, and curvature information of the object surface.
[0036] Furthermore, assuming that the number of channels is 3, the multi-channel fusion module includes:
[0037] The fusion features of the three channels Splicing to generate sequence Input sequence F into the multi-channel fusion module to obtain outputs h1, h2 and h3: the first element of sequence F Obtain the output and hidden state h1 through GRU; the second element of the sequence F Obtain the output and hidden state h2 through GRU; the third element of the sequence F Obtain the output and hidden state h3 through GRU;
[0038] Extract vibration information and stress information fusion feature v through attention mechanism:
[0039] e i =w i tanh(W i h i +b i ),1≤i≤3
[0040]
[0041] Among them, e i Indicates the relevant information h corresponding to the i-th eigenvector i The determined attention probability distribution value, w i and W i represents the weight coefficient matrix of the i-th fusion feature vector, b i Indicates the offset corresponding to the i-th fusion feature.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] (1) Compared with traditional touch and slip sensors based on resistance, voltage, and piezoelectricity, the present invention adopts a fiber Bragg grating sensor with high anti-electromagnetic interference capability, which can achieve high-precision measurement of both vibration and stress modes, and can obtain spatial information such as hardness, texture, shape, and curvature through multi-channel acquisition points located at different positions.
[0044] (2) Compared with the traditional single-modal single-channel object surface recognition algorithm, the model adopted by the present invention integrates the vibration signals and stress signals obtained by the three channels, thereby capturing more comprehensive information about the object surface. At the same time, it has higher recognition accuracy and can achieve more accurate object surface recognition function. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flow chart of the specific operation of the present invention.
[0046] Figure 2 It is a schematic diagram of the optical path of the touch-slip sensor of the present invention.
[0047] Figure 3 It is a schematic diagram of optical fiber packaging of the present invention.
[0048] Figure 4 It is a schematic diagram of the process of constructing a multi-channel touch and slip dataset in the present invention.
[0049] Figure 5 This is an example of visualization of multi-channel tactile information obtained by the present invention.
[0050] Figure 6It is a structural diagram of the single-channel multimodal mutual attention fusion module adopted in the present invention.
[0051] Figure 7 It is a schematic diagram of the structure of the multi-channel fusion module adopted in the present invention.
[0052] Figure 8 This is the five-fold cross-validation result of the multi-channel touch-slip multimodal mutual fusion model adopted by the present invention. DETAILED DESCRIPTION
[0053] The following further illustrates how to use the multi-channel fiber optic sensing system proposed in the present invention to perform object surface recognition and perception in conjunction with the accompanying drawings and specific embodiments.
[0054] The present invention is achieved through the following technical solutions, as shown in the flow chart: Figure 1 As shown, the specific steps of the present invention are:
[0055] The first step is to collect multimodal information of touch and slip
[0056] The present invention uses a fiber Bragg grating sensor with strong anti-electromagnetic interference capability as a touch-slip sensor to obtain touch-slip information. The system optical path is as follows: Figure 2 As shown, it includes a swept light source module 1, a 2×2 coupler 2, a 2×2 coupler 6, a 2×2 coupler 7 and a 2×2 coupler 8, a fiber Bragg grating 3, a fiber Bragg grating 4 and a fiber Bragg grating 5 with a reflection bandwidth of 1 nm, a photodetector 9, a photodetector 10 and a photodetector 11, a Faraday rotator 12, a Faraday rotator 13 and a Faraday rotator 14, a data acquisition card 15, and a host computer data processing 16.
[0057] A fiber Bragg grating (FBG), a Faraday rotator, a sweeping light module, and a photodetector are connected to the four ports of a 2×2 coupler to form a Michelson interferometer structure. The fiber Bragg grating serves as the detection arm, while the Faraday rotator serves as the reference arm. During a scanning cycle, when the wavelength of the sweeping light source matches the reflected wavelength of the fiber Bragg grating (FBG), interference occurs in the 2×2 coupler, generating the desired interference signal. Three Michelson interferometer structures are connected in parallel to form a multichannel optical path. The sweeping light module serves as the input and is connected to the three 2×2 couplers via a 2×2 coupler. The three photodetectors serve as the multichannel outputs and are connected to a data acquisition card via SMA cables. This connection is then connected to a host computer via the PCIE interface on the data acquisition card.
[0058] The packaging of the fiber grating used in the present invention is as follows Figure 3As shown. Three fiber Bragg gratings are encapsulated in three silicone sleeves, which are placed on three resin rods at the end of the robotic arm. Stress and vibration information is obtained by controlling the robotic arm to contact the surface of the object to be sampled. The robotic arm's downward pressure distance is 10mm, and the sliding speed is 160mm / s. The object to be sampled is a silicone block with different hardness, texture, shape, and curvature. The hardness is Shore A30 and Shore A70, the texture includes cement texture and silicone texture, and there are two curvatures: Curvature 1 is 0.1599 and Curvature 2 is 0.4777.
[0059] The parameters used for each module in the system's optical path are as follows: the sweeping light source has a scanning rate of 50 Hz, the data acquisition card has an acquisition rate of 100 kSa / s, and the optical path difference between the two arms is set to 5 cm. Each corresponding scanning frame has a length of 20 ms, each scanning frame contains approximately 80 interference signal wavenumbers, and the interference signal length of each scanning frame is approximately 2000 scanning points. According to the Nyquist theorem, the maximum frequency for system stress information detection is 25 Hz, and the maximum frequency for vibration information detection is 1 kHz.
[0060] The second step is to build a multi-channel touch-slip dataset.
[0061] The multi-channel interference signal is collected by the data acquisition card, and the interference signal of each channel is expressed as:
[0062]
[0063] The envelope of the interference signal is represented by the reflection spectrum R(λ) of the fiber Bragg grating, L is the optical path difference of the fiber Bragg grating, and n is the refractive index of the fiber Bragg grating. is the phase change of the interference signal caused by vibration. Stress information can be extracted from R(λ). The vibration information is extracted from the sensor. The interference signal with multi-channel touch and slip data is transmitted to the host computer through the data acquisition card for demodulation.
[0064] The interference signal of the current frame is subjected to wavelet denoising, the interference signal envelope is extracted by spline difference, and then Gaussian smoothing is performed on the envelope. Next, the vibration information and stress information are obtained by performing different processing on the envelope: the offset of the envelope peak in the scanning frame is calculated, the strain waveform is fitted according to the offset, and the waveform is used as stress information; the 3dB area in the envelope is extracted, and the envelope in the area is subjected to SWT time-frequency analysis, and then the wavelet ridge is extracted, and the instantaneous phase is restored by integration to obtain the vibration waveform, and the vibration waveform is downsampled to obtain the vibration information. Figure 5As shown in the figure, the touch and slip information is visualized. The left side is the visualization of three-channel vibration information, and the right side is the visualization of three-channel stress information. It can be seen that the touch and slip information of the three channels has different performances at different positions. By integrating these differences, the overall hardness, texture, shape, curvature and other information of the object surface can be obtained.
[0065] Multi-channel stress information, vibration information and labels form a data group Data i =(S 1i , S 2i , S 3i , D 1i , D 2i , D 3i , L i 0, where Data i represents the i-th group of data, S 1i , S 2i , S 3i Represents the stress information data corresponding to each channel, D 1i , D 2i , D 3i Indicates the vibration information data corresponding to each channel, L i The label representing the surface condition of the object is obtained through the characteristics of the object to be sampled.
[0066] The third step is to build and train a multi-channel touch-slip multimodal fusion model.
[0067] The multimodal fusion algorithm currently used can only fuse single-point touch and slip information data, and cannot directly process multi-channel data. At the same time, the network structure used for the multimodal fusion algorithm can often only integrate partial features of a certain modality into another feature, which leads to insufficient feature fusion. As the network depth increases, the recognition performance of the sensor system decreases.
[0068] The multi-channel touch-slip multi-modal mutual fusion model adopted by the present invention can better improve the above problems, can process multi-channel touch-slip information data, and fully integrate the two modal data during the fusion process. The network structure of the model consists of two parts. The first part is as follows: Figure 6 The single-channel multimodal mutual attention fusion module shown in the figure is as follows. Figure 7 The multi-channel fusion module shown.
[0069] The single-channel multimodal mutual attention fusion module performs a preliminary fusion of single-channel tactile information to obtain fusion features. The module consists of an N-layer multimodal mutual decoder module and N-1 iterative multimodal interaction modules. The multimodal mutual decoder module fuses vibration and stress information, but stress information is easily lost during the iteration process. Introducing the iterative multimodal interaction module can enhance stress information. Each layer of the multimodal mutual decoder module has an identical structure, and each layer outputs stress information features fused with vibration information features. The output of each layer is input into the iterative multimodal interaction module to obtain the fusion features of the vibration and stress information of that layer. After iteration, the fusion features of vibration and stress information are finally obtained.
[0070] In the first step, the stress information and vibration information are input into the multimodal mutual decoder module. The Transformer encoder is used to capture the dependencies between different positions in the vibration information data, namely the deep vibration feature F D The input of the first layer of the multimodal mutual decoder module is stress information and deep vibration features F D , stress information generates stress information features F through multi-head self-attention S , projecting stress information features onto the bond through a linear layer Sum Similarly, project vibration information characteristics to bonds Sum Next, generate the interaction attention matrix:
[0071]
[0072] Normalizing the interaction attention matrix, we get:
[0073]
[0074] in, In order to integrate the vibration information characteristics of the strain mode, To fuse the stress information features of the vibration mode. Each element in is considered as a weight of the linear layer and applied to Above is the initial fusion feature:
[0075]
[0076] Will Perform linear mapping and then project it to the key of multi-head cross attention and value F S Projected to the weight of multi-head cross attention d k for The dimension of is obtained, and the output result of the first layer of the multimodal mutual decoder module is obtained, that is, the stress information feature of the first layer fused with the vibration information feature:
[0077]
[0078] The second step is to integrate the stress information features of the first layer into the fusion vibration information features and stress information characteristics F S Input iterative multimodal interaction module. First, use linear layer to map F S , generate the stress characteristics of the current layer Next, we use the linear layer to project the multimodal fusion features of touch and slip First, we map F with a linear layer 1 S Generate the output of the current layer of the iterative multimodal interaction module Next, we use the linear layer 3 to project the multimodal fusion features of touch and slip to obtain Calculate the attention matrix of the reorganized stress features and normalize it:
[0079]
[0080] Then pass and mapped with a linear layer 2 Generated Get the new stress feature:
[0081]
[0082] Next, Inject into middle:
[0083]
[0084] in, is a learnable weight. As each layer is iterated, BN stands for batch normalization and the output is is the fusion feature of the first layer of vibration information and stress information, which is combined with the deep vibration feature F D Together they serve as the input to the second layer of the multimodal mutual decoder module. Generate stress information feature F through multi-head self-attention dec1 , through the linear layer F dec1 Projection to key Sum Similarly, project vibration information characteristics to bonds Sum Next, generate the interaction attention matrix:
[0085]
[0086] Normalizing the interaction attention matrix, we get:
[0087]
[0088] In order to integrate the vibration information characteristics of the strain mode, is the multiple fusion features of the fused vibration modes. Each element in is considered as a weight of the linear layer and applied to Above is the initial fusion feature:
[0089]
[0090] Will Perform linear mapping and then project it to the key of multi-head cross attention and value F S Projected to the weight of multi-head cross attention d k for The dimension of is obtained, and the output result of the second layer of the multimodal mutual decoder module is obtained, that is, the stress feature of the second layer fused with the vibration feature:
[0091]
[0092] Then, the stress features of the second layer are fused with the vibration features and the first floor Input iterative multimodal interaction module. First, use linear layer to map Generate stress characteristics for the current layer First, we use a linear layer 1 to map Generate the current layer output of the iterative multimodal interaction module Next, we use the linear layer 3 to project the multimodal fusion features of touch and slip to obtain Calculate the attention matrix of the reorganized stress features and normalize it:
[0093]
[0094] Then pass and mapped with a linear layer 2 Generated Get the new stress feature:
[0095]
[0096] Next, Inject into In the example, the fusion features of the vibration information and stress information of the second layer are obtained:
[0097]
[0098] Next, Inject into In the example, the fusion features of the vibration information and stress information of the second layer are obtained:
[0099]
[0100] Repeat the above steps iteratively to obtain the fusion features of vibration information and stress information of the N-1th layer.
[0101]
[0102] Input it into the last layer of the multimodal mutual decoder module to obtain the final fusion features of vibration information and stress information
[0103]
[0104] The multi-channel fusion module combines the fusion features of the three channels The multi-channel touch-slip multimodal fusion features are further fused and normalized to achieve the recognition of object surface information. The key to the multi-channel fusion module is the GRU-attention mechanism model. Among them, GRU is a recurrent neural network structure similar to LSTM, which can further extract the fusion features of the three channels. The information relationship between the two fusion features before and after the sequence, such as the fusion feature The information given to the feature vector Enhance the relationship between the fusion features of different channels input into the attention mechanism. The attention mechanism further extracts the fusion features of vibration information and stress information, highlighting the key part of the fusion features, that is, the feature vector containing the fusion of multi-channel vibration signal and stress signal information.
[0105] The first step is to combine the fusion features of the three channels Splicing to generate sequence
[0106] In the second step, the sequence F is input into the multi-channel fusion module to obtain the outputs h1, h2 and h3.
[0107] The first element of sequence F Get the output and hidden state h1 through GRU. First, get two gate states:
[0108]
[0109] Among them, r is the gate for controlling reset, and z is the gate for controlling update. Then use the reset gate to get the reset data O*r1, and then add O*r1 to the current fusion feature vector Connect and scale the data to the range of (-1, 1) through a tanh activation function:
[0110]
[0111] Here h1 mainly contains the current fusion feature vector We also need to remember the current state, which is achieved by adding h1 to the current hidden state in a targeted manner. Next, we use update gating to forget and remember:
[0112] h1=z1*h1
[0113] The update gate z is in the range of (0,1). The closer z is to 1, the more data is remembered, and the closer z is to 0, the more data is forgotten.
[0114] The second element of sequence F The output and hidden state h2 are obtained through GRU. First, the previous hidden state and Get two gate states:
[0115]
[0116] Then use the reset gate to get the reset data h1*r2, and then combine h1*r2 with the current fusion feature vector Connect and scale the data to the range of (-1, 1) through a tanh activation function:
[0117]
[0118] Here h2 mainly contains the current fusion feature vector We also need to remember the current state, which is achieved by adding h2 to the current hidden state in a targeted manner. Next, we use update gating to forget and remember to obtain the output:
[0119] h2=(1-z2)*h1+z2*h2
[0120] The third element of sequence F Similarly, we get the output h3
[0121] h3=(1-z3)*h2+z3*h3
[0122] The third step is to extract the fusion features of vibration information and stress information through the attention mechanism:
[0123] e i =wi tanh(W i h i +b i ),1≤i≤3
[0124]
[0125] Among them, e i Is a verification model, e i Indicates the relevant information h corresponding to the i-th eigenvector i The determined attention probability distribution value, w i and W i represents the weight coefficient matrix of the i-th fusion feature vector, b i Indicates the offset corresponding to the i-th fusion feature.
[0126] The third step is to perform normalization calculation to obtain the attention weight matrix:
[0127] A=softmax(v)
[0128] The label corresponding to the feature vector is determined according to the attention weight matrix, thereby achieving overall recognition of information such as the hardness, texture, and curvature of the object surface.
[0129] The multi-channel touch-slip multimodal fusion model was trained and validated using a five-fold crossover method. In the first step, the dataset was evenly divided into five subsets. In the second step, four of these subsets were used as training sets, and the remaining subset was used as the test set. These two steps were repeated five times, each time using a different subset as the test set. The recognition accuracy results from each training and validation were averaged to obtain the recognition accuracy of the multi-channel touch-slip multimodal fusion model.
[0130] In the process of single-channel touch-slip multimodal fusion, the present invention adopts Figure 6 The multimodal mutual attention fusion module shown in FIG is composed of a multimodal mutual decoder module and an iterative multimodal interaction module. This module fuses the stress and vibration modes to obtain a single-channel fusion feature. Figure 7 The multi-channel touch-slip multimodal fusion model structure shown here uses multi-channel touch-slip multimodal fusion of three single-channel fusion features to achieve surface recognition. Using a dataset containing 800 data sets obtained from actual touch-slip multimodal information collection as an example, the multi-modal mutual decoder module was set to 8 layers. The multi-channel touch-slip multimodal mutual fusion model was trained and validated using a five-fold crossover method, achieving a recognition accuracy of 87.19%. These results demonstrate that the multi-channel touch-slip multimodal mutual fusion model employed in the present invention achieves good recognition accuracy.
Claims
1. A multi-channel touch-slip optical fiber sensing recognition method based on multi-modal fusion, comprising the following steps: The first step is to collect multimodal information of touch and slip A multi-channel fiber optic sensing system for touch and slip information based on fiber Bragg gratings was built. The fiber Bragg grating sensors were encapsulated in different silicone sleeves, which were placed at the end of the robotic arm to collect multimodal information of touch and slip. The second step is to build a multi-channel touch and slip dataset A robotic arm was used to contact silicone blocks of varying hardness, texture, shape, and curvature, collecting multi-channel interference signals. The center wavelength offset of each interference signal was extracted as stress information, and the instantaneous frequency of the interference signal was demodulated to obtain vibration information. The multi-channel stress information, vibration information, and labels were combined to form a multi-channel tactile-slip dataset. The tactile-slip information of each channel exhibited different manifestations at corresponding locations, and the data from each channel was combined to obtain information on the hardness, texture, shape, and curvature of the object's surface. The third step is to build and train a multi-channel touch-slip multimodal mutual fusion model. The network structure of the multi-channel touch-slip multimodal mutual fusion model consists of two parts. The first part is a single-channel multimodal mutual attention fusion module, and the second part is a multi-channel fusion module. (1) Single-channel multimodal mutual attention fusion module, which performs preliminary fusion of single-channel touch and slip information to obtain fusion features. The module includes an N-layer multimodal mutual decoder module and N-1 iterative multimodal interaction modules. The multimodal mutual decoder module fuses vibration information and stress information, and introduces an iterative multimodal interaction module to enhance stress information. The structure of each layer of the multimodal mutual decoder module is the same, and the output of each layer is the stress information feature fused with the vibration information feature. The output result of each layer is input into the iterative multimodal interaction module to obtain the fusion feature of the vibration information and stress information of the layer. After iteration, the fusion feature of the vibration information and stress information is finally obtained. (2) The multi-channel fusion module, based on the GRU attention mechanism, further fuses the fusion features of each channel to form a multi-channel touch-slip multimodal fusion feature, and normalizes it to achieve the recognition of object surface information; The single-channel multimodal mutual attention fusion module includes: Input the stress information and vibration information into the multimodal mutual decoder module; use the Transformer encoder to capture the dependencies between different positions in the vibration information data, namely the deep vibration feature F D The input of the first layer of the multimodal mutual decoder module is stress information and deep vibration features F D , stress information generates stress information features F through multi-head self-attention S , projecting stress information features onto the bond through a linear layer Sum Similarly, the deep vibration feature F is transformed into D Projection to key Sum Then the interactive attention matrix is generated and normalized to obtain the vibration information features of the fused strain mode. and fusion of the stress information characteristics of the vibration mode Will Each element in is considered as a weight of the linear layer and applied to On the top, obtain the initial fusion features Will Perform linear mapping and then project it to the key of multi-head cross attention and value F S Projected to the weight of multi-head cross attention d k for The dimension of the multimodal mutual decoder module is obtained by the first layer output result, that is, the stress information feature of the first layer fused with the vibration information feature The stress information features of the first layer are fused with the vibration information features and stress information characteristics F S Input iterative multimodal interaction module, through the linear layer mapping F S , generate the stress characteristics of the current layer Then use the linear layer to project the touch-slip multimodal fusion features First, we map F with a linear layer 1 S Generate the current layer output of the iterative multimodal interaction module Next, we use the linear layer 3 to project the multimodal fusion features of touch and slip to obtain Calculate the attention matrix of the reorganized stress feature and normalize it to obtain Then pass and mapped with a linear layer 2 Generated Obtaining new stress characteristics Will Inject into middle: in, is a learnable weight. As each layer is iterated, BN stands for batch normalization and the output is is the fusion feature of the first layer of vibration information and stress information, which is combined with the deep vibration feature F D Together they serve as the input to the second layer of the multimodal mutual decoder module; Generate stress information feature F through multi-head self-attention dec1 , through the linear layer F dec1 Projection to key Sum Similarly, project vibration information characteristics to bonds Sum Generate an interactive attention matrix; normalize the interactive attention matrix to obtain the vibration information characteristics of the fused strain mode and multiple fusion features of fused vibration modes Will Each element in is considered as a weight of the linear layer and applied to Above, that is, the initial fusion feature Will Perform linear mapping and then project it to the key of multi-head cross attention and value F S Projected to the weight of multi-head cross attention d k for The dimension of the multimodal mutual decoder module is obtained by the second layer output result, that is, the stress feature of the second layer fused with the vibration feature The stress characteristics of the second layer are fused with the vibration characteristics and the first floor Input iterative multimodal interaction module; first use linear layer 1 to map Generate the current layer output of the iterative multimodal interaction module Next, we use the linear layer 3 to project the multimodal fusion features of touch and slip to obtain Calculate the attention matrix of the reorganized stress feature and normalize it to obtain Then pass and mapped with a linear layer 2 Generated Obtaining new stress characteristics Will Inject into In the second layer, the fusion features of vibration information and stress information are obtained Repeat the above steps iteratively to obtain the fusion features of vibration information and stress information of the N-1th layer. Input it into the last layer of the single-channel multimodal mutual decoder module to obtain the final fusion features of vibration information and stress information 2. The multi-channel touch-slip optical fiber sensing identification method according to claim 1, characterized in that: The multi-channel fusion module includes: 1) The fusion features of the vibration information and stress information obtained from each channel are finally concatenated into a multi-channel fusion feature sequence F; 2) Input the sequence F into the multi-channel fusion module to obtain the output and hidden state of each dimension through GRU, 3) Extract vibration information and stress information fusion feature v through attention mechanism: 4) Perform normalization calculation to obtain the attention weight matrix: A=softmax(v) The label corresponding to the eigenvector is determined according to the attention weight matrix to achieve overall recognition of the hardness, texture, and curvature information of the object surface.
3. The multi-channel touch-slip optical fiber sensing identification method according to claim 2, characterized in that: Assuming the number of channels is 3, the multi-channel fusion module includes: The fusion features of the three channels Splicing to generate sequence Input sequence F into the multi-channel fusion module to obtain outputs h1, h2 and h3: the first element of sequence F Obtain the output and hidden state h1 through GRU; the second element of the sequence F Obtain the output and hidden state h2 through GRU; the third element of the sequence F Obtain the output and hidden state h3 through GRU; Extract vibration information and stress information fusion feature v through attention mechanism: e i =w i fishy(W i h i +b i ),1≤i≤3 Among them, e i Indicates the relevant information h corresponding to the i-th eigenvector i The determined attention probability distribution value, w i and W i represents the weight coefficient matrix of the i-th fusion feature vector, b i Indicates the offset corresponding to the i-th fusion feature.
Citation Information
Patent Citations
Stress, temperature and vibration composite detection optical fiber sensor and signal processing method
CN110132329A
Flexible fingertip tactile-slip sensor based on fiber bragg grating and detection method of flexible fingertip tactile-slip sensor
CN116046031A