Method for classifying material surface texture based on fusion of planar texture force touch and sound signals
By combining the plane texture force tactile and sound signals, MFCC feature vectors are extracted and the texture recognition model is established using the SVM algorithm, the problem of difficulty in comprehensive classification of single-modal tactile information in the existing technology is solved, and higher material surface texture classification accuracy and model stability are achieved.
Patent Information
- Application Number
- CN202310645224.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-06-02
AI Technical Summary
The existing material surface texture classification methods mainly rely on single-modal tactile information, making it difficult to fully grasp the surface characteristics of the material, and lack of combination of visual or auditory information, resulting in low classification accuracy.
By combining plane texture force tactile and sound signals, three-axis accelerometer, three-axis force sensor and microphone data are collected, Mel frequency cepspectral coefficient (MFCC) is extracted as feature vectors, and a texture recognition model is established using the support vector machine (SVM) algorithm to realize the classification of material surface textures.
It improves the accuracy and reliability of material surface texture classification, enhances the generalization ability and stability of the model, can adjust the weight of the feature vector according to the situation, and improves the robustness of the classification model.
Smart Images

Figure CN116738281B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of virtual reality and human-computer interaction, and particularly relates to a planar texture force tactile and sound signal device for fast and accurate measurement and a material surface texture classification method based on the fusion of planar texture force tactile and sound signals. Background Art
[0002] Among various tactile information, the delicate and complex material surface texture information is one of the key bases for humans to recognize and identify objects and master material properties. In the process of interacting with objects, touch is one of the main means for humans to understand the surface properties of items in daily life, and the contact between the skin and the object provides important information for object recognition. In the field of industrial manufacturing and production, the material surface texture is a very important factor. The surface texture and texture will affect the appearance and performance of the product, so it is necessary to inspect and classify the surface.
[0003] Currently, there are already some tactile classification methods for material surface textures.
[0004] The 2011 paper "Vibrotactile recognition and categorization of surfaces by a humanoid robot" used a built-in triaxial accelerometer on an artificial nail to collect vibration acceleration data of the material surface texture, extracted spectral features, and performed surface recognition and classification.
[0005] The 2019 paper "DCNN for tactile sensory data classification based on transfer learning" used a transfer learning method formalized by a convolutional neural network (CNN) to convert three-dimensional tensile tactile data generated during interaction on an e-skin into a two-dimensional image, and used a pre-trained CNN to achieve the classification of tactile data.
[0006] In the 2020 Chinese patent CN112146701A, by controlling the movement of a two-axis motion platform and a pressing module, the electrical signal of a tactile sensor is obtained, the topography data and deformation distance of the surface of the object to be measured are obtained, and further the topography and soft hardness of the object surface are obtained.
[0007] In the 2022 Chinese patent CN114898219A, pressure data (tactile data) is used to perform object recognition by using an SVM machine learning method including a multi-head encoded self-attention layer and a feed-forward neural network layer.
[0008] Through the analysis of the research status of the tactile classification method for the texture of the material surface, most of the existing data applied to texture classification is single-modal tactile information, which can only provide information in one aspect. It is very difficult to comprehensively grasp the characteristics of the material surface. Due to the lack of combination with visual or auditory information, it is impossible to provide richer and more accurate information, nor can it further improve the accuracy of texture classification. Summary of the Invention
[0009] The present invention provides a method for classifying the texture of a material surface based on the fusion of planar texture force touch and sound signals, which improves the accuracy and reliability of the texture classification of the material surface by combining two types of information, namely texture force touch and sound signals. By collecting the data of a triaxial accelerometer, a triaxial force sensor, and a microphone during the scanning of the object to be measured, extracting the features of these data, and combining them with the SVM machine learning method to establish a texture recognition model, texture classification is achieved.
[0010] The technical solution adopted by the present invention includes the following steps:
[0011] (1) The robotic arm control module uses the robotic arm to drive and control the probe to apply a pressing force to the object to be measured, and can achieve scanning at different moving speeds, scan the length of the object to be measured, and can perform multiple scans;
[0012] (2) The signal acquisition module includes a triaxial accelerometer embedded inside the probe, which records the acceleration information of the three axes when the probe scans the object to be measured, a triaxial force sensor located below the object to be measured, which records the friction information of the three axes during the scanning process, and a microphone placed in front of the object to be measured, which records the sound signal generated during the scanning process;
[0013] (3) The signal processing module receives the signals from the signal acquisition module and displays the time-domain diagrams of the signals of the triaxial accelerometer, the triaxial force sensor, and the microphone on the display screen in real time;
[0014] (4) The signal processing module receives the signals from the signal acquisition module, performs data pre-emphasis using first-order high-pass filtering, frames, windows, and extracts the Mel frequency cepstral coefficients (MFCCs) as feature vectors;
[0015] (5) Fuse the feature vectors extracted from the triaxial accelerometer data, the triaxial force sensor data, and the microphone data. Use linear weighted fusion to assign a weight to the feature vectors of the triaxial accelerometer, the triaxial force sensor, and the microphone, and then add them weighted to obtain the total feature vector. Store the total feature vector together with the corresponding object to be measured in the sample library;
[0016] (6) The training of the classification model adopts the support vector machine (SVM) algorithm, and the newly collected samples are classified and recognized through the trained texture surface recognition model in the sample library.
[0017] In step (1) of the present invention, the probe of the robotic arm control module scans the surface of the object to be measured, specifically including:
[0018] The robotic arm drive control module uses Visual Studio 2019 to control the robotic arm to apply a pressing force F 1 of 1N to 2N on the object to be measured and scan it at different moving speeds V 1 of 60mm / s to 140mm / s, set the length L 1 of 200mm to 300mm of the object to be measured, and set the number of scans N 1 .
[0019] In step (2) of the present invention, the signal acquisition module acquires signals, specifically including:
[0020] Three-axis acceleration data a is obtained in real time through a three-axis accelerometer embedded in the probe;
[0021] Three-axis friction data f is obtained in real time through a three-axis force sensor located below the object to be measured;
[0022] Audio signal s is obtained in real time through a microphone placed in front of the object to be measured;
[0023] Then the acquired data is transmitted to the signal processing unit.
[0024] In step (3) of the present invention, the acquired signals are received and displayed in real time, specifically including:
[0025] The signal processing unit receives the data (a, f, s) from the three-axis accelerometer, the three-axis force sensor, and the microphone. Subsequently, the time-domain diagrams of the three-axis accelerometer data, the time-domain diagrams of the three-axis force sensor data, and the time-domain diagrams of the microphone data are displayed on the display screen in real time to help the operator analyze and judge the signals. By detecting the change of the signal data in real time, the signal processing unit discovers abnormal situations and responds in a timely manner, so as to realize the full-range monitoring and diagnosis of the object to be measured.
[0026] In step (4) of the present invention, signal pre-emphasis, framing, windowing and feature vector extraction are performed, specifically including:
[0027] Step 401, combine the components a x , a y , a z of the x, y, and z axes of the three-axis acceleration data a into a total a', and the components f of the x, y, and z axes of the three-axis force sensor data fx , f y , f z Combine into a total f', and complete it according to the following formula of L2 norm:
[0028]
[0029]
[0030] where represents the magnitude at the j-th sampling point of the x-axis of the triaxial accelerometer data a, represents the magnitude at the j-th sampling point of the y-axis of the triaxial accelerometer data a, represents the magnitude at the j-th sampling point of the z-axis of the triaxial accelerometer data a; represents the magnitude at the j-th sampling point of the x-axis of the triaxial force sensor data f, represents the magnitude at the j-th sampling point of the y-axis of the triaxial force sensor data f, represents the magnitude at the j-th sampling point of the z-axis of the triaxial force sensor data f;
[0031] Step 402, perform pre-emphasis on the triaxial acceleration data a', triaxial force sensor data f', and microphone data s according to the following formula of the first-order high-pass filter to obtain a 1 , f 1 and s 1 :
[0032] a 1 (j) = b × a'(j) + (1 - b) × a 1 (j - 1)
[0033] f 1 (j) = b × f'(j) + (1 - b) × f 1 (j - 1)
[0034] s 1 (j) = b × s(j) + (1 - b) × s 1 (j - 1)
[0035] where j represents the value at the j-th sampling point, b is the parameter of the filter, indicating that the output value of the current sampling point is affected by the output value of the previous sampling point to a certain extent, and the value range of b should be between 0 and 1;
[0036] Step 403, frame a 1 , f 1 and s 1 according to a fixed time window, use a window with a length of 20 ms, and connect the frames in a 50% overlapping manner, and apply a Hamming window to each frame of the signal;
[0037]
[0038]
[0039] a 1w (n) = a 1k (n)w(n)
[0040]
[0041]
[0042] f 1w (n) = f 1k (n)w(n)
[0043]
[0044]
[0045] s 1w (n) = s 1k (n)w(n)
[0046] Among them, the window length is N, a 1k , f 1k , s 1k represent obtaining the k-th frame signal, w(n) represents the Hamming window function, a 1w , f 1w , s 1w represent the k-th frame obtained after windowing;
[0047] Step 404, for each frame signal a 1w , f 1w , s 1w after windowing, perform short-time Fourier transform to convert the time-domain signals a 1w , f 1w , s 1w to the frequency domain A 1w , F 1w , S 1w , and dot-multiply the Mel filter bank with A 1w , F 1w , S 1w to obtain the Mel spectrum energy coefficients AE 1w (i), FE 1w (i), SE 1w (i) of the k-th frame. The Mel filter bank is as follows;
[0048]
[0049]
[0050]
[0051]
[0052] Among them, the center frequency of the i-th Mel filter is f(i), and H 0 (m) represents the response value of the i-th Mel filter in terms of frequency, and m is the m-th triangular filter in the filter bank;
[0053] Step 405: Take the logarithm and perform discrete cosine transform DFT to convert the 13 Mel frequency coefficients of each frame into 13 MFCC coefficients;
[0054] Step 406: Extract the MFCC feature matrix and retain the first 13 coefficients of the discrete cosine transform:
[0055]
[0056]
[0057]
[0058] Among them, M is the number of Mel filter banks, l is the MFCC coefficient serial number, and AC k (l) is the MFCC coefficient of the triaxial accelerometer data, and PC k (l) is the MFCC coefficient of the triaxial force sensor data, and SC k (l) is the MFCC coefficient of the microphone data. They are arranged in the order of each frame, and the first 13 MFCC coefficients of each frame are arranged in order to form a feature vector, obtaining the triaxial accelerometer feature vector f 1 、the triaxial force sensor feature vector f 2 、the microphone data feature vector f 3 。
[0059] In step (5) of the present invention, linear weighted fusion is performed according to the feature vectors extracted in step (4), specifically including:
[0060] Calculate the signal-to-noise ratios of the triaxial acceleration data a, the triaxial force sensor data f, and the audio data s according to the following formula:
[0061]
[0062]
[0063]
[0064] Among them, SNR a represents the signal-to-noise ratio of the triaxial acceleration data a, and SNR fIndicates the signal-to-noise ratio of the three-axis force sensor data f, SNR s Indicates the signal-to-noise ratio of the microphone data s, Ps a Indicates the signal power of the three-axis acceleration data a, Ps f Indicates the signal power of the three-axis force sensor data f, Ps s The signal power of the microphone data s, Pn a Indicates the noise power of the three-axis acceleration data a, Pn f Indicates the noise power of the three-axis force sensor data f, Pn s The noise power of the microphone data s, calculate the mean mu and standard deviation sigma of the three types of data (a, f, s); calculate the power of the signal and noise, that is, Ps = mu 2 and Pn = sigma 2 ;
[0065] Calculate the feature vector f of the three-axis acceleration data through the following formula 1 、the feature vector f of the three-axis force sensor data 2 and the feature vector f of the microphone data 3 of the weighted sum:
[0066] F = w 1 ×f 1 +w 2 ×f 2 +w 3 ×f 3
[0067]
[0068]
[0069]
[0070] Among them, w 1 Indicates the weight value of the feature vector f 1 , w 2 Indicates the weight value of the feature vector f 2 , w 3 Indicates the weight value of the feature vector f 3 , SNR a Indicates the signal-to-noise ratio of the three-axis acceleration data a, SNR f Indicates the signal-to-noise ratio of the three-axis force sensor data f, SNR s Indicates the signal-to-noise ratio of the microphone data s, according to the determined weight value, for f 1 、f 2 、f 3 Perform weighted calculation to obtain the final fusion feature vector F.
[0071] Step (6) of the present invention specifically includes:
[0072] Divide the data in the sample library into a training set and a test set, and the ratio of the training set to the test set is 8:2;
[0073] Select the radial basis function RBF kernel for the SVM algorithm kernel function, and use cross-validation to adjust the regularization penalty parameter C and the parameter gamma that controls the width of the radial basis function kernel function;
[0074] Use the training set in the sample library to train the SVM algorithm to obtain a classifier model;
[0075] Use the test set in the sample library to evaluate the performance of the classifier model and calculate the accuracy of the classifier model;
[0076] Use the trained SVM classifier to classify newly collected sample data.
[0077] The present invention has the following beneficial effects:
[0078] 1. Use planar texture force tactile data and sound data for object texture recognition. The fusion of planar texture force tactile and sound information can provide more information for material surface texture classification, thereby improving the classification accuracy;
[0079] 2. Collect texture force tactile and sound data with different pressing forces and different scanning speeds, which increases the sample diversity and effectively improves the generalization ability and stability of the model;
[0080] 3. The weights of the texture force tactile data and sound data feature vectors can be adjusted according to the situation to improve the robustness of the texture force tactile and sound signal classification model. Description of the Drawings
[0081] Figure 1 is a schematic diagram of the planar texture force tactile and sound signal acquisition device of the present invention;
[0082] Figure 2 is the working flow chart of the planar texture force tactile and sound signal acquisition device of the present invention;
[0083] Figure 3 is the flow chart of the texture force tactile and sound data recognition method extracted by the present invention;
[0084] Figure 4 is the three-axis acceleration frequency domain diagram when collecting the interaction with the real texture;
[0085] Figure 5 is the schematic diagram of the texture force tactile and sound signals when collecting the interaction with the real texture;
[0086] Figure 6It is a flowchart for calculating the feature vectors of force-tactile and sound data;
[0087] Figure 7 It is a flowchart for establishing a recognition model of planar texture force-tactile and sound signals;
[0088] Figure 8 It is a normalized confusion matrix diagram of the classification results of the sample library data. Specific implementation manners
[0089] It includes the following steps:
[0090] (1). The robotic arm control module uses the robotic arm to drive and control the probe to apply a pressing force F to the object to be measured 1 , and can achieve scanning at different moving speeds V 1 , and scan the length L of the object to be measured 1 and perform N 1 scans;
[0091] (2). The signal acquisition module includes a triaxial accelerometer embedded inside the probe, which records the acceleration information of the three axes when the probe scans the object to be measured, a triaxial force sensor located below the object to be measured, which records the friction information of the three axes during the scanning process, and a microphone placed in front of the object to be measured, which records the sound signal generated during the scanning process;
[0092] (3). The signal processing module receives the signals from the signal acquisition module, and displays the time-domain diagrams of the signals of the triaxial accelerometer, the triaxial force sensor and the microphone on the display screen in real time;
[0093] (4). The signal processing module receives the signals from the signal acquisition module, performs data pre-emphasis using first-order high-pass filtering, frames, windows, and extracts the Mel-frequency cepstral coefficients (MFCCs) as feature vectors;
[0094] (5). The feature vectors extracted from the triaxial accelerometer data, the triaxial force sensor data and the microphone data are fused. Linear weighted fusion is used to assign a weight to the feature vectors of the triaxial accelerometer, the triaxial force sensor and the microphone, and then they are weighted and added together to obtain the total feature vector, and the total feature vector is stored in the sample library together with the corresponding object to be measured;
[0095] (6). The training of the classification model adopts the support vector machine (SVM) algorithm, and the newly acquired samples are classified and recognized by the texture surface recognition model trained in the sample library.
[0096] In step (1) of the present invention, the probe of the robotic arm control module scans the surface of the object to be measured, which specifically includes:
[0097] The robotic arm drive control module uses Visual Studio 2019 to control the robotic arm, achieving scanning at different pressing forces F1 (1N, 1.5N, 2N) and different moving speeds V 1 (60mm / s, 80mm / s, 100mm / s, 120mm / s, 140mm / s), setting the length L of the object to be scanned 1 (300mm), setting the number of scans to 5 times. The objects to be scanned include A4 paper, leather, blanket, wood board, glass, and sandpaper.
[0098] In step (2) of the present invention, the signal acquisition module acquires signals, specifically including:
[0099] Real-time acquisition of three-dimensional acceleration data a through a triaxial accelerometer embedded inside the probe;
[0100] Real-time acquisition of three-dimensional friction force data f through a triaxial force sensor located below the object to be measured;
[0101] Real-time acquisition of audio signal s through a microphone placed in front of the object to be measured;
[0102] Then, the acquired data is transmitted to the signal processing unit.
[0103] In step (3) of the present invention, the acquired signals are received and displayed in real time, specifically including:
[0104] The signal processing unit receives the data (a, f, s) from the triaxial accelerometer, triaxial force sensor, and microphone. Subsequently, the time-domain diagrams of the triaxial accelerometer data, the time-domain diagrams of the triaxial force sensor data, and the time-domain diagrams of the microphone data are displayed on the display screen in real time to assist the operator in analyzing and judging the signals. By detecting the changes in the signal data in real time, the signal processing unit can discover abnormal situations and respond in a timely manner, thereby achieving all-round monitoring and diagnosis of the object to be measured.
[0105] In step (4) of the present invention, signal pre-emphasis, framing, windowing, and feature vector extraction are performed, specifically including:
[0106] Step 401, combining the x, y, and z components a x 、a y 、a z of the triaxial accelerometer data a into a total a', and combining the x, y, and z components f x 、f y 、f z of the triaxial force sensor data f into a total f', which is completed according to the following formula of L2 norm:
[0107]
[0108]
[0109] Among them, represents the magnitude at the j-th sampling point of the x-axis of the triaxial accelerometer data a, represents the magnitude at the j-th sampling point of the y-axis of the triaxial accelerometer data a, represents the magnitude at the j-th sampling point of the z-axis of the triaxial accelerometer data a; represents the magnitude at the j-th sampling point of the x-axis of the triaxial force sensor data f, represents the magnitude at the j-th sampling point of the y-axis of the triaxial force sensor data f, represents the magnitude at the j-th sampling point of the z-axis of the triaxial force sensor data f;
[0110] Step 402, perform pre-emphasis on the triaxial acceleration data a', triaxial force sensor data f', and microphone data s according to the following formula using a first-order high-pass filter to obtain a 1 , f 1 and s 1 :
[0111] a 1 (j) = b × a'(j) + (1 - b) × a 1 (j - 1)
[0112] f 1 (j) = b × f'(j) + (1 - b) × f 1 (j - 1)
[0113] s 1 (j) = b × s(j) + (1 - b) × s 1 (j - 1)
[0114] Among them, j represents the value at the j-th sampling point, b is the parameter of the filter, indicating that the output value of the current sampling point is affected by the output value of the previous sampling point to a certain extent, and the value range of b should be between 0 and 1;
[0115] Step 403, frame a, 1 , f 1 and s 1 according to a fixed time window, use a window with a length of 20 ms, and use a 50% overlap method to connect between frames, and apply a Hamming window to each frame of the signal;
[0116]
[0117]
[0118] a 1w (n) = a 1k(n)w(n)
[0119]
[0120]
[0121] f 1w (n) = f 1k (n)w(n)
[0122]
[0123]
[0124] s 1w (n) = s 1k (n)w(n)
[0125] where the window length is N, a 1k , f 1k , s 1k represent obtaining the k-th frame signal, w(n) represents the Hamming window function, a 1w , f 1w , s 1w represent the k-th frame obtained after windowing;
[0126] Step 404, for each frame signal a 1w , f 1w , s 1w after windowing, perform short-time Fourier transform to convert the time-domain signals a 1w , f 1w , s 1w to the frequency domain A 1w , F 1w , S 1w , and perform dot product of the Mel filter bank with A 1w , F 1w , S 1w to obtain the Mel spectral energy coefficients AE 1w (i), FE 1w (i), SE 1w (i) of the k-th frame. The Mel filter bank is as follows:
[0127]
[0128]
[0129]
[0130]
[0131] where the center frequency of the i-th Mel filter is f(i), Hi (m) represents the response value of the i-th Mel filter in terms of frequency, where m is the m-th triangular filter in the filter bank;
[0132] Step 405: Take the logarithm and perform the discrete cosine transform (DFT) to convert the 13 Mel frequency coefficients of each frame into 13 MFCC coefficients;
[0133] Step 406: Extract the MFCC feature matrix and retain the first 13 coefficients of the discrete cosine transform:
[0134]
[0135]
[0136]
[0137] where M is the number of Mel filter banks, l is the MFCC coefficient serial number, and AC k (l) is the MFCC coefficient of the triaxial accelerometer data, and PC k (l) is the MFCC coefficient of the triaxial force sensor data, and SC k (l) is the MFCC coefficient of the microphone data. They are arranged in the order of each frame, and the first 13 MFCC coefficients of each frame are arranged in order to form a feature vector, obtaining the triaxial accelerometer feature vector f 1 and the triaxial force sensor feature vector f 2 and the microphone data feature vector f 3 .
[0138] In step (5) of the present invention, linear weighted fusion is performed according to the feature vectors extracted in step (4), which specifically includes:
[0139] Calculate the signal-to-noise ratios of the triaxial acceleration data a, the triaxial force sensor data f, and the audio data s according to the following formula:
[0140]
[0141]
[0142]
[0143] where SNR a represents the signal-to-noise ratio of the triaxial acceleration data a, and SNR f represents the signal-to-noise ratio of the triaxial force sensor data f, and SNR s represents the signal-to-noise ratio of the microphone data s, and Ps a represents the signal power of the triaxial acceleration data a, and Ps fThe signal power of the three-axis force sensor data f, Ps s The signal power of the microphone data s, Pn a The noise power representing the three-axis acceleration data a, Pn f The noise power representing the three-axis force sensor data f, Pn s The noise power of the microphone data s, calculate the mean mu and standard deviation sigma of the three kinds of data (a, f, s); calculate the power of the signal and the noise, that is, Ps = mu 2 and Pn = sigma 2 ;
[0144] Calculate the weighted sum of the three-axis acceleration data eigenvector f 1 、the three-axis force sensor data eigenvector f 2 and the microphone data eigenvector f 3 by the following formula:
[0145] F = w 1 ×f 1 +w 2 ×f 2 +w 3 ×f 3
[0146]
[0147]
[0148]
[0149] where, w 1 represents the weight of the eigenvector f 1 w 2 represents the weight of the eigenvector f 2 w 3 represents the weight of the eigenvector f 3 SNR a represents the signal-to-noise ratio of the three-axis acceleration data a, SNR f represents the signal-to-noise ratio of the three-axis force sensor data f, SNR s represents the signal-to-noise ratio of the microphone data s, according to the determined weights, for f 1 、f 2 、f 3 perform weighted calculation to obtain the final fused eigenvector F.
[0150] Step (6) of the present invention specifically includes:
[0151] Divide the data in the sample library into a training set and a test set, and the ratio of the training set to the test set is 8:2;
[0152] The kernel function of the SVM algorithm selects the radial basis function (RBF) kernel, and uses cross-validation to adjust the regularization penalty parameter C and the parameter gamma that controls the width of the RBF kernel function;
[0153] Use the training set in the sample library to train the SVM algorithm to obtain a classifier model;
[0154] Use the test set in the sample library to evaluate the performance of the classifier model and calculate the accuracy of the classifier model;
[0155] Use the trained SVM classifier to classify newly collected sample data.
[0156] The present invention will be further described below in conjunction with the accompanying drawings and specific examples. It should be noted that these are not used to limit the scope of the disclosure of the present invention.
[0157] Figure 1 The following is a schematic diagram of an automated texture tactile and sound signal acquisition device according to the present invention, including: a robotic arm 101, a probe 102 embedded with a three-axis accelerometer, a three-axis force sensor 103, a microphone 104, and a computer 105. The specific working process is as follows:
[0158] The probe 102 embedded with a three-axis accelerometer is fixed on the robotic arm 101, and corresponding scanning parameters are set through the computer 105: scanning speed, pressing force, number of scans, and scanning length. Then the robotic arm 101 controls the probe 102 to scan on the three-axis force sensor 103 on which the planar texture sample is placed, and real-time collects the three-axis accelerometer data a, the three-axis force sensor data f, and the microphone data s. Finally, these data are sent to the computer 105 for processing and analysis. The computer 105 will display the time-domain waveforms of the three-axis accelerometer data a, the three-axis force sensor data f, and the microphone data s at the current moment in real time; the computer will perform filter pre-emphasis, frame segmentation, windowing, and extract 13-bit Mel-frequency cepstral coefficients (MFCCs) as feature vectors; then linearly weight and fuse the feature vectors of the three-axis accelerometer data, the three-axis force sensor data, and the microphone data; then use the SVM algorithm for classification processing.
[0159] Figure 2 It is a flow chart of the material surface texture classification method based on planar texture force tactile and sound signal fusion of the present invention. The method of the present invention includes steps 1 to 6;
[0160] Step 1, collect planar texture force tactile data and sound data when interacting with real textures, including: three-axis accelerometer data a, three-axis force sensor data f, microphone data s, as Figure 3 shown: The robotic arm 301 controls the probe 302 embedded with an accelerometer with a preset pressing force F 1, Scanning speed v 1 Scan the object to be measured along the x-axis direction. The probe 302 with an embedded accelerometer, the three-axis force sensor 303, and the microphone 304 record the scanning information of the scanning point 305(x 1 , y 1 ). Scan with different pressing forces F1 (1N, 1.5N, 2N) and different moving speeds V 1 (60mm / s, 80mm / s, 100mm / s, 120mm / s, 140mm / s), set the number of scans N 1 to 5, scan 6 different samples, including A4 paper, leather, blanket, wooden board, glass, and sandpaper, and obtain a total of 450 (3×5×5×6) pieces of data;
[0161] Step 2, The signal acquisition module includes a three-axis accelerometer embedded inside the probe, which records the acceleration information of the three axes when the probe scans the object to be measured, a three-axis force sensor located below the object to be measured, which records the friction information of the three axes during the scanning process, and a microphone placed in front of the object to be measured, which records the sound information generated during the scanning process;
[0162] In this step, the three-axis accelerometer data a is obtained in real time through the three-axis accelerometer 302 fixed on the robotic arm 301, and the sampling rate is 4096; the three-axis force sensor data f is obtained in real time through the three-axis force sensor 303 fixed below the planar texture sample, and the sampling rate is 4096; the microphone data s is obtained in real time through the microphone 304, and the sampling rate is 4096;
[0163] Step 3, Analyze the collected three-axis acceleration data a, three-axis force sensor data f, and microphone data s;
[0164] In this step, perform Fourier transform on the three-axis acceleration data a, as shown in Figure 4 (a), 401 is the acceleration signal spectrum diagram of the x-axis, as shown in Figure 4 (b), 402 is the acceleration signal spectrum diagram of the y-axis, as shown in Figure 4 (c), 403 is the acceleration signal spectrum diagram of the z-axis; display the three-axis force sensor data, microphone data, and three-axis accelerometer data, and by detecting the change of the signal data, discover abnormal situations and respond in a timely manner, as shown in Figure 5 , 501 is the acceleration signals of the x, y, and z axes collected by the three-axis accelerometer, 502 is the signal collected by the microphone, and 503 is the x, y, and z axis signals collected by the three-axis force sensor.
[0165] Step 4, Preprocess the three-axis accelerometer data a, three-axis force sensor data f, and microphone data s, frame, window, and extract feature vectors; specifically, as shown in Figure 6As shown in:
[0166] Step 401: Combine the x, y, and z components a x , a y , a z of the three-axis accelerometer data a into a total a', and combine the x, y, and z components f x , f y , f z of the three-axis force sensor data f into a total f'. This is completed according to the following L2 norm formula:
[0167]
[0168]
[0169] where represents the magnitude at the j-th sampling point on the x-axis of the three-axis accelerometer data a, represents the magnitude at the j-th sampling point on the y-axis of the three-axis accelerometer data a, represents the magnitude at the j-th sampling point on the z-axis of the three-axis accelerometer data a; represents the magnitude at the j-th sampling point on the x-axis of the three-axis force sensor data f, represents the magnitude at the j-th sampling point on the y-axis of the three-axis force sensor data f, represents the magnitude at the j-th sampling point on the z-axis of the three-axis force sensor data f;
[0170] Step 402: Perform pre-emphasis on the three-axis acceleration data a', three-axis force sensor data f', and microphone data s using the following first-order high-pass filter formula to obtain a 1 , f 1 , and s 1 :
[0171] a 1 (j) = b × a'(j) + (1 - b) × a 1 (j - 1)
[0172] f 1 (j) = b × f'(j) + (1 - b) × f 1 (j - 1)
[0173] s 1 (j) = b × s(j) + (1 - b) × s 1 (j - 1)
[0174] where j represents the value at the j-th sampling point, b is the parameter of the filter, indicating that the output value at the current sampling point is affected to a certain extent by the output value at the previous sampling point, and the value range of b should be between 0 and 1;
[0175] Step 403: Segment a 1 , f 1 , and s 1 into frames according to a fixed time window. Use a window with a length of 20 ms and connect the frames with 50% overlap. Apply a Hamming window to each frame of the signal;
[0176]
[0177]
[0178] a 1w (n) = a 1k (n)w(n)
[0179]
[0180]
[0181] f 1w (n) = f 1k (n)w(n)
[0182]
[0183]
[0184] s 1w (n) = s 1k (n)w(n)
[0185] where the window length is N, a 1k , f 1k , and s 1k represent obtaining the k-th frame of the signal, w(n) represents the Hamming window function, and a 1w , f 1w , and s 1w represent the k-th frame obtained after windowing;
[0186] Step 404: Perform a short-time Fourier transform on each frame of the windowed signal a 1w , f 1w , and s 1w to convert the time-domain signals a 1w , f 1w , and s 1w to the frequency domain A 1w , F 1w , and S 1w . Multiply the Mel filter bank with A 1w , F 1w , and S 1w to obtain the Mel spectral energy coefficients AE 1w (i), FE 1w(i), SE 1w (i), the Mel filter bank is as follows;
[0187]
[0188]
[0189]
[0190]
[0191] where the center frequency of the i-th Mel filter is f(i), and H i (m) represents the response value of the i-th Mel filter in frequency, and m is the m-th triangular filter in the filter bank;
[0192] Step 405: Take the logarithm and perform the discrete cosine transform DFT to convert the 13 Mel frequency coefficients of each frame into 13 MFCC coefficients;
[0193] Step 406: Extract the MFCC feature matrix and retain the first 13 coefficients of the discrete cosine transform:
[0194]
[0195]
[0196]
[0197] where M is the number of Mel filter banks, l is the MFCC coefficient serial number, and AC k (l) is the MFCC coefficient of the triaxial accelerometer data, and FC k (l) is the MFCC coefficient of the triaxial force sensor data, and SC k (l) is the MFCC coefficient of the microphone data, arranged in the order of each frame, and the first 13 MFCC coefficients of each frame are arranged in order as a feature vector to obtain the triaxial accelerometer feature vector f 1 、the triaxial force sensor feature vector f 2 、the microphone data feature vector f 3 ;
[0198] Step 5: Perform linear weighted fusion according to the feature vectors extracted in Step 4;
[0199] Calculate the signal-to-noise ratios of the triaxial acceleration data a, the triaxial force sensor data f, and the microphone data s according to the following formula:
[0200]
[0201]
[0202]
[0203] Among them, SNR a represents the signal-to-noise ratio of the three-axis acceleration data a, SNR f represents the signal-to-noise ratio of the three-axis force sensor data f, SNR s represents the signal-to-noise ratio of the microphone data s, Ps a represents the signal power of the three-axis acceleration data a, Ps f represents the signal power of the three-axis force sensor data f, Ps s the signal power of the microphone data s, Pn a represents the noise power of the three-axis acceleration data a, Pn f represents the noise power of the three-axis force sensor data f, Pn s the noise power of the microphone data s, calculate the mean mu and standard deviation sigma of the three kinds of data (a, f, s); calculate the power of the signal and the noise, that is, Ps = mu 2 and Pn = sigma 2 ;
[0204] Through the following formula, the weighted sum of the three-axis acceleration data eigenvector f 1 、the three-axis force sensor data eigenvector f 2 and the microphone data eigenvector f 3 :
[0205] F = w 1 × f 1 + w 2 × f 2 + w 3 × f 3
[0206]
[0207]
[0208]
[0209] Among them, w 1 represents the weight value of the eigenvector f 1 ,w 2 represents the weight value of the eigenvector f 2 ,w 3 represents the weight value of the eigenvector f 3 ,SNR a represents the signal-to-noise ratio of the three-axis acceleration data a, SNR f represents the signal-to-noise ratio of the three-axis force sensor data f, SNR sDenote the signal-to-noise ratio of the microphone data s. According to the determined weights, perform weighted calculations on f 1 and f 2 and f 3 to obtain the final fused feature vector F. Store the obtained final feature vector F together with the corresponding object to be measured in the sample library, and establish a sample library including A4 paper, leather, blanket, wooden board, glass, and sandpaper;
[0210] Step 6: Classify the data in the sample library into a training set and a test set. The ratio of the training set to the test set is 8:2. Sample the support vector machine (SVM) as the classifier, and sample the radial basis function (RBF) kernel function for classification. In the training stage, perform model training on the training set, select the best SVM parameter configuration, and then use this parameter configuration for SVM training to obtain a trained classifier; in the test stage, input the feature vector to be measured into the trained classifier for classification to obtain the classification result. The process is as Figure 7 shown.
[0211] In this step, sample the grid search method, facilitate multiple groups of possible parameter combinations, and use the cross-validation method to evaluate the performance under different parameter configurations. Select the best regularization parameter C and kernel function parameter gamma to obtain the optimal SVM texture force touch and sound recognition model. Finally, prepare 100 collected sample data as the data set in the classification experiment. Among them, 90 are used to train the classifier, and the remaining 10 are used to test the classifier. The normalized confusion matrix of the summary results of 10 tests is as Figure 8 shown. It can be seen that the classification accuracy is 89.8%.
[0212] Use the obtained texture force touch and sound recognition model to classify new objects to be measured.
[0213] In this step, after steps 1 and 2, collect the triaxial accelerometer signal, triaxial force sensor signal, and microphone signal of the probe scanning the object to be measured. After step 4, extract the features of the object to be measured. After step 5, obtain the linearly weighted fused feature vector. Finally, input the feature vector into the SVM texture force touch and sound recognition model to obtain the corresponding category of the object to be measured.
[0214] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for classifying material surface textures based on the fusion of planar texture force touch and sound signals, characterized in that, it includes the following steps: (1). The robotic arm control module uses the robotic arm to drive and control the probe to apply a pressing force to the object to be measured, and can achieve scanning at different moving speeds, scan the length of the object to be measured, and can perform multiple scans; (2). The signal acquisition module includes a triaxial accelerometer embedded inside the probe to record the acceleration information of the three axes when the probe scans the object to be measured, a triaxial force sensor located below the object to be measured to record the friction force information of the three axes during the scanning process, and a microphone placed in front of the object to be measured to record the sound signal generated during the scanning process; (3). The signal processing module receives the signals from the signal acquisition module and displays the time-domain diagrams of the signals from the triaxial accelerometer, triaxial force sensor, and microphone on the display screen in real time; (4). The signal processing module receives the signals from the signal acquisition module, performs data pre-emphasis using first-order high-pass filtering, frames, windows, and extracts Mel Frequency Cepstral Coefficients (MFCCs) as feature vectors; (5). Fuse the feature vectors extracted from the triaxial accelerometer data, triaxial force sensor data, and microphone data. Use linear weighted fusion to assign a weight to the feature vectors of the triaxial accelerometer, triaxial force sensor, and microphone, and then add them weighted to obtain the total feature vector. Store the total feature vector together with the corresponding object to be measured in the sample library; (6). The training of the classification model uses the Support Vector Machine (SVM) algorithm to classify and identify newly collected samples through the texture surface recognition model trained in the sample library.
2. The method for classifying material surface textures based on the fusion of planar texture force touch and sound signals according to claim 1, characterized in that, in the step (1), the probe of the robotic arm control module scans the surface of the object to be measured, specifically including: The robotic arm drive control module uses Visual Studio 2019 to control the robotic arm to apply a pressing force F to the object to be measured 1 , from 1 N to 2 N, at different moving speeds V 1 , from 60 mm / s to 140 mm / s for scanning, and set the length L of the object to be scanned 1 , from 200 mm to 300 mm, and set the number of scans N 1 .
3. The method for classifying material surface textures based on the fusion of planar texture force touch and sound signals according to claim 1, characterized in that, in the step (2), the signal acquisition module acquires signals, specifically including: obtain three-dimensional acceleration data a in real time through a triaxial accelerometer embedded inside the probe; obtain three-dimensional friction force data f in real time through a triaxial force sensor located below the object to be measured; obtain an audio signal s in real time through a microphone placed in front of the object to be measured; then transmit the collected data to the signal processing unit.
4. The method for classifying material surface textures based on the fusion of planar texture force touch and sound signals according to claim 1, characterized in that, in the step (3), receiving and displaying the collected signals in real time, specifically including: The signal processing unit receives data (a, f, s) from a triaxial accelerometer, a triaxial force sensor, and a microphone. Subsequently, the time-domain diagrams of the triaxial accelerometer data, the triaxial force sensor data, and the microphone data are displayed in real time on a display screen to assist the operator in analyzing and judging the signals. By detecting the changes in the signal data in real time, the signal processing unit discovers abnormal situations and responds in a timely manner, thereby achieving all-round monitoring and diagnosis of the object under test.
5. The method for classifying material surface textures based on the fusion of planar texture force touch and sound signals according to claim 1, characterized in that, in step (4), signal pre-emphasis, framing, windowing, and feature vector extraction are performed, specifically including: Step 401, combine the x, y, and z axis components a x 、a y 、a z of the three-axis accelerometer data a into a total a', and combine the x, y, and z axis components f x 、f y 、f z of the three-axis force sensor data f into a total f', and complete it according to the following formula L2 norm: Among them, represents the magnitude at the j-th sampling point on the x-axis of the triaxial accelerometer data a, represents the magnitude at the j-th sampling point on the y-axis of the triaxial accelerometer data a, represents the magnitude at the j-th sampling point on the z-axis of the triaxial accelerometer data a; represents the magnitude at the j-th sampling point on the x-axis of the triaxial force sensor data f, represents the magnitude at the j-th sampling point on the y-axis of the triaxial force sensor data f, represents the magnitude at the j-th sampling point on the z-axis of the triaxial force sensor data f; Step 402, perform pre-emphasis on the three-axis acceleration data a’, three-axis force sensor data f’ and microphone data s according to the following formula with a first-order high-pass filter to obtain a 1 , f 1 and s 1 : a 1 (j) = b × a’(j) + (1 - b) × a 1 (j - 1) f 1 (j) = b × f’(j) + (1 - b) × f 1 (j - 1) s 1 (j) = b × s(j) + (1 - b) × s 1 (j - 1) where j represents the value at the j-th sampling point, b is the parameter of the filter, indicating that the output value of the current sampling point is affected by the output value of the previous sampling point to a certain extent, and the value range of b should be between 0 and 1; Step 403, for a 1 , f 1 and s 1 Frame them according to a fixed time window, using a window of 20 ms in length, and connecting the frames in a 50% overlapping manner, and applying a Hamming window to each frame of the signal; a 1w (n) = a 1k (n)w(n) f 1w f(n) = 1k f(n)w(n) s 1w (n) = s 1k (n)w(n) Among them, the window length is N, a 1k , f 1k , s 1k represent obtaining the k-th frame signal, w(n) represents the Hamming window function, a 1w , f 1w , s 1w represent the k-th frame obtained after windowing; Step 404, for each windowed frame of signal a 1w 、f 1w 、s 1w perform short-time Fourier transform to convert the time-domain signals a 1w 、f 1w 、s 1w to the frequency domain A 1w 、F 1w 、S 1w , multiply the Mel filter bank with A 1w 、F 1w 、S 1w to obtain the Mel spectrum energy coefficients AE 1w (i), FE 1w (i), SE 1w (i) of the k-th frame. The Mel filter bank is given by the following formula; Among them, the center frequency of the i-th Mel filter is f(i), and H i (m) represents the response value of the i-th Mel filter in terms of frequency, where m is the m-th triangular filter in the filter bank; Step 405: Take the logarithm and perform a discrete cosine transform DFT to convert the 13 Mel frequency coefficients of each frame into 13 MFCC coefficients; Step 406: Extract the MFCC feature matrix and retain the first 13 coefficients of the discrete cosine transform: Among them, M is the number of Mel filter banks, l is the MFCC coefficient serial number, and AC k (l) is the MFCC coefficient of the triaxial accelerometer data, and FC k (l) is the MFCC coefficient of the triaxial force sensor data, and SC k (l) is the MFCC coefficient of the microphone data, arranged in the order of each frame, and the first 13 MFCC coefficients of each frame are arranged in order to form a feature vector, obtaining the triaxial accelerometer feature vector f 1 , the triaxial force sensor feature vector f 2 , and the microphone data feature vector f 3 .
6. The method for classifying material surface textures based on the fusion of planar texture force touch and sound signals according to claim 1, characterized in that, in step (5), linear weighted fusion is performed according to the feature vectors extracted in step (4), specifically including: Calculate the signal-to-noise ratios of the triaxial acceleration data a, the triaxial force sensor data f, and the audio data s according to the following formula: Among them, SNR a represents the signal-to-noise ratio of the three-axis acceleration data a, SNR f represents the signal-to-noise ratio of the three-axis force sensor data f, SNR s represents the signal-to-noise ratio of the microphone data s, Ps a represents the signal power of the three-axis acceleration data a, Ps f represents the signal power of the three-axis force sensor data f, Ps s The signal power of the microphone data s, Pn a represents the noise power of the three-axis acceleration data a, Pn f represents the noise power of the three-axis force sensor data f, Pn s The noise power of the microphone data s, calculate the mean mu and standard deviation sigma of the three kinds of data (a, f, s); calculate the power of the signal and the noise, that is, Ps = mu 2 and Pn = sigma 2 ; Calculate the weighted sum of the eigenvector f of the three-axis acceleration data, 1 the eigenvector f of the three-axis force sensor data, 2 and the eigenvector f of the microphone data 3 as follows: F = w 1 × f 1 + w 2 × f 2 + w 3 × f 3 Among them, w 1 represents the weight of the feature vector f 1 ; w 2 represents the weight of the feature vector f 2 ; w 3 represents the weight of the feature vector f 3 ; SNR a represents the signal-to-noise ratio of the three-axis acceleration data a; SNR f represents the signal-to-noise ratio of the three-axis force sensor data f; SNR s represents the signal-to-noise ratio of the microphone data s. According to the determined weights, f 1 , f 2 , f 3 are weighted and calculated to obtain the final fused feature vector F.
7. The method for classifying material surface textures based on the fusion of planar texture force touch and sound signals according to claim 1, characterized in that, step (6) specifically includes: Divide the data in the sample library into a training set and a test set, and the ratio of the training set to the test set is 8:2; Select the radial basis function RBF kernel for the SVM algorithm kernel function, and use cross-validation to adjust the regularization penalty parameter C and the parameter gamma that controls the width of the radial basis function kernel function; Use the training set in the sample library to train the SVM algorithm to obtain a classifier model; Use the test set in the sample library to evaluate the performance of the classifier model and calculate the accuracy of the classifier model; Use the trained SVM classifier to classify the newly collected sample data.
Citation Information
Patent Citations
Tactile measurement device and method
CN112146701A
Manipulator tactile data representation recognition method based on SVM (Support Vector Machine)
CN114898219A
Method and apparatus for identifying object material based on voice features
CN107545902A
Tactile mode recognition method based on kernel method
CN112668609A