Quadruped robot terrain classification method and system based on sound feature and IMU information fusion

By fusing sound features and IMU information, generating depth maps and inputting neural network models for terrain classification, the problem of poor performance of traditional visual perception in complex environments is solved, and the ability of four-legged robots to recognize complex terrain is improved.

CN120196988AActive Publication Date: 2025-06-24WUHAN UNIV OF TECH

Patent Information

Application Number
CN202510264446.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

Traditional four-legged robot perception methods mainly rely on visual perception, making it difficult to effectively classify terrain in low light or inclement weather conditions, and it is difficult to cope with the variability of complex terrain environments.

Method used

Using a method based on the fusion of sound characteristics and IMU information, the Mel spectrogram and IMU infographic are extracted by collecting audio signals and IMU signals, and splicing them in chronological order, and inputting them to the neural network model for terrain classification.

Benefits of technology

It improves the terrain recognition ability of four-legged robots in complex environments, enhances the ability to classify unstructured outdoor terrain, and can more accurately classify terrain under changing environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196988A_ABST
    Figure CN120196988A_ABST
Patent Text Reader

Abstract

The invention provides a quadruped robot terrain classification method and system based on sound feature and IMU information fusion, and relates to the field of terrain perception, and the method specifically comprises the steps: collecting an audio signal, processing the audio signal to obtain a Mel spectrogram, collecting an IMU signal, drawing a triaxial IMU signal time domain graph, dividing the time domain graph into a plurality of IMU information graphs, and obtaining a quadruped robot terrain classification result. The feature value of each time domain graph window is calculated, an IMU information graph corresponding to the Mel-frequency spectrogram in time relation can be established, then the Mel-frequency spectrogram and the IMU information graph are spliced according to the time sequence, and the spliced Mel-frequency spectrogram and IMU information graph are input into a neural network model for terrain classification; according to the method, the classification capability of the unstructured outdoor terrain is enhanced in a mode of fusing sound perception and ontology perception, and the spliced input image not only contains audio features, but also retains IMU information, so that a neural network model can learn from two different perception sources, terrain features can be comprehensively understood, and the classification efficiency of the unstructured outdoor terrain is improved. Therefore, the recognition capability of the complex environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of terrain perception, and in particular, to a method and system for classifying terrains of a quadruped robot based on the fusion of sound features and IMU information. Background Art

[0002] With the development of scientific research, the technology of quadruped robots has become increasingly perfect, the industrial scale and application prospects have been continuously expanded, and the terrain perception technology of quadruped robots has received more and more attention.

[0003] However, traditional perception means of quadruped robots mainly use visual perception for terrain classification, such as optical sensors like lidar and cameras. However, this method performs poorly under low light or bad weather conditions. In order to achieve autonomous detection of complex terrains by quadruped robots, relying solely on visual perception of terrain cannot enable quadruped robots to smoothly pass through complex terrains, and it is difficult to cope with the variability of complex terrain environments. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for classifying terrains of a quadruped robot based on the fusion of sound features and IMU information, so as to solve the problem that the prior art in the above-mentioned background art is difficult to cope with the variability of complex terrain environments.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for classifying terrains of a quadruped robot based on the fusion of sound features and IMU information, the steps include: collecting audio signals, dividing the audio signals into several audio segments, preprocessing the audio segments, and extracting audio features and converting them into Mel spectrograms; collecting IMU signals, plotting the time-domain diagrams of the three-axis IMU signals, dividing the time-domain diagrams into several time-domain diagram windows, calculating the feature values of each time-domain diagram window, and establishing an IMU information diagram; splicing the Mel spectrograms and the IMU information diagrams with corresponding time relationships in chronological order to obtain a depth diagram containing sound features and IMU features; constructing a neural network model, and inputting the depth diagram into the neural network model to extract feature vectors and perform terrain classification.

[0006] Optionally, the step of preprocessing the audio segments specifically includes: framing and windowing the audio segments, where there is an overlap of 70%-80% between adjacent frames.

[0007] Optionally, the step of extracting audio features and converting them into Mel spectrograms specifically includes: extracting the audio features of each audio segment, converting the time-domain signal of the audio features into a frequency-domain signal through short-time Fourier transform, and then generating a spectrogram; filtering the spectrogram with a Mel filter to obtain a Mel spectrogram.

[0008] Optionally, the step of dividing the time-domain graph into several time-domain graph windows specifically includes: dividing the time-domain graph of the IMU signal by using a sliding window method to obtain several time-domain graph windows, and splicing the time-domain graph windows as time frames in chronological order; where the start time of the i-th window is: t i = i·s, and the end time of the i-th window is: t i+1 = t i + ω; in the formula, the t i is the start time of the i-th window, the s is the step size of window sliding, the t i+1 is the end time of the i-th window, and the ω is the length of each window; the IMU signal corresponding to the i-th window is: a x (t i , t i+1 ) = [a x (t i ),..., a x (t i+1 )], a y (t i , t i+1 ) = [a y (t i ),..., a y (t i+1 )], a z (t i , t i+1 ) = [a z (t i ),..., a z (t i+1 )]; in the formula, the a x (t) is the x-direction IMU signal component, the a y (t) is the y-direction IMU signal component, the a z (t) is the z-direction IMU signal component, the a x (t i , t i+1 ) is the x-direction IMU signal component corresponding to the i-th window, the a y (t i , t i+1 ) is the y-direction IMU signal component corresponding to the i-th window, and the a z (t i , t i+1 ) is the z-direction IMU signal component corresponding to the i-th window.

[0009] Optionally, the step of calculating the eigenvalue of each time-domain graph window includes: normalizing the data of all obtained windows: a′ x (ti , t i+1 ) = log(a x (t i , t i+1 ) + 1), a′ y (t i , t i+1 ) = log(a y (t i , t i+1 ) + 1), a′ z (t i , t i+1 ) = log(a z (t i , t i+1 ) + 1); Wherein, the a x (t i , t i+1 ) is the x-direction IMU signal component corresponding to the i-th window, and the a y (t i , t i+1 ) is the y-direction IMU signal component corresponding to the i-th window, and the a z (t i , t i+1 ) is the z-direction IMU signal component corresponding to the i-th window, and the a′ x (t i , t i+1 ) is the normalized data of the x-direction IMU signal component corresponding to the i-th window, and the a′ y (t i , t i+1 ) is the normalized data of the y-direction IMU signal component corresponding to the i-th window, and the a′ z (t i , t i+1 ) is the normalized data of the z-direction IMU signal component corresponding to the i-th window; Calculate the mean value of the IMU signal for each time frame: Wherein, the μ x is the mean value of the IMU signal for each time frame in the x direction, and the μ y is the mean value of the IMU signal for each time frame in the y direction, and the μ z is the mean value of the IMU signal for each time frame in the z direction; Calculate the variance of the IMU signal for each time frame:

[0010] Wherein, the is the variance of the IMU signal for each time frame in the x direction, and the is the variance of the IMU signal for each time frame in the y direction, and the δ z2 is the variance of the IMU signal for each time frame in the z direction.

[0011] Optionally, the step of establishing the IMU information map specifically includes: using time as the horizontal axis, feature name as the vertical axis, and color depth as the feature value size to establish the IMU information map.

[0012] Optionally, the step of inputting the depth map into the neural network model to extract feature vectors and perform terrain classification specifically includes: adjusting the size and number of channels of the depth map and inputting it into the neural network model, training the neural network model by minimizing the cross-entropy loss, compressing the feature map output by the neural network model into a 1x1 feature vector through a global average pooling layer, and performing classification through a fully connected layer, and calculating the output class probability through a Softmax layer.

[0013] Optionally, the step of training the neural network model by minimizing the cross-entropy loss further includes: calculating and recording the loss and accuracy of the current model on the training set every time it is trained.

[0014] Optionally, the accuracy calculation formula is: where N is the number of samples, is an indicator function, when the predicted label is equal to the true label y i takes 1, otherwise takes 0.

[0015] On the other hand, the present invention also provides a quadruped robot terrain classification system based on the fusion of sound features and IMU information, including: an audio map generation module for collecting audio signals, dividing the audio signals into several audio segments, preprocessing the audio segments, and extracting audio features and converting them into a Mel spectrogram; an IMU map generation module for collecting IMU signals, plotting the time-domain diagram of the three-axis IMU signals, dividing the time-domain diagram into several time-domain diagram windows, calculating the feature values of each time-domain diagram window, and establishing an IMU information map; a splicing module for splicing the Mel spectrogram and the IMU information map with corresponding time relationships in chronological order to obtain a depth map containing sound features and IMU features; a terrain classification module for constructing a neural network model, inputting the depth map into the neural network model to extract feature vectors and perform terrain classification.

[0016] Compared with the prior art, the beneficial effects of the present invention are:

[0017] The present invention collects audio signals, obtains a Mel spectrogram after processing, and collects IMU signals to draw a time-domain diagram of the three-axis IMU signals. The time-domain diagram is divided into several to obtain an IMU information diagram, and the eigenvalue of each time-domain diagram window is calculated, so as to establish an IMU information diagram corresponding to the Mel spectrogram in terms of time relationship. Then, the Mel spectrogram and the IMU information diagram are spliced in chronological order and input into a neural network model for terrain classification. The present application enhances the classification ability of unstructured outdoor terrains by integrating sound perception and proprioception. The spliced input image contains both audio features and retains IMU information, enabling the neural network model to learn from two different perception sources, comprehensively understand terrain features, and thus improve the recognition ability for complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic flowchart of the method steps of the present invention.

[0019] Figure 2 It is a schematic diagram of the overall architecture of the method of the present invention.

[0020] Figure 3 It is a schematic diagram of the IMU information diagram of the present invention.

[0021] Figure 4 It is a schematic diagram of the depth map after splicing the Mel spectrogram and the IMU information diagram of the present invention.

[0022] Figure 5 It is a schematic diagram of the system function modules of the present invention.

[0023] In the figure, 10 - audio diagram generation module, 20 - IMU diagram generation module, 30 - splicing module, 40 - terrain classification module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Next, the solution of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0025] It should be noted that in the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances for the embodiments of this application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0026] Those skilled in the art of this technology can understand that unless specifically stated, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0027] Those skilled in the art of this technology can understand that unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0028] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0029] It should be noted that, without conflict, the embodiments and features in the embodiments of this application can be combined with each other. The following will refer to the drawings and combine with embodiments to detail this application.

[0030] Please refer to Figures 1 - 5 , a terrain classification method for a quadruped robot based on the fusion of sound features and IMU information of the present invention, the steps include:

[0031] S100. Collect the audio signal, divide the audio signal into several audio segments, preprocess the audio segments, and extract audio features and convert them into a Mel spectrogram.

[0032] Specifically, place the quadruped robot on different sample terrains, such as hard pavement, grassland or sandy land, to move; collect the sound signals generated when its legs interact with the terrain through a directional microphone fixed on the robot's body. Subsequently, perform short-time Fourier transform processing on the sound signals and calculate through a Mel filter bank. Finally, obtain the Mel spectrogram. It is found through exploration that the spectrogram of the audio signal has a good effect in environmental sound classification, and after filtering through the Mel filter bank on this basis, the Mel spectrogram is more sensitive to changes in the low-frequency part.

[0033] S200. Collect the IMU signal, draw the time-domain diagram of the three-axis IMU signal, divide the time-domain diagram into several time-domain diagram windows, calculate the eigenvalue of each time-domain diagram window, and establish an IMU information diagram.

[0034] Specifically, collect the IMU signal through the inertial measurement unit (IMU), draw the time-domain diagrams of the IMU signals in the x, y, and z directions, window the time-domain diagrams of the IMU signals, and then slide the window to obtain several time-domain diagram windows. By dividing the time-domain diagram windows and splicing the time-domain diagram windows as time frames in chronological order, the IMU information diagram can be corresponded to the Mel spectrogram in terms of time relationship.

[0035] Furthermore, calculate the mean and variance values of the IMU signals in the x, y, and z directions of each window. Take the mean and variance values of the IMU signals as the eigenvalues of the time-domain diagram window. Use time as the horizontal axis and the feature name as the vertical axis, and the color depth represents the numerical value of the eigenvalue to establish an IMU information diagram; where the feature name includes the mean and / or variance value of the IMU signal in the x, y, and z directions, and the numerical value of the eigenvalue includes the numerical size of the mean value of the IMU signal and the numerical size of the variance value of the IMU signal in the x, y, and z directions. In this application, by calculating the mean and variance values of the IMU signals in the x, y, and z directions of each window and using the mean and variance values for data normalization, the neural network can be trained and learned more efficiently, greatly improving the efficiency and stability of model training and learning.

[0036] S300. Splice the Mel spectrogram and the IMU information diagram with corresponding time relationships in chronological order to obtain a depth map containing sound features and IMU features.

[0037] Specifically, this application combines the signal data collected by the IMU and the audio signals collected by the directional microphone. After processing, an IMU information map and a Mel spectrogram are obtained. By fusing these two perception methods of proprioception and sound perception, the classification ability for unstructured outdoor terrains is enhanced, greatly increasing the accuracy of terrain perception of the quadruped robot.

[0038] S400. Construct a neural network model, input the depth map into the neural network model to extract feature vectors and perform terrain classification.

[0039] Specifically, in order to learn the spatial features in these depth maps, we use the deep residual network ResNet for training; this network can effectively extract features in the picture and finally output a 512-dimensional vector, and perform classification through a fully connected layer, using the softmax activation function to calculate the classification probability.

[0040] It can be understood that in the present invention, by collecting audio signals, a Mel spectrogram is obtained after processing, and by collecting IMU signals, a time-domain diagram of the three-axis IMU signals is drawn. The time-domain diagram is divided into several to obtain an IMU information map, and the eigenvalue of each time-domain diagram window is calculated, so as to establish an IMU information map corresponding to the Mel spectrogram in terms of time relationship. Then, the Mel spectrogram and the IMU information map are spliced in chronological order and input into the neural network model for terrain classification; this application enhances the classification ability for unstructured outdoor terrains by fusing the ways of sound perception and proprioception. The spliced input image contains both audio features and retains IMU information, so that the neural network model can learn from two different perception sources, can more comprehensively understand the terrain features, and thus improve the recognition ability for complex environments.

[0041] In some embodiments, the step of preprocessing the audio segment specifically includes: framing and windowing the audio segment, where there is an overlap of 70%-80% between two adjacent frames.

[0042] Specifically, for traditional audio framing, in order to ensure the continuity of sound, the overlap between two adjacent frames is about 50%. In this application, by setting an overlap of 70%-80% between two adjacent frames, not only can the signal continuity be ensured, but also some small signal changes can be better captured. Since the quadruped robot collects the sound of the foot contacting the ground, if the overlap is too low, it will be insensitive to the small changes in the sound. If the overlap is too high and the processed data is too large, it will affect the model performance. Therefore, an overlap of 70%-80% between two adjacent frames can capture more detailed sounds, greatly improving the accuracy of complex terrain perception.

[0043] In some embodiments, the step of extracting audio features and converting them into a Mel spectrogram specifically includes: extracting the audio features of each audio segment, converting the time-domain signal of the audio features into a frequency-domain signal through short-time Fourier transform, and then generating a spectrogram; filtering the spectrogram with a Mel filter to obtain a Mel spectrogram.

[0044] Specifically, first, the audio signals of various categories are segmented into small segments with a duration of t seconds, features are extracted from each segment and converted into a Mel spectrogram, which is used as a new sample for classification; the audio segments are frame-processed, divided into M frames, each frame has a length of N, and there is an overlap of 70%-80% between frames. Let x(n) be the original audio signal of each frame. Next, by windowing each frame signal and applying short-time Fourier transform STFT to obtain the spectrum, the formula is as follows: In the formula, the X(t, f) is a complex value in the time-frequency domain, representing the signal information at time t and frequency f, the x(n) is the sampling value of the original signal at the nth, w(n) is a window function, w(n - t) is a window centered at time t, and as time t changes, the window slides and always weights the signal around the current time t, the is a complex exponential, representing the rotation factor of the signal in the frequency domain, and the N is the total number of sampling points per frame.

[0045] Further, a Hanning window is used to reduce spectral leakage and Gibbs effect, and its calculation formula is as follows: In the formula, the w(n) is the Hanning window.

[0046] Further, calculate the power spectrum at linear frequencies, and its calculation formula is as follows: P(f) = |X(t, f)| 2 ; in the formula, the P(f) is the power spectrum, and the f is the linear frequency.

[0047] Further, in order to convert the frequency to the Mel frequency scale, the Mel frequency conversion formula is adopted: In the formula, the M(f) is the Mel frequency scale, and the f is the linear frequency.

[0048] Further, apply a Mel filter bank to map the result of the short-time Fourier transform to the Mel frequency, so as to obtain the energy value of each time frame at the Mel frequency scale, and its calculation formula is as follows: In the formula, f ∈ M means that the linear frequency f is mapped to the Mel frequency M.

[0049] Further, calculate the logarithmic form of the power spectrogram, and its calculation formula is as follows: S(t, M(f)) = 10·log 10(P(t, M(f))), which provides the basis for obtaining the Mel spectrogram and subsequent image stitching and feature input of the neural network model.

[0050] In some embodiments, the step of dividing the time-domain graph into a plurality of time-domain graph windows specifically includes: dividing the time-domain graph of the acceleration signal by using a sliding window method to obtain a plurality of time-domain graph windows, and splicing the time-domain graph windows as time frames in chronological order; wherein, the starting time of the i-th window is: t i = i·s, and the ending time of the i-th window is: t i+1 = t i + ω; in the formula, the t i is the starting time of the i-th window, the s is the step size of window sliding, the t i+1 is the ending time of the i-th window, and the ω is the length of each window; the acceleration signal corresponding to the i-th window is: a x (t i , t i+1 ) = [a x (t i ),..., a x (t i+1 )], a y (t i , t i+1 ) = [a y (t i ),..., a y (t i+1 )], a z (t i , t i+1 ) = [a z (t i ),..., a z (t i+1 )]; in the formula, the a x (t) is the acceleration component in the x direction, the a y (t) is the acceleration component in the y direction, the a z (t) is the acceleration component in the z direction, the a x (t i , t i+1 ) is the acceleration component in the x direction corresponding to the i-th window, the a y (t i , t i+1 ) is the acceleration component in the y direction corresponding to the i-th window, and the a z (t i , t i+1 ) is the acceleration component in the z direction corresponding to the i-th window.

[0051] Specifically, the IMU signals include acceleration signals, angular velocity signals, and velocity signals. In this application, the processing process of the acceleration signals will be listed, and the processing of other signals can be implemented similarly. In this application, the acceleration signals are collected by the inertial measurement unit (IMU), the time-domain diagrams of the acceleration signals in the x, y, and z directions are plotted, the time-domain diagrams of the acceleration signals are windowed, and then the window is slid to obtain a number of time-domain diagram windows. By dividing the time-domain diagram windows, the time-domain diagram windows are used as time frames and stitched together in chronological order, so that the IMU information diagram can be corresponded to the Mel spectrogram in terms of time relationship.

[0052] In some embodiments, the step of calculating the eigenvalue of each time-domain diagram window includes: normalizing the data of all the obtained windows: a′ x (t i ,t i+1 )=log(a x (t i ,t i+1 )+1), a′ y (t i ,t i+1 )=log(a y (t i ,t i+1 )+1), a′ z (t i ,t i+1 )=log(a z (t i ,t i+1 )+1); the a x (t i ,t i+1 ) is the x-direction IMU signal component corresponding to the i-th window, the a y (t i ,t i+1 ) is the y-direction IMU signal component corresponding to the i-th window, the a z (t i ,t i+1 ) is the z-direction IMU signal component corresponding to the i-th window, the a′ x (t i ,t i+1 ) is the normalized data of the x-direction IMU signal component corresponding to the i-th window, the a′ y (t i ,t i+1 ) is the normalized data of the y-direction IMU signal component corresponding to the i-th window, the a′ z (t i ,t i+1) is the normalized data of the z-direction IMU signal component corresponding to the i-th window; calculate the mean acceleration of each time frame: In the formula, the μ x is the mean value of the IMU signal of each time frame in the x direction, and the μ y is the mean value of the IMU signal of each time frame in the y direction, and the μ z is the mean value of the IMU signal of each time frame in the z direction; calculate the acceleration variance of each time frame: In the formula, the is the variance of the IMU signal of each time frame in the x direction, and the is the variance of the IMU signal of each time frame in the y direction, and the is the variance of the IMU signal of each time frame in the z direction.

[0053] Specifically, in this application, by calculating the mean value and variance value of the acceleration signals in the x, y, and z directions of each window, and using the mean value and variance value for data standardization, the neural network can train and learn more efficiently, greatly improving the efficiency and stability of model training and learning.

[0054] In some embodiments, the step of establishing the IMU information map specifically includes: using time as the horizontal axis, feature name as the vertical axis, and color depth as the feature value size to establish the IMU information map.

[0055] Specifically, the feature names include the mean value of the IMU signal and / or the variance value of the IMU signal in the x, y, and z directions, and the numerical values of the feature values include the numerical size of the mean value of the IMU signal and the numerical size of the variance value of the IMU signal in the x, y, and z directions; in this application, by calculating the mean value and variance value of the IMU signal in the x, y, and z directions of each window, and using the mean value and variance value for data standardization, the neural network can train and learn more efficiently, greatly improving the efficiency and stability of model training and learning.

[0056] In some embodiments, the step of inputting the depth map into the neural network model to extract feature vectors and perform terrain classification specifically includes: adjusting the size and number of channels of the depth map and inputting it into the neural network model, training the neural network model by minimizing the cross-entropy loss, compressing the feature map output by the neural network model into a 1x1 feature vector through the global average pooling layer, and performing classification through the fully connected layer, and calculating the output category probability through the Softmax layer.

[0057] Specifically, we first resize the obtained Mel spectrogram to 224×224 with 3 channels, which can enhance the expressive power of the model through multi-dimensional feature encoding while taking into account computational efficiency and generality.

[0058] Furthermore, the deep residual network ResNet-18 is trained by minimizing the cross-entropy loss; the neural network adopted in this application includes a 7x7 convolutional layer with a stride of 2, a 3x3 max pooling layer with a stride of 2, and four residual block layers, namely Layer 1 to Layer 4; each residual block consists of two convolutional layers and realizes the addition of the input and output through a shortcut connection.

[0059] Specifically, by minimizing the cross-entropy loss in this application, the predicted probability distribution of the neural network model can be made as close as possible to the distribution of the true labels, thereby improving the classification accuracy; and the global features of the image are captured through larger convolutional layers, and the output feature map is reduced to half of the input image by moving the sliding window two grids each time, and the image size is further compressed through the max pooling layer to retain the most significant features in the local area and enhance the robustness of the neural network model; the four residual blocks can improve the information flow efficiency and achieve efficient feature learning through deep stacking and each residual block reusing the features of the previous layer through a shortcut.

[0060] Furthermore, the neural network model compresses the feature map into a 1x1 feature vector through a global average pooling layer and performs classification through a fully connected layer, and the output class probability is calculated through a Softmax layer.

[0061] Specifically, compressing the multi-dimensional feature map into a 1x1 feature vector, classifying the feature vector after global average pooling through a fully connected layer, and calculating the output class probability through a Softmax layer can provide an intuitive and optimizable basis for the classification task.

[0062] Furthermore, let fi(si,θ) be the activation value of spectrogram si and class j, and θ be the parameters (weights W and biases b) of the network. The Softmax function and the cross-entropy loss function are calculated as follows: In the formula, P(y=j|s i ; θ) is the predicted probability that spectrogram s i belongs to the j-th class, and the loss can be calculated as:

[0063] In some embodiments, the step of training the neural network model by minimizing the cross-entropy loss further includes: calculating and recording the loss and accuracy of the current model on the training set every time it is trained.

[0064] Specifically, the present application also dynamically adjusts the learning rate of each parameter after each training of the neural network model; during the training process, the learning rate is set to 0.0001, which is a relatively small initial value. During one iteration (epoch) of each training data set, the neural network model performs forward propagation on the input data, calculates the loss between the predicted result and the true label, and then calculates the gradient through backpropagation and dynamically updates the model parameters; after one iteration of each training data set ends, the loss and accuracy of the current model on the training set are calculated and recorded, thereby improving the accuracy of recognizing complex environments.

[0065] In some embodiments, the accuracy calculation formula is: In the formula, N is the number of samples, is an indicator function. When the predicted label is equal to the true label y i it takes 1, otherwise it takes 0.

[0066] Specifically, by calculating and recording the accuracy of the current model on the training set, the present application can guide the optimization of hyperparameters. If the accuracy stagnates for a long time, the learning rate can be reduced, enabling the neural network model to converge more precisely. When the accuracy no longer improves for multiple consecutive epochs, the training is terminated in advance.

[0067] On the other hand, the present invention also provides a quadruped robot terrain classification system based on the fusion of sound features and IMU information, including: an audio map generation module 10, configured to collect audio signals, divide the audio signals into several audio segments, preprocess the audio segments, and extract audio features and convert them into Mel spectrograms; an IMU map generation module 20, configured to collect IMU signals, plot the time-domain diagrams of the three-axis IMU signals, divide the time-domain diagrams into several time-domain diagram windows, calculate the eigenvalue of each time-domain diagram window, and establish an IMU information map; a splicing module 30, configured to splice the Mel spectrograms and the IMU information maps corresponding to the time relationship in chronological order to obtain a depth map containing sound features and IMU features; a terrain classification module 40, configured to construct a neural network model, input the depth map into the neural network model to extract feature vectors and perform terrain classification.

[0068] The technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0069] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-described embodiment methods, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-described methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0070] The above are only the embodiments of the present invention and do not limit the patent scope of the present invention. All equivalent transformations made by using the content of the specification and drawings of the present invention, directly or indirectly applied in the related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. A quadruped robot terrain classification method based on sound feature and IMU information fusion, characterized in that the steps include: Collecting audio signals, dividing the audio signals into a plurality of audio segments, preprocessing the audio segments, and extracting audio features and converting them into mel-spectrograms; Collect IMU signals, draw a time domain graph of the three-axis IMU signals, divide the time domain graph into a number of time domain graph windows, calculate the eigenvalue of each time domain graph window, and establish an IMU information graph; The Mel spectrum graph and IMU information graph of the corresponding time relationship are spliced ​​in chronological order to obtain a depth map containing sound features and IMU features; A neural network model is constructed, and the depth map is input into the neural network model to extract feature vectors and perform terrain classification.

2. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 1 is characterized in that: The step of preprocessing the audio clip specifically includes: The audio segment is framed and windowed, wherein two adjacent frames overlap by 70%-80%.

3. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 1 is characterized in that: The step of extracting audio features and converting them into a Mel-spectrogram specifically includes: Extracting audio features of each audio clip, converting the time domain signal of the audio features into a frequency domain signal by short-time Fourier transform, and then generating a spectrogram; The spectrogram is filtered using a Mel filter to obtain a Mel spectrogram.

4. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 1 is characterized in that: The step of dividing the time domain graph into a plurality of time domain graph windows specifically includes: The time domain image of the IMU signal is divided by a sliding window method to obtain a plurality of time domain image windows, and the time domain image windows are spliced ​​as time frames in chronological order; Among them, the starting time of the i-th window is: t i =i·s, the end time of the i-th window is: t i+1 =t i +ω; In the formula, the t i is the starting time of the i-th window, s is the step size of the window sliding, and t i+1 is the end time of the i-th window, and ω is the length of each window; The IMU signal corresponding to the i-th window is: a x (t i , t i+1 )=[a x (t i ), ..., a x (t i+1 )], In the formula, the a x (t) is the IMU signal component in the x direction, and the a y (t) is the IMU signal component in the y direction, and the a z (t) is the IMU signal component in the z direction, and the a x (t i , t i+1 ) is the x-direction IMU signal component corresponding to the i-th window, and the a y (t i , t i+1 ) is the y-direction IMU signal component corresponding to the i-th window, and the a z (t i , t i+1 ) is the z-direction IMU signal component corresponding to the i-th window.

5. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 4 is characterized in that: The step of calculating the characteristic value of each time domain image window comprises: Normalize the data of all windows obtained: a′ x (t i , t i+1 )=log(a x (t i , t i+1 )+1), a′ y (t i , t i+1 )=log(a y (t i , t i+1 )+1), a′ z (t i , t i+1 )=log(a z (t i , t i+1 )+1); In the formula, the a x (t i , t i+1 ) is the x-direction IMU signal component corresponding to the i-th window, and the a y (t i , t i+1 ) is the y-direction IMU signal component corresponding to the i-th window, and the a z (t i , t i+1 ) is the z-direction IMU signal component corresponding to the i-th window, and the a′ x (t i , t i+1 ) is the normalized data of the IMU signal component in the x direction corresponding to the i-th window, and the a′ y (t i , t i+1 ) is the normalized data of the IMU signal component in the y direction corresponding to the i-th window, and the a′ z (t i , t i+1 ) is the normalized data of the IMU signal component in the z direction corresponding to the i-th window; Calculate the mean IMU signal for each time frame: In the formula, the μ x is the mean value of the IMU signal in each time frame in the x direction, and μ y is the mean value of the IMU signal in each time frame in the y direction, and μ z is the mean value of the IMU signal in each time frame in the z direction; Calculate the IMU signal variance for each time frame: In the formula, is the IMU signal variance of each time frame in the x direction, is the IMU signal variance of each time frame in the y direction, is the IMU signal variance for each time frame in the z direction.

6. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 1 is characterized in that: The steps of establishing the IMU information graph specifically include: An IMU information graph is established with time as the horizontal axis, feature name as the vertical axis, and color depth as the feature value.

7. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 1 is characterized in that: The step of inputting the depth map into the neural network model to extract feature vectors and perform terrain classification specifically includes: The size and number of channels of the depth map are adjusted and input into the neural network model, the neural network model is trained by minimizing the cross entropy loss, the feature map output by the neural network model is compressed into a 1x1 feature vector through a global average pooling layer, and classified through a fully connected layer, and the output category probability is calculated through a Softmax layer.

8. The method for terrain classification of a quadruped robot based on the fusion of sound features and IMU information according to claim 7, characterized in that: The step of training the neural network model by minimizing the cross entropy loss also includes: After each training, the loss and accuracy of the current model on the training set are calculated and recorded.

9. The quadruped robot terrain classification method based on sound feature and IMU information fusion according to claim 8, characterized in that: The accuracy calculation formula is: In the formula, N is the number of samples, is an indicator function, when the predicted label and the true label y i If they are equal, it is 1, otherwise it is 0.

10. A quadruped robot terrain classification system based on the fusion of sound features and IMU information, characterized in that: include: An audio graph generation module is used to collect audio signals, divide the audio signals into a number of audio segments, pre-process the audio segments, and extract audio features and convert them into mel-spectrograms; An IMU map generation module is used to collect IMU signals, draw a time domain map of the three-axis IMU signal, divide the time domain map into a number of time domain map windows, calculate the eigenvalue of each time domain map window, and establish an IMU information map; A splicing module is used to splice the Mel spectrum map and the IMU information map of the corresponding time relationship in chronological order to obtain a depth map containing sound features and IMU features; The terrain classification module is used to construct a neural network model, input the depth map into the neural network model to extract feature vectors and perform terrain classification.

Citation Information

Patent Citations

  • Footstep detecting method according to acceleration and audio information

    CN106531186A

  • Vibration information terrain classification and identification method based on CNN-LSTM

    CN110956154A

  • Terrain recognition method and device based on sound interaction between mobile robot and ground

    CN112288870A

  • Transform-based multi-sensor data fusion robot terrain sensing method

    CN117409264A

  • Fusion feature classification and identification method based on noise reduction underwater acoustic signals

    CN118351881A

Cited By

  • Terrain recognition method, device and equipment for foot-type robot and medium

    CN121682451A