Gait recognition method and device based on ultrasonic sensing, medium, program product and terminal

Through the gait recognition method based on ultrasonic perception, the twin neural network is trained using ultrasonic reflected signals and historical gait signals, which solves the problems of low recognition accuracy and insufficient recognition ability of foreign individuals in complex environments, and achieves efficient and accurate gait recognition and rapid expansion of new users.

CN119939210APending Publication Date: 2025-05-06SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411893084.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing gait recognition technology is sensitive to environmental changes and has limited identification capabilities for foreign individuals, making it difficult to achieve efficient and accurate identity verification in complex scenarios.

Method used

The gait recognition method based on ultrasonic perception is adopted. By transmitting ultrasonic waves and collecting reflected signals, the initial gait signal and multiple historical gait signals are obtained in real time. The positive and negative sample pairs are generated based on these signals, and the model is trained through the twin neural network to generate the updated gait recognition results.

Benefits of technology

It significantly improves the accuracy and stability of gait recognition, especially in complex and dynamic environments, shows good adaptability, can achieve rapid addition and accurate identification of new users under low sample size conditions, and has strong outsider identification and exclusion capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939210A_ABST
    Figure CN119939210A_ABST
Patent Text Reader

Abstract

The invention provides a gait recognition method and device based on ultrasonic sensing, a medium, a program product and a terminal, gait information is collected through a microphone and a loudspeaker of common commercial audio equipment, and feature extraction and recognition are performed by using a convolutional neural network. And then realizing user expansion and identification under a small sample condition through a twin neural network multiplexing pre-training feature extraction layer and introduction of a data placeholder generation technology. The problems that an existing gait recognition technology mainly depends on visual information and is easily influenced by illumination and background changes, and large limitation exists in the aspect of foreign individual recognition are solved. High accuracy and stability are kept in a complex dynamic environment, the new user expansion and foreign person recognition capability is effectively improved, meanwhile, the implementation threshold is reduced through use of common audio equipment, and the overall recognition efficiency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gait recognition, and in particular to a gait recognition method, device, medium, program product and terminal based on ultrasonic sensing. Background Art

[0002] Smart space refers to a work or living environment that integrates computing devices and multimodal sensors, aiming to provide users with convenient interactive experience and personalized services through advanced technology. With the rapid development of the Internet of Things and human-computer interaction technology, contactless identity recognition technology has gradually become a research focus, especially in terms of ensuring security. Traditional methods of identity recognition usually include knowledge-based (such as passwords), object-based (such as tokens) and biometric-based (such as fingerprints, facial features, etc.) methods. Knowledge-based and object-based methods are prone to the risk of being forgotten, lost or stolen, while biometric-based identity authentication methods can be divided into physiological characteristics and biological behavior patterns. Although the physiological feature method provides convenience, it can usually only achieve one-time authentication and requires users to actively cooperate under specific conditions.

[0003] In contrast, identity authentication methods based on biological behavior patterns (such as gait recognition) have gradually gained attention due to their advantages of non-contact, difficult to disguise, and no need for active cooperation from users. As a dynamic feature, gait has individual specificity and stability, and can achieve secure authentication within a certain distance, providing an ideal continuous identity authentication solution for smart spaces. In the field of gait recognition, existing solutions usually rely on image processing and algorithm analysis to extract relevant information by observing the gait characteristics of individuals.

[0004] However, existing technologies still have limitations in terms of accuracy and applicability. First, some technologies are sensitive to environmental changes (such as changes in lighting and viewing angle), resulting in reduced recognition accuracy. Especially in complex scenarios (such as crowded areas), the system often has difficulty in achieving fast and high-accuracy identity authentication, resulting in misjudgments and missed judgments. In addition, the expansion of new users of these technologies requires a large number of samples, and it is difficult to achieve accurate recognition under small sample conditions. The recognition ability of foreign individuals is limited, which increases the challenges in data security and accuracy, and limits its effectiveness in real-world applications. Summary of the invention

[0005] In view of the shortcomings of the prior art mentioned above, the purpose of the present application is to provide a gait recognition method, device, medium, program product and terminal based on ultrasonic sensing, which is used to solve the problem that the shortcomings of the prior art are that it is sensitive to environmental changes and has limited ability to recognize foreign individuals, making it difficult to achieve efficient and accurate identity authentication in complex scenarios.

[0006] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides a gait recognition method based on ultrasonic sensing, including: by emitting ultrasonic waves and collecting their reflected signals, obtaining the initial gait signal to be detected in real time; simultaneously obtaining multiple historical gait signals; based on the multiple historical gait signals, pre-training the gait recognition model, and extracting the first feature extraction network from the trained gait recognition model; based on the initial gait signal and the multiple historical gait signals, generating multiple groups of positive and negative sample pairs; inputting the multiple groups of positive and negative sample pairs into a twin neural network including a first feature extraction network for model training to generate an updated twin neural network; and generating a gait recognition result corresponding to the initial gait signal through the updated twin neural network.

[0007] In some embodiments of the first aspect of the present application, the multiple groups of positive and negative sample pairs include a first type of positive and negative sample pairs and a second type of positive and negative sample pairs; the first type of positive and negative sample pairs are composed of one or more gait signal combinations in historical gait signals; the second type of positive and negative sample pairs are composed of an initial gait signal and any combination of historical gait signals, and sample pairs are composed of an initial gait signal and a combination of an initial gait signal.

[0008] In some embodiments of the first aspect of the present application, the process of inputting multiple groups of positive and negative sample pairs into a twin neural network including a first feature extraction network for model training includes: inputting the first type of positive and negative sample pairs and the second type of positive and negative sample pairs into the twin neural network including the first feature extraction network, and performing an optimization operation on the first feature extraction network in the twin neural network based on the contrast loss function to update it to a twin neural network including a second feature extraction network.

[0009] In some embodiments of the first aspect of the present application, the process of generating a gait recognition result corresponding to the initial gait signal through an updated twin neural network includes: replacing the first feature extraction network in the trained gait recognition model with the second feature extraction network to generate an optimized gait recognition model; inputting the initial gait signal into the optimized gait recognition model to generate an optimized classification probability; if the highest probability value in the optimized classification probability is lower than a first threshold, calculating the centroid similarity of each category; if the centroid similarities of all categories are lower than a second threshold, marking the gait category of the initial gait signal as an outsider.

[0010] In some embodiments of the first aspect of the present application, after acquiring multiple historical gait signals, the following operation is also performed: performing linear interpolation operations on the multiple historical gait signals, and inserting randomly generated placeholder data into the interpolation results to simulate potential outsider characteristics.

[0011] In some embodiments of the first aspect of the present application, an initial gait signal to be detected is acquired in real time by emitting ultrasonic waves and collecting their reflected signals; after acquiring multiple historical gait signals at the same time, the following operations are also performed: one or more of mixing operations, filtering operations, frequency domain conversion operations, and outlier removal operations are performed on the initial gait signal and multiple historical gait signals in sequence.

[0012] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides a gait recognition device based on ultrasonic sensing, including: a data acquisition module: used to acquire the initial gait signal to be detected in real time by emitting ultrasonic waves and collecting their reflected signals; and simultaneously acquire multiple historical gait signals; a pre-training module: used to pre-train the gait recognition model based on the multiple historical gait signals, and extract the first feature extraction network from the trained gait recognition model; a gait recognition module: used to generate multiple groups of positive and negative sample pairs based on the initial gait signal and the multiple historical gait signals; input the multiple groups of positive and negative sample pairs into a twin neural network including a first feature extraction network for model training to generate an updated twin neural network; and generate a gait recognition result corresponding to the initial gait signal through the updated twin neural network.

[0013] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the gait recognition method based on ultrasonic sensing is implemented.

[0014] To achieve the above-mentioned purpose and other related purposes, the fourth aspect of the present application provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer implements the gait recognition method based on ultrasonic sensing.

[0015] To achieve the above-mentioned purpose and other related purposes, the fifth aspect of the present application provides an electronic terminal, including a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the gait recognition method based on ultrasonic sensing.

[0016] As described above, the ultrasonic sensing-based gait recognition method, device, medium, program product and terminal of the present application have the following beneficial effects: significantly improving the accuracy and stability of gait recognition, especially showing good adaptability in complex and dynamic environments. The complexity of user use is reduced by collecting gait data through ordinary audio equipment, and the demand for training samples is effectively reduced by reusing the pre-trained feature extraction layer, thereby optimizing the overall efficiency. Under low sample volume conditions, it can achieve rapid addition and accurate identification of new users, and at the same time has powerful outsider identification and exclusion capabilities. This technology is widely used in many fields such as security monitoring, health monitoring and smart home, further expanding the application prospects of gait recognition technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flow chart of an embodiment of a gait recognition method based on ultrasonic sensing of the present application is shown.

[0018] Figure 2 The figure shows the time-frequency diagram of the preprocessed gait signal in an embodiment of the gait recognition method based on ultrasonic sensing of the present application.

[0019] Figure 3 A flow chart of another embodiment of the gait recognition method based on ultrasonic sensing of the present application is shown.

[0020] Figure 4 A structural schematic diagram of an embodiment of a gait recognition device based on ultrasonic sensing of the present application is shown.

[0021] Figure 5 A structural schematic diagram of an embodiment of a gait recognition terminal based on ultrasonic sensing of the present application is shown. DETAILED DESCRIPTION

[0022] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0023] Before further describing the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are applicable to the following interpretations:

[0024] <1> Ultrasonic reflection signal: When using ultrasound for measurement or detection, the signal formed by the ultrasound wave propagating from the transmitting source to the target object and reflecting back contains the characteristic information of the object being measured, such as shape, material and distance.

[0025] <2> Gait signal: The signal generated by an individual during walking, usually contains information such as body movement, footstep frequency, stride, speed and direction, and also includes the difference frequency signal related to the distance, speed and gait of the human body caused by ultrasonic reflection.

[0026] <3> Difference frequency signal: The frequency difference between two signals in the mixing process, which is often used to extract specific frequency components. For gait signals, the difference frequency signal provides information about the dynamic and motion characteristics of the individual gait.

[0027] <4> Mixing operation: The process of generating new frequency components by multiplying or combining two signals of different frequencies. In this operation, the sum frequency (i.e., frequency addition) and difference frequency (i.e., frequency subtraction) of the two frequencies are usually generated. Mixing operations are widely used in wireless communications, radar systems, and audio processing.

[0028] <5> Butterworth bandpass filter: It has a smooth frequency response and can effectively pass signals within a specific frequency range while suppressing other frequency components. Its characteristics make it widely used in many signal processing applications.

[0029] <6> Time-frequency diagram: An image that simultaneously displays the signal information in the time and frequency domains, revealing the pattern of how the signal's spectral characteristics change over time. It is often used for signal analysis and feature extraction.

[0030] <7> .npz format: A data storage file format, commonly used to save numpy arrays. It supports storing multiple arrays in the same file, making it easier to organize and read data.

[0031] <8> Data placeholder generation: In machine learning, a technique used to expand a training dataset, typically creating synthetic samples, to improve the generalization and robustness of the model.

[0032] <9> Frequency Modulated Continuous Wave (FMCW): A modulation technique that linearly modulates the frequency of the transmitted signal (i.e., the frequency changes over time) to improve distance resolution. It is widely used in radar and ultrasonic applications.

[0033] <10> Chirp signal: A signal whose frequency changes over time. Chirp signals can improve the distance resolution and velocity resolution of target detection by sending pulses with constantly changing frequencies. This is because when receiving the return signal, the frequency change can provide richer Doppler information, which helps to identify and estimate the position and velocity of the target.

[0034] <11> Doppler shift: The change in wave frequency due to the relative motion between the observer and the source of the wave, often used to measure the speed of a moving target and is widely used in radar, ultrasound and sonar technology.

[0035] <12> Convolutional Neural Network (CNN): A deep learning model that is particularly suitable for image and sequence data, which extracts local features through convolutional layers and then performs classification or regression tasks through fully connected layers.

[0036] <13> Twin network: It consists of two or more sub-networks that share weights and have the same structure and parameters to ensure consistent feature representation when processing different inputs. Each sub-network receives different input samples and extracts features through the same neural network structure at the same time. The key to the twin network is that the feature vectors of its output are compared by calculating the distance, thereby learning the similarity measure between the inputs. This design makes the twin network particularly suitable for tasks that require judging the similarity or difference of input data.

[0037] <14> Feature extraction layer: In a neural network, a layer responsible for extracting high-level features from input data, usually including convolutional layers and pooling layers, to facilitate subsequent more complex processing.

[0038] <15> Fully connected layer: A layer in a neural network that receives the output of all neurons in the previous layer, multiplies it by its own weights, and then sums them up. It is usually used to integrate features and output the final result.

[0039] <16> Positive and negative sample pairs: In contrastive learning, sample pairs consisting of positive samples (similar or similar data) and negative samples (different types of data) are used to train the model to improve feature discrimination capabilities.

[0040] <17> Contrastive loss function: A loss function used to evaluate the performance of the model in processing positive and negative sample pairs, which optimizes the learning process of the model by minimizing the distance between positive sample pairs and maximizing the distance between negative sample pairs.

[0041] To facilitate understanding of the embodiments of the present application, first Figure 1 Detailed description. Figure 1The following is a flow chart of a gait recognition method based on ultrasonic sensing in an embodiment of the present invention. The gait recognition method based on ultrasonic sensing in this embodiment mainly includes the following steps:

[0042] Step S11: Acquire the initial gait signal to be detected in real time by emitting ultrasonic waves and collecting their reflected signals; and acquire multiple historical gait signals at the same time.

[0043] In one embodiment of the present invention, an initial gait signal to be detected is acquired in real time by emitting ultrasonic waves and collecting their reflected signals; after acquiring multiple historical gait signals at the same time, the following operations are performed: one or more of mixing operations, filtering operations, frequency domain conversion operations, and outlier removal operations are performed on the initial gait signal and the multiple historical gait signals in sequence.

[0044] In this embodiment, the collected gait signal is mixed to extract the difference frequency signal from the original signal, and the difference frequency signal contains information related to the distance, speed and gait characteristics of the human body. In the mixing process, the frequency of the signal is effectively reduced by mixing and synchronizing the signal with the reference signal, and low-frequency information related to the target movement is obtained. By selecting a suitable carrier frequency, the high-frequency signal can be more accurately converted into the difference frequency signal, making the gait characteristics therein more obvious and easy to extract and analyze later.

[0045] Subsequently, the mixed difference frequency signal is filtered using a Butterworth bandpass filter to eliminate high-frequency and low-frequency noise components in the signal. The Butterworth bandpass filter can provide a smooth passband and set a specific upper and lower frequency range to ensure that only the key frequency components of the gait signal are retained. In this embodiment, the focus of the filter parameter setting is to cover the frequency band related to the gait characteristics of the human body. This step greatly reduces the impact of non-gait components such as environmental noise and equipment interference on the signal, and improves the purity and reliability of the data.

[0046] Next, the filtered signal is processed by fast Fourier transform (FFT) to convert the time domain signal into a frequency domain signal. Through FFT, the complex time domain waveform can be decomposed into the amplitude and phase of each frequency component, so as to analyze the spectral characteristics of the signal. In this step, the frequency characteristic information closely related to the gait characteristics is extracted, such as the in-phase component (I) and orthogonal component (Q) of the target signal, to characterize the periodicity, amplitude and dynamic changes of the gait movement.

[0047] Preferably, after the frequency domain signal is extracted, in order to further improve the credibility and accuracy of the data, the signal is detected and processed for outliers. A method combining the mean and standard deviation is used to identify and remove outliers in the signal, such as equipment interference burst signals, noise generated by non-target motion, or abnormal sampling points that appear during data acquisition. By removing these abnormal data, it is possible to effectively avoid interference with gait feature extraction and analysis, and maintain the integrity and consistency of the signal data.

[0048] Finally, the processed signal is stored in a compressed format. After the above mixing, filtering, frequency domain analysis and outlier removal, the processed signal data is integrated and saved as a .npz format file. This file contains the main in-phase (I) and quadrature (Q) component information, which is convenient for subsequent gait recognition tasks and model training. By adopting a compressed storage format, not only storage space is saved, but also the process of data loading and reading is simplified. The time-frequency diagram after processing and saving is shown in the figure below. Figure 2 As shown, the main characteristic content and spectrum distribution of the signal are displayed.

[0049] It should be noted that the historical gait signals are also collected through the above steps. For example, gait data of 10 different subjects are collected, and labels are assigned to each subject sample. The acquisition method of these historical gait signals is consistent with the initial gait signal acquisition operation and preprocessing operation of real-time collection, thereby ensuring the consistency and reliability of the data.

[0050] In one embodiment of the present invention, after acquiring the multiple historical gait signals, the following operation is further performed: performing a linear interpolation operation on the multiple historical gait signals, and inserting randomly generated placeholder data into the interpolation results to simulate the characteristics of potential outsiders.

[0051] In this embodiment, a data placeholder generation method is innovatively introduced to simulate unknown class samples, i.e., outsider data. This embodiment uses linear interpolation technology to process existing historical gait signals. Specifically, linear interpolation operations are performed on various types of historical gait signals to generate new sample points. During interpolation, interpolation processing is performed between feature points of two or more known gait signals, and new sample features located between these known feature points are calculated using a specific interpolation algorithm (e.g., linear interpolation or Lagrange interpolation). In order to enhance the diversity of the data set, randomly generated placeholder data are inserted into the interpolated data set. The generation process of these placeholder data includes randomly selecting some feature values, the range of which is similar to the actual gait signal, to ensure that the placeholder data can cover the potential unknown feature space. This process is implemented by a random number generation algorithm to maintain the rationality of the signal features. By combining linear interpolation with placeholder data insertion, the generated enhanced data set provides more training samples to support subsequent gait recognition tasks.

[0052] In one embodiment of the present invention, a frequency modulated continuous wave (FMCW) ultrasonic signal is used for gait recognition. The FMCW signal is composed of multiple repeated chirp signals. The chirp signal is a signal whose frequency changes with time. Its advantages are high time-frequency resolution and strong noise resistance. These characteristics make the FMCW signal particularly suitable for gait recognition in complex environments.

[0053] It should be noted that FMCW signals mainly have two modulation forms: sawtooth wave and triangle wave. The sawtooth wave is suitable for situations where the Doppler frequency shift caused by the relative speed of the detection target can be ignored, and the sawtooth wave can achieve the maximum detection distance. The triangle wave is suitable for scenarios where both distance and speed measurement are required, and can better demodulate the Doppler frequency of the target reflection signal. However, the implementation of the sawtooth wave is relatively simple, but there may be ambiguity in speed and distance measurement. The reason is that the same frequency difference may correspond to multiple different distances and speeds, which will affect the accuracy of the measurement results. The advantage of the triangle wave is that it can eliminate the ambiguity in the measurement of speed and distance. Specifically, each frequency difference corresponds to only one unique distance and speed. In addition, the frequency change of the triangle wave is smoother, which helps to improve the stability and accuracy of the signal. At the same time, the triangle wave performs well in anti-multipath interference and can more effectively process reflected signals in complex environments.

[0054] Furthermore, the ranging principle of FMCW signals is based on the frequency difference between the transmitted signal and the echo signal. When the signal is transmitted, due to the distance difference between the target and the transmitting device, the echo signal will have a certain delay relative to the transmitted signal, resulting in a frequency difference between the two. This frequency difference is proportional to the distance of the target. In addition, if there is relative motion between the target and the transmitting device, Doppler frequency shift will be introduced, further affecting the frequency difference. By comprehensively analyzing the frequency difference between the transmitted signal and the echo signal, the distance and speed of the target can be calculated at the same time.

[0055] Preferably, a triangle wave FMCW signal of 18-22kHz is used, thereby effectively eliminating the measurement ambiguity of speed and distance, and having good anti-interference ability in complex environments. By using the triangle wave FMCW signal, high-precision gait recognition in complex environments is achieved.

[0056] In one embodiment of the present invention, the data collection method includes: using a smart phone to transmit an 18-22kHz FMCW ultrasonic signal and receiving the signal reflected by the human body when walking. This process is carried out in multiple scenes, including conference rooms, bedrooms, etc., to collect gait data of different users. In the data collection process, various factors such as different clothing, shoes, gait speed, and handheld mobile phones are considered to ensure the diversity and representativeness of the data.

[0057] The reason why ultrasound can effectively identify multiple signals by emitting ultrasound for gait data collection is that ultrasound itself has high sensitivity and penetration ability, and can capture subtle changes in gait. When participants walk in different types of clothing (such as formal and casual clothing) and various shoes (such as sneakers, leather shoes and slippers), ultrasound can keenly record the differences in reflected signals caused by these changes. These different wearing conditions may affect gait patterns, and the high-frequency signals of ultrasound can effectively reflect these differences during the collection process.

[0058] In addition, participants walked at different speeds in a set area, including fast and slow walking, and ultrasound was able to capture gait characteristics at different speeds by analyzing the frequency changes and delays of the echo signal. This flexibility enables the gait recognition system to adapt to a variety of walking conditions, thereby enhancing the system's application range.

[0059] Furthermore, when walking, the different ways in which participants hold their phones may also affect gait characteristics. By analyzing the changes in these signals using ultrasound, we can study how the relative position between the phone and the human body affects signal reception. In summary, the efficiency and adaptability of ultrasound technology enable it to fully capture and identify various gait signals in complex environments and changing conditions, thereby providing a reliable data basis for gait analysis.

[0060] Step S12: pre-training a gait recognition model based on the multiple historical gait signals, and extracting a first feature extraction network from the trained gait recognition model.

[0061] In one embodiment of the present invention, in this embodiment, the gait recognition model adopts a convolutional neural network (CNN). The network extracts features of gait signals through multiple convolution layers, each of which contains multiple convolution kernels of different sizes, and slides the convolution kernels in the time and frequency dimensions to extract the spatiotemporal features in the gait signal. The convolution operation can capture key gait features such as stride length and step frequency, and identify local detail features and global motion features at different scales.

[0062] Furthermore, the maximum pooling layer in the CNN network is used to perform dimensionality reduction sampling on the feature map output by the convolution layer. By selecting the maximum value in the feature map for downsampling, the data dimension is reduced while retaining the main feature information, reducing the computational complexity. This feature dimensionality reduction operation helps the network obtain more representative feature expressions. In the model training stage, the parameters are optimized by combining forward propagation and back propagation.

[0063] In the forward propagation process of the CNN network, the input gait data is calculated in turn through the convolution layer, pooling layer, batch normalization layer, activation layer and fully connected layer, and finally the classification probability is output through the Softmax layer. The introduction of the activation layer introduces nonlinearity to the network and increases the expressiveness of the model, while the batch normalization layer helps to accelerate the training process and improve the stability of the model. In the back propagation process, the error between the predicted result and the true label is calculated, and the weight parameters in the network are updated using the gradient descent algorithm to minimize the loss function.

[0064] When the trained model is used for gait recognition, the input gait data is extracted through the trained convolution kernel, processed by the activation layer and batch normalization layer, and then reduced in dimension through the pooling layer, and finally further integrated through the fully connected layer. Low-level features (such as the basic morphological features of the signal) are gradually transformed into high-level features (such as the complete gait pattern features), and finally the recognition results are output through the Softmax layer to complete the preliminary classification of the gait data samples.

[0065] Step S13: generating multiple groups of positive and negative sample pairs based on the initial gait signal and the multiple historical gait signals;

[0066] In one embodiment of the present invention, the multiple groups of positive and negative sample pairs include a first type of positive and negative sample pairs and a second type of positive and negative sample pairs; the first type of positive and negative sample pairs are composed of one or more gait signal combinations in historical gait signals; the second type of positive and negative sample pairs are composed of an initial gait signal and any combination of historical gait signals, and sample pairs are composed of an initial gait signal and a combination of an initial gait signal.

[0067] In this embodiment, the historical gait signal set is assumed to be H = {h1, h2, ..., h n}, where h i represents the i-th historical gait signal. Assume that the initial gait signal is s. Then the composition of the first type of positive and negative sample pairs can be expressed as, the positive sample pair is (h i ,h i ), indicating that the same historical gait signal is paired with itself; the negative sample pair is (h i ,h j ), where i≠j, represents the pairing between different types of historical gait signals.

[0068] Furthermore, the second type of positive and negative sample pairs can be expressed as follows: the positive sample pair is (s, s), which represents the pairing of the initial gait signal and itself; the negative sample pair is (s, h i ), indicating that the initial gait signal is paired with any historical gait signal. The positive and negative sample construction method in this embodiment ensures that the sample pairs contain the similarities between signals of the same type (positive sample pairs) and the differences between signals of different types (negative sample pairs).

[0069] Step S14: inputting the multiple groups of positive and negative sample pairs into the twin neural network including the first feature extraction network for model training to generate an updated twin neural network.

[0070] In one embodiment of the present invention, the process of inputting multiple groups of positive and negative sample pairs into a twin neural network including a first feature extraction network for model training includes: inputting the first category of positive and negative sample pairs and the second category of positive and negative sample pairs into the twin neural network including the first feature extraction network, and performing an optimization operation on the first feature extraction network in the twin neural network based on the contrast loss function to update it to a twin neural network including a second feature extraction network.

[0071] In this embodiment, multiple groups of positive and negative sample pairs are input into a twin neural network including a first feature extraction network to generate multiple gait similarities, and the feature extraction network in the twin network is reversely trained according to the gait similarities to update the first feature extraction network to the second feature extraction network. Among them, the update of the second feature extraction network realizes the separation of samples of different categories and the proximity of samples of the same category through the twin neural network. The purpose of reusing the first feature extraction network is to reduce the amount of training samples required to add new users and perform metric learning through the twin neural network (SiameseNetwork).

[0072] Specifically, positive sample pairs (consisting of two samples of the same category) and negative sample pairs (consisting of samples of different categories) are generated for each category. By constructing these positive and negative sample pairs, the twin neural network can learn the ability to recognize features of different categories with a limited number of samples. In addition, the feature extraction network (first feature extraction network) of the existing CNN model is used to create two feature extraction networks with shared weights, thereby extracting feature representations of input samples without retraining.

[0073] The shared weight structure ensures that the two input samples get consistent feature representations through the same feature extraction process, which can improve the accuracy and efficiency of similarity calculation. Next, compare the feature differences of the two input samples, map them to a shared feature space, and calculate the Euclidean distance between the samples to determine whether the two samples are from the same category. The smaller the distance, the more similar the samples are; the larger the distance, the more significant the difference between the samples.

[0074] In this embodiment, the twin neural network is trained using a contrast loss function. The loss function is designed to effectively learn the features that distinguish different categories by minimizing the distance between pairs of samples of the same category and maximizing the distance between pairs of samples of different categories. The mathematical expression of the loss function is shown in Formula 1. Where D represents the Euclidean distance between two samples, y represents whether the samples belong to the same category, and margin represents the preset boundary distance. By minimizing this loss, the model can better learn the difference between different categories.

[0075]

[0076] Step S15: Generate a gait recognition result corresponding to the initial gait signal through the updated twin neural network.

[0077] In one embodiment of the present invention, the process of generating a gait recognition result corresponding to the initial gait signal through the updated twin neural network includes: replacing the first feature extraction network in the trained gait recognition model with the second feature extraction network to generate an optimized gait recognition model; inputting the initial gait signal into the optimized gait recognition model to generate an optimized classification probability; if the highest probability value in the optimized classification probability is lower than the first threshold, calculating the centroid similarity of each category; if the centroid similarities of all categories are lower than the second threshold, marking the gait category of the initial gait signal as an outsider.

[0078] It should be noted that foreigner identification, as a key issue in the gait recognition task, has been further optimized. Traditional classification models can usually only accurately classify predefined known categories, while foreigner identification requires the model to have two capabilities at the same time: one is to correctly classify users of known categories, and the other is to effectively identify unseen foreigners. The main challenge of foreigner identification is that the model not only needs to have excellent known category classification capabilities, but also must have the generalization ability to distinguish known classes from unknown classes.

[0079] In this embodiment, the first threshold is used to determine whether the highest value of the optimized classification probability is reliable enough to confirm that the gait signal belongs to a specific category. If the highest optimized classification probability is lower than the threshold, it means that the model has insufficient confidence in the classification of the signal, which means that the gait signal may not meet the characteristics of any known category. In this case, the similarity between the center of mass of the signal and each category will be further analyzed to determine its classification direction. The second threshold is used to evaluate the similarity between the current gait signal and the center of mass of all known categories when the first threshold is judged to be insufficiently confident. If the center of mass similarity of all categories is found to be lower than this threshold when calculating the center of mass similarity, it means that the characteristics of the signal are significantly different from any known category. At this time, it can be reasonably judged that the gait signal does not meet the characteristics of the internal known categories and is marked as an outsider. The setting and application of the double-layer threshold helps to enhance the robustness of the gait signal classification, thereby improving the accuracy and safety of gait recognition. The selected threshold can be adjusted according to the actual application scenario and the characteristics of the training data to achieve the best effect.

[0080] In this embodiment, the centroid refers to the center point of a group of samples in the feature space, which is usually obtained by calculating the average value of the sample features. The centroid represents the typical features of each category and helps to determine the relationship between new samples and known categories.

[0081] Therefore, in the above embodiment, first, in the training stage of the gait recognition model, a data placeholder generation method is used to simulate potential unknown classes, so that the model initially has the data basis for distinguishing known classes and outsider samples. Furthermore, in this embodiment, the twin neural network further enhances the model's ability to distinguish, specifically through the learning of positive and negative sample pairs. During the training process, by comparing samples of the same category (positive sample pairs), the similarities between them are learned, and by comparing samples of different categories (negative sample pairs), the differences between them are learned. With the help of the contrast loss function, the feature distance between samples of the same category is gradually minimized, and the feature distance between samples of different categories is maximized. In addition, by analyzing the distribution of similarities between samples of known categories, a threshold for distinguishing known classes from outsiders is set, and ultimately the model's ability to recognize outsiders in different scenarios is improved.

[0082] In one embodiment of the present invention, an existing data set is input into the twin network in the form of positive and negative sample pairs. By calculating the similarity between the input samples, metric learning techniques (such as cosine similarity or Euclidean distance) are applied to determine the correlation between samples, thereby identifying and extracting important sample relationship features. Subsequently, a variety of loss functions are used to minimize the distance between similar samples and maximize the distance between heterogeneous samples, including but not limited to contrast loss and triple loss, to effectively guide the learning process of the twin network.

[0083] Furthermore, through the back-propagation algorithm, the weights of the twin network are updated according to the calculated loss value, and the network parameters are adjusted using optimization algorithms such as Adam or SGD to improve its feature extraction ability. In this process, multiple rounds of iterative training are performed, new data is continuously input, and the parameters of the twin network are updated in real time to gradually improve its ability to capture data features. After optimization, the twin network generates a more accurate and robust feature representation. Finally, a shared feature extraction network is extracted from the updated twin network and used as the second feature extraction network. Through the above steps, the optimized twin network can more effectively capture and utilize the relational features between data, significantly enhancing the performance of the model in subsequent tasks.

[0084] Figure 3 A flow chart of another embodiment of the present invention is shown. First, data preprocessing and data placeholder and generation operations are performed on the acquired initial gait data signal and historical gait data in sequence, and then the historical gait data is input into the convolution layer, batch normalization layer, activation layer and pooling layer in sequence, and finally the feature image with a depth of 7 is input into the fully connected layer, and the CNN network is updated by back propagation according to the final classification result and the standard classification result of the historical gait data, so as to generate a first feature extraction layer including the convolution layer, batch normalization layer, activation layer and pooling layer.

[0085] Furthermore, the historical gait data is input into the twin network, and the shared network in the twin network is constructed with the weight coefficient of the first feature extraction layer to construct the feature space of the sample data in the twin network, and the data in the feature extraction layer is updated according to the calculated distance and the similarity of the actual sample, thereby achieving the goal of shortening the feature space distance of similar samples and pushing the feature space distance of different samples farther. Subsequently, the updated feature extraction layer is named the second feature extraction layer, and based on the second feature extraction layer.

[0086] In one embodiment of the present invention, Figure 3 As shown, after the second feature extraction layer, the second gait recognition model is connected using a fully connected layer to generate a probability distribution for each category. Next, after normalization, the category with the highest probability is determined as the final classification result. In the current example, it is assumed that the collected real-time gait data comes from an internal person. Since its gait characteristics have a certain consistency with the known categories, it is suitable for the information entry process of the internal person.

[0087] In another embodiment of the present invention, an anomaly detection network is constructed based on the second feature extraction layer. Specifically, the second feature extraction layer is replaced by the first feature extraction layer in the original CNN network, and a threshold judgment is performed on the classification results output by the model. If the highest classification probability among all categories is lower than 0.2, it means that all classification probabilities are lower than the preset threshold. This situation usually occurs in outsiders because the gait characteristics of outsiders may be significantly different from the categories known to the system. Therefore, at this time, the current gait data is marked as an outsider, which is suitable for the scenario of anomaly detection. Such a design can help improve safety and identify and handle potential abnormal behaviors in a timely manner.

[0088] In the embodiments of the present application, words such as "first" and "second" are used to distinguish the same or similar items with basically the same functions and effects. For example, the first feature extraction network and the second feature extraction network are only used to distinguish different feature extraction networks, and their order is not limited. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.

[0089] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" represent examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0090] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.

[0091] Figure 4 is a schematic block diagram of a gait recognition device 400 based on ultrasonic sensing provided in an embodiment of the present application. Figure 4 As shown, the device includes a data acquisition module 401, a pre-training module 402 and a gait recognition module 403.

[0092] Data acquisition module 401: used to acquire the initial gait signal to be detected in real time by emitting ultrasonic waves and collecting their reflected signals; and to acquire multiple historical gait signals at the same time;

[0093] Pre-training module 402: used to pre-train a gait recognition model based on the multiple historical gait signals, and extract a first feature extraction network from the trained gait recognition model;

[0094] Gait recognition module 403: used to generate multiple groups of positive and negative sample pairs based on the initial gait signal and multiple historical gait signals; input the multiple groups of positive and negative sample pairs into the twin neural network including the first feature extraction network for model training to generate an updated twin neural network; and generate a gait recognition result corresponding to the initial gait signal through the updated twin neural network.

[0095] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.

[0096] It should also be understood that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application may be integrated into a processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0097] Figure 5is a schematic block diagram of an electronic terminal provided in an embodiment of the present application. Figure 5 As shown, the electronic terminal includes: at least one processor 501, a memory 502, at least one network interface 503 and a user interface 505. The various components in the device are coupled together through a bus system 504. It can be understood that the bus system 504 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 In the specification, various buses are labeled as bus systems.

[0098] The user interface 505 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.

[0099] It is understood that the memory 502 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include but is not limited to these and any other suitable categories of memory.

[0100] The memory 502 in the embodiment of the present invention is used to store various categories of data to support the operation of the electronic terminal 500. Examples of these data include: any executable program for operating on the electronic terminal 500, such as an operating system 5021 and an application 5022; the operating system 5021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 5022 may include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The gait recognition method based on ultrasonic sensing provided in the embodiment of the present invention may be included in the application 5022.

[0101] The method disclosed in the above embodiment of the present invention can be applied to the processor 501, or implemented by the processor 501. The processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 501 or the instruction in the form of software. The above processor 501 may be a general processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 501 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general processor 501 may be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0102] In an exemplary embodiment, the electronic terminal 500 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD) to execute the aforementioned method.

[0103] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: a computer program code, when the computer program code is run on a computer, the computer executes the gait recognition method based on ultrasonic sensing of any one of the embodiments shown above.

[0104] According to the method provided in the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores a program code. When the program code runs on a computer, the computer executes the gait recognition method based on ultrasonic perception of any one of the embodiments shown above.

[0105] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program and / or a computer. By way of illustration, both applications running on a computing device and a computing device can be components. One or more components may reside in a process and / or an execution thread, and a component may be located on a computer and / or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may, for example, communicate through local and / or remote processes according to signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system and / or a network, such as the Internet interacting with other systems through signals).

[0106] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0107] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0108] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0109] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0111] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When loading and executing computer program instructions (programs) on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that contains one or more available media integrated. Available media may be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid state disks (SSDs)).

[0112] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0113] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0114] In summary, the present application provides a gait recognition method, device, medium, program product and terminal based on ultrasonic sensing. The present invention provides a method for improving the efficiency of gait recognition, which collects gait information through the microphone and speaker of ordinary commercial audio equipment, and uses a convolutional neural network for feature extraction and recognition, and then reuses the pre-trained feature extraction layer through the twin neural network and introduces data placeholder generation technology to achieve user expansion and recognition under small sample conditions. It solves the problem that the existing gait recognition technology mainly relies on visual information, is easily affected by lighting and background changes, and has great limitations in the recognition of foreign individuals. It maintains high accuracy and stability in complex dynamic environments, effectively improves the ability to expand new users and identify outsiders, and at the same time reduces the implementation threshold and optimizes the overall recognition efficiency through the use of ordinary audio equipment. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.

[0115] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.

Claims

1. A gait recognition method based on ultrasonic sensing, characterized in that: include: By emitting ultrasonic waves and collecting their reflected signals, the initial gait signal to be detected is obtained in real time; and multiple historical gait signals are obtained at the same time; Pre-training a gait recognition model based on the multiple historical gait signals, and extracting a first feature extraction network from the trained gait recognition model; Based on the initial gait signal and the multiple historical gait signals, generating multiple groups of positive and negative sample pairs; Input multiple groups of positive and negative sample pairs into the twin neural network including the first feature extraction network for model training to generate an updated twin neural network.

2. The gait recognition method based on ultrasonic sensing according to claim 1, characterized in that: The multiple groups of positive and negative sample pairs include a first type of positive and negative sample pairs and a second type of positive and negative sample pairs; The first type of positive and negative sample pairs is composed of one or more gait signal combinations in historical gait signals; The second type of positive and negative sample pairs is composed of an initial gait signal and any combination of historical gait signals, and a sample pair is composed of an initial gait signal and a combination of an initial gait signal.

3. The gait recognition method based on ultrasonic sensing according to claim 2 is characterized in that: The process of inputting the plurality of groups of positive and negative sample pairs into the twin neural network including the first feature extraction network for model training includes: The first category of positive and negative sample pairs and the second category of positive and negative sample pairs are input into a twin neural network including a first feature extraction network, and an optimization operation is performed on the first feature extraction network in the twin neural network based on a contrast loss function to update it to a twin neural network including a second feature extraction network.

4. The gait recognition method based on ultrasonic sensing according to claim 3 is characterized in that: The process of generating a gait recognition result corresponding to the initial gait signal through the updated twin neural network includes: Replacing the first feature extraction network in the trained gait recognition model with the second feature extraction network to generate an optimized gait recognition model; Inputting the initial gait signal into the optimized gait recognition model to generate an optimized classification probability; If the highest probability value in the optimized classification probability is lower than the first threshold, the centroid similarity of each category is calculated; if the centroid similarities of all categories are lower than the second threshold, the gait category of the initial gait signal is marked as an outsider.

5. The gait recognition method based on ultrasonic sensing according to claim 1, characterized in that: After obtaining a plurality of historical gait signals, the following operations are further performed: a linear interpolation operation is performed on the plurality of historical gait signals, and randomly generated placeholder data is inserted into the interpolation results to simulate the characteristics of potential outsiders.

6. The gait recognition method based on ultrasonic sensing according to claim 1, characterized in that: By emitting ultrasonic waves and collecting their reflected signals, the initial gait signal to be detected is obtained in real time; after obtaining multiple historical gait signals, the following operations are also performed: one or more of mixing operations, filtering operations, frequency domain conversion operations and outlier removal operations are performed on the initial gait signal and multiple historical gait signals in sequence.

7. A gait recognition device based on ultrasonic sensing, characterized in that: include: Data acquisition module: used to acquire the initial gait signal to be detected in real time by emitting ultrasonic waves and collecting their reflected signals; and to acquire multiple historical gait signals at the same time; Pre-training module: used for pre-training a gait recognition model based on a plurality of the historical gait signals, and extracting a first feature extraction network from the trained gait recognition model; Gait recognition module: used to generate multiple groups of positive and negative sample pairs based on the initial gait signal and multiple historical gait signals; input the multiple groups of positive and negative sample pairs into the twin neural network including the first feature extraction network for model training to generate an updated twin neural network; generate a gait recognition result corresponding to the initial gait signal through the updated twin neural network.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the gait recognition method based on ultrasonic sensing described in any one of claims 1 to 6 is implemented.

9. A computer program product, characterized in that The computer program product includes computer program codes, and when the computer program codes are executed on a computer, the computer is enabled to implement the gait recognition method based on ultrasonic sensing as claimed in any one of claims 1 to 6.

10. An electronic terminal comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the gait recognition method based on ultrasonic sensing as described in any one of claims 1 to 6.