Device and method for gesture recognition and / or person identification based on acoustomyographic signals

A bracelet-based device with acoustic sensors and signal processing techniques addresses the limitations of existing gesture recognition technologies, enabling reliable and accurate gesture and person recognition through enhanced acoustomyographic signal capture and analysis.

FR3164811A1Pending Publication Date: 2026-01-23CENT NAT DE LA RECH SCI (C N R S) +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2024007760
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing gesture recognition technologies using electromyographic (EMG), sonomyographic (SMG), and acoustomyographic (AMG) signals face issues such as electrode detachment, high cost, complexity, noise sensitivity, and poor signal quality, limiting accurate and robust gesture and person recognition.

Method used

A device comprising a bracelet with cavities and acoustic sensors positioned around a limb to capture acoustomyographic signals, combined with signal processing methods like Empirical Mode Decomposition and Mel Frequency Cepstral Coefficients, to enhance signal-to-noise ratio and enable reliable gesture and person recognition.

Benefits of technology

The device and method provide robust and accurate recognition of gestures and individuals by leveraging acoustomyographic signals, offering improved signal quality and unique muscle sound signatures for authentication and control applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A gesture recognition device is proposed, comprising: a bracelet configured to be positioned around a person's limb; a plurality of cavities formed in the bracelet, each cavity having an entrance intended to be in contact with the person's skin and delimited by a base opposite the entrance; and a plurality of acoustic sensors, each acoustic sensor being housed within or adjacent to the base of one of the cavities to receive acoustomyographic signals produced by the muscles of the user's limb. Abstract figure: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Device and method for gesture recognition and / or person identification from acoustomyographic signals. Technical field

[0001] This disclosure relates to devices and methods for detecting and interpreting a person's gestures and movements. More specifically, this disclosure relates to a device and method for recognizing a person's gestures and / or identifying a person from acoustomyographic signals. Previous technique

[0002] Gesture recognition and interpretation is typically used in the medical field, for example, to control a prosthesis, to monitor muscle fatigue and aging, or to estimate a person's strength. Such gesture recognition and interpretation also finds use in virtual reality, for controlling objects or within a game. Furthermore, personal identification can be used, as biometric data, for authentication and to secure access in the security sector.

[0003] It is known to detect a person's gestures and movements using electromyographic (EMG) signals. Electromyographic signals are electrical signals generated in the person's muscle fibers when the muscles contract. To detect these signals, surface electrodes are attached to the skin over the targeted muscle to provide a signal emitted by a set of muscle fibers. However, acquiring an EMG signal requires precise placement of the electrodes on the person. During movement, the electrodes become detached and do not allow for obtaining a satisfactory signal. Furthermore, the EMG signal is affected by changes in skin impedance.

[0004] It is also known to perform gesture detection from sonomyographic signals (SMG). Such signals use ultrasound to detect muscle movements and changes in shape. To detect these signals, an ultrasound sensor is placed on the skin over the muscles. However, acquiring SMG signals is expensive, the sensors are not very portable, and analyzing the acquired signals is complex, requiring advanced algorithms.

[0005] Finally, it has been proposed to use gesture detection based on acoustomyographic (AMG) signals, based on the acoustic vibrations generated by muscles when they contract. However, the proposed devices and methods do not They are not accurate. Acoustic sensors are sensitive to noise, which prevents reliable and robust gesture recognition. Summary

[0006] A gesture and / or person recognition device is proposed comprising: - a bracelet configured to be positioned around a person's limb; - a plurality of cavities formed in the bracelet, each cavity having an entrance intended to be in contact with the skin of said limb of the person and being delimited by a base, opposite the entrance; and - a plurality of acoustic sensors, each acoustic sensor being housed within or adjacent to the bottom of a cavity in the plurality of cavities to receive acoustomyographic signals produced by the muscles of the person's limb.

[0007] Such a device makes it possible to capture the sounds emitted by muscles during their contraction (acoustomyographic signal). In particular, it is possible to detect muscle sounds with considerable gain and a very good signal-to-noise ratio. The acoustomyographic signals acquired with this device make it possible to recognize a particular gesture of the person, or even to distinguish between different people performing the same gesture. Indeed, for the same movement, the sound emitted by a person's muscles constitutes a signature specific to that person.

[0008] Optionally, each cavity may be spherical, cylindrical, conical, or Gaussian in shape. The cavity shape is chosen to improve gain and increase the signal-to-noise ratio of the acoustomyographic signal.

[0009] Optionally, each cavity can be spherical in shape and have a radius-to-height ratio of approximately 1, or each cavity can be conical in shape and have a radius-to-height ratio of approximately 2. Thus, the gain and signal-to-noise ratio are optimized by playing on the volume variations in the cavity.

[0010] Optionally, each cavity may have a radius between 6 mm and 9 mm and a height between 6 mm and 9 mm. The cavity radius is chosen to increase the surface area of ​​the cavity in contact with the person's skin, improving the perception of the acoustomyographic signal by the acoustic sensor housed within or adjacent to the bottom of the cavity. The cavity height is chosen to improve the amplification of the acoustomyographic signal within the cavity.

[0011] Optionally, each acoustic sensor can be sensitive to audible sound waves, i.e. waves having a frequency below 20 kHz.

[0012] Optionally, each acoustic sensor can be inserted into a housing extending along a central axis of the corresponding cavity, said housing opening into a narrow opening formed in the bottom of the cavity. The sensor thus positioned in the housing is located at the bottom of the corresponding cavity and is adapted to acquire the signal from the person's muscles, amplified by the cavity.

[0013] Optionally, the device may include between 2 and 8, preferably between 4 and 6, acoustic sensors. Such a number of acoustic sensors makes it possible to cover all the extensor and flexor muscles of the person's limb and improves the robustness of gesture and / or person recognition.

[0014] Optionally, the device may further include a plurality of amplification and filtering boards, and each acoustic sensor is connected to one of the plurality of amplification and filtering boards. Thus, the acoustomyographic signal is conditioned following its acquisition by the associated acoustic sensor.

[0015] Optionally, the device may further comprise a plurality of rigid elements, each rigid element comprising one of the plurality of cavities, and an elastic element connecting the plurality of rigid elements. The elastic element allows the bracelet to be tightened against the user's skin, so that the entrance of the cavity formed in the rigid element extends over the user's limb and the sound emitted by the muscles is captured by the acoustic sensor housed within or adjacent to the bottom of the cavity.

[0016] According to another aspect, a method for recognizing gestures and / or people is proposed comprising: - an acquisition of an acoustomyographic signal from the muscles of a person performing a gesture; - an analysis of the acquired acoustomyographic signal to extract characteristics of the acquired acoustomyographic signal; - an input of the extracted features into a classifier, said classifier comprising features associated with a particular gesture or a particular person, the features of the classifier having been recorded beforehand, the input of the extracted features into the classifier allowing the extracted features to be associated with the particular gesture performed or the particular person performing the gesture.

[0017] Such a method reliably and robustly recognizes a person's gestures. In particular, it is possible to distinguish between different people performing the same gesture. Indeed, for the same movement, the sound emitted by a person's muscles constitutes a unique signature for that person. Thus, the method makes it possible to recognize a specific individual, notably for authentication or access authorization. It is also possible to recognize a specific gesture for the purpose of controlling a prosthesis or an object in virtual reality.

[0018] The acquisition of the acoustomyographic signal from the person's muscles can be performed by means of acquisition in conjunction with one or more acoustic sensors. For example, each acoustic sensor can be integrated into equipment that includes these acquisition means and forms an interface with a limb / body part of the person. The method can be implemented by a fixed apparatus / installation or a mobile device. The method can use a wrist device, as described above, or an installation with one or more receiver compartments, where a limb of the person undergoing the recognition phase is received and held. For example, each sensor used during the acquisition is housed within or adjacent to a cavity, at the level of a cavity boundary.Preferably, each cavity has a central axis passing through the sensor and opening through an opening / entrance of the cavity, which entrance is delimited by an annular portion which allows contact (possibly annular) with a limb of the person.

[0019] Optionally, the acoustomyographic signal may combine several signals (sub-signals) due to the involvement of several acoustic sensors, which may, for example, have a distribution, possibly annular, around a muscular area emitting a characteristic sound mark or imprint. During the acquisition of the acoustomyographic signal, the method may include the generation of a signal vector, which is representative of this characteristic sound imprint, measured or resulting from measurements in the frequency range below 20 kHz (audible sounds).

[0020] Optionally, the signal analysis may include obtaining the frequency spectrum of the acquired acoustomyographic signal.

[0021] Optionally, obtaining the frequency spectrum may include: windowing the acquired acoustomyographic signal to obtain time frames of the acoustomyographic signal; and applying a short-time Fourier transform to each time frame to obtain the frequency spectrum of each time frame. Thus, the frequency spectrum can be obtained using a simple and fast implementation method. This method is particularly suitable for processing on relatively short windows (less than 128 samples).

[0022] Optionally, obtaining the frequency spectrum may include: decomposing the acquired signal into intrinsic oscillation modes (IMFs) using the Empirical Mode Decomposition (EMD) adaptive decomposition method; applying a nonlinear TKEO operator and a discrete energy separation algorithm to obtain an instantaneous amplitude and frequency of the acquired signal; windowing the instantaneous amplitude and frequency to obtain frames of the instantaneous amplitude and frequency; and calculating a marginal spectrum. Hilbert's method, MHS, uses the amplitude and instantaneous windowed frequency to obtain the frequency spectrum of each frame. Thus, the frequency spectrum can be obtained using a self-adaptive method. This method is particularly well-suited for processing relatively long windows (greater than 128 samples).

[0023] Optionally, other signal analysis methods may be considered, in particular other time-frequency analysis methods may be suitable, such as the wavelet transform, the Gabor transform or the Ville-Wigner transform.

[0024] Optionally, the signal analysis further includes a calculation of the MEL coefficients of the acoustomyographic signal. The MEL coefficients and their derivatives can form the features used in the classifier. Optionally, other features can be considered, for example, the fundamental frequency, formants, zero-crossing rate, and statistical features of the spectra. It is also possible to extract features using deep networks.

[0025] Optionally, the calculation of the Mel coefficients may include: passing the frequency spectrum through a bank of filters arranged according to the Mel scale to obtain a set of coefficients representing the signal power in different frequency bands, each frequency band being defined according to the Mel scale; taking the logarithm of the output of each filter in the filter bank; and applying a discrete cosine transform (DCT) to the logarithm to calculate the Mel coefficients. The calculation of the Mel coefficients makes it particularly useful to efficiently represent the relevant information of the acoustomyographic signal while reducing the amount of data to be processed.

[0026] Optionally, the filter bank may comprise between 5 and 15 filters, preferably between 7 and 13 filters, and even more preferably 7 filters. Such a number of filters makes it possible to obtain more characteristics (Mel coefficients, derivatives of the coefficients and second derivatives of the coefficients) of the signal, improving the robustness and reliability of gesture recognition.

[0027] Optionally, the filter bank is spread over a frequency range between 0 Hz and 300 Hz, preferably between 0 Hz and 100 Hz. Such a frequency range is suitable for processing acoustomyographic signals.

[0028] Optionally, the classifier may be one of the following: an artificial neural network (ANN) model; a Support Vector Machine (SVM) model; or a k-nearest neighbors (KNN) algorithm. Such models are relatively simple and can be implemented on low-cost boards. Other models accessible to those skilled in the art may also be considered. Brief description of the drawings

[0029] Other features, details and advantages will become apparent from the detailed description below and from the analysis of the accompanying drawings, in which: Fig. 1

[0030] [Fig-1] schematically illustrates a gesture recognition device and / or persons arranged on a person's forearm, according to one embodiment. Fig. 2

[0031] [Fig.2] schematically illustrates a portion of the gesture and / or person recognition device of [Fig.1] according to one embodiment. Fig. 3

[0032] [Fig.3] schematically illustrates a cross-sectional view of a detail of the portion of the device in [Fig.2]. Fig. 4A and Fig. 4B

[0033] [Fig. 4A] and [Fig. 4B] schematically illustrate a top view and a bottom view of an example of a rigid element that can be implemented in a gesture and / or person recognition device of [Fig. 1] according to one embodiment. Fig. 5A, Fig. 5B, Fig. 5C, Fig. 5D, Fig. 5E

[0034] [Fig.5A], [Fig.5B], [Fig.5C], [Fig.5D], [Fig.5E] schematically illustrate views cross-section of different examples of rigid elements that can be implemented in a gesture and / or person recognition device of the [Fig.1] according to one embodiment. Fig. 6

[0035] [Fig.6] illustrates a flowchart of a gesture and / or person recognition process according to an embodiment Fig. 7

[0036] [Fig.7] illustrates a flowchart of a first method of acquiring an acoustomyographic signal according to one embodiment. Fig. 8

[0037] [Fig. 8] illustrates a flowchart of a first method for obtaining a frequency spectrum of an acoustomyographic signal according to one embodiment. Fig. 9

[0038] [Fig. 9] illustrates a flowchart of a second method for obtaining a frequency spectrum of an acoustomyographic signal according to one embodiment. Fig. 10

[0039] [Fig. 10] illustrates a flowchart for calculating the Mel coefficient according to one embodiment. Fig. 11

[0040] [Fig. 11] illustrates a flowchart for the input of Mel coefficients into a classifier according to one embodiment. Description of the implementation methods

[0041] Figure 1 schematically illustrates an example of a device 10 for recognizing a person's gestures, and / or for recognizing a person based on the gesture performed by the person, from acoustomyographic (AMG) signals. The device 10 comprises a bracelet 12, a plurality of cavities 14 formed in the bracelet 12, and a plurality of acoustic sensors 16 housed in the cavities 14.

[0042] In the example illustrated in [Fig. 1], the bracelet 12 is positioned on the person's forearm 18. Thus, the gesture and / or person recognition device 10 is adapted to recognize hand gestures and movements. The recognition device 10 is also adapted to identify a person based on their hand gestures. For example, the device 10 is adapted to recognize when the person closes their hand, or to distinguish the person from others based on their hand closure. Alternatively, the bracelet 12 could be positioned around any other limb, for example, the person's leg.

[0043] The bracelet 12 comprises a plurality of rigid elements 20 connected to each other by an elastic element 22. The rigid elements 20 are regularly spaced around the entire circumference of the bracelet 12. Each rigid element 20 may, for example, be made of a plastic material, such as PVC, or of a metallic material. The elastic element 22 may be one or more elastic bands. The elastic element 22 allows the bracelet 12 to be tightened against the user's skin. Each rigid element 20 may include one or more openings 24 into which each elastic band is inserted to secure the elastic band to the rigid element 20.

[0044] Each rigid element 20 comprises one cavity from among the plurality of cavities 14. The cavity is delimited by the walls of the rigid element. The cavity extends from a first face 26 of the rigid element 20 intended to be in contact with the person's skin. The cavity 14 comprises an entrance 28, intended to be in contact with the person's skin. The entrance 28 of the cavity is then delimited by a portion of the rigid element in contact with the person's limb. The cavity is delimited by a bottom 30, opposite the entrance 28. The bottom is defined by the wall of the rigid element. An acoustic sensor from among the plurality of acoustic sensors 16 is housed within or adjacent to the bottom 30 of the cavity. The cavity amplifies the sounds produced by the muscles during a movement of the person, which can be captured by the acoustic sensor housed in or near the bottom 30 of the cavity.

[0045] Figures 5A to 5E illustrate different cavity shapes 14. The cavity may, for example, have a spherical, cylindrical, conical, or Gaussian shape. The shape The cavity can be a hemisphere or a quarter sphere. Its shape is chosen to reduce the signal-to-noise ratio of the signal perceived by the acoustic sensor and to maximize signal gain. Preferably, the cavity is spherical or conical.

[0046] A cavity height h is, for example, between 6 mm and 9 mm. The height h is chosen to improve the amplification of the acoustomyographic signal within the cavity. A cavity radius r is, for example, between 6 mm and 9 mm. The radius r is chosen to increase the surface area of ​​the cavity in contact with the person's skin, improving the perception of the acoustomyographic signal by the acoustic sensor located towards the back of the cavity. Note that the radius r is measured in the plane of the cavity entrance 28, and the height h is measured between the entrance 28 and the back 30 of the cavity, perpendicular to the plane of the cavity entrance 28.

[0047] When the cavity is spherical, a ratio between the radius r of the cavity and the height h of the cavity, r / h, is about 1. When the cavity is conical, a ratio between a radius of the cavity and a height of the cavity, r / h, is about 2. These r / h ratios particularly improve the gain and increase the signal-to-noise ratio of the acoustomyographic signal.

[0048] The acoustic sensors 16 can be any type of sensor capable of detecting acoustic vibrations in the environment. Each acoustic sensor can, for example, be sensitive to waves with a frequency below 20 kHz. The device 10 comprises, for example, between 2 and 8 acoustic sensors housed in the cavities 14, preferably between 4 and 6 acoustic sensors. The acoustic sensors 16 are regularly spaced around the person's limb when the bracelet 12 is positioned around the limb. The plurality of acoustic sensors 16 makes it possible to cover all the flexor and extensor muscles of the person and thus improves the robustness of gesture and / or person recognition.

[0049] Each acoustic sensor is housed within or adjacent to the bottom 30 of a cavity of the plurality of cavities 14, for example by being inserted into a housing 32 extending from a second face 34 of the rigid element 20, opposite the first face 26, to the bottom 30 of the cavity. The housing 32 may be cylindrical, a central axis of each cavity coinciding with a central axis of the housing 32. The bottom 30 of the cavity may include a narrow opening communicating with the housing 32. The narrow opening is formed in the wall delimiting the bottom of the cavity, to communicate with the housing 32. The sounds emitted by the muscles can be amplified in the cavity and captured by the acoustic sensor.

[0050] The device 10 may further comprise a plurality of amplification and filtering boards 36. For example, an amplification and filtering board may be associated with each acoustic sensor 16. As shown in [Fig. 2], the amplification board and filtering can be mounted on each rigid element 20. For example, each rigid element 20 may include mounting holes 38, for example four mounting holes 38, and each amplification and filtering board may include rods 40, in particular four rods 42, intended to be inserted into the mounting holes 38. The acoustic sensor in the cavity of the rigid element can then be electronically connected to the amplification and filtering board.

[0051] Each amplification and filtering board 36 may include at least one bandpass filter and at least one amplifier. For example, each amplification and filtering board 36 may include two amplification and filtering stages, or even three amplification and filtering stages. Each amplification and filtering board 36 may further include an analog-to-digital converter.

[0052] The device 10 further includes an acquisition module configured to combine several signals from different sensors 16. For example, these signals can each be preprocessed by the amplification and filtering board 36 associated with the sensor that detected the corresponding signal. In non-limiting embodiments, the acquisition module is carried by the bracelet 12.

[0053] The device may also include a communication module configured to transmit the acquired acoustomyographic signal, notably by the acquisition module. The communication module communicates wirelessly with an analysis processing system. For example, the wireless communication is of the Bluetooth or Wi-Fi type. Note that the processing and analysis system could also be integrated into the wristband 12, or be wired to each acoustic sensor 16.

[0054] A method for processing and analyzing electromyographic signals to recognize a person's gestures is subsequently described ([Fig. 6]). The device above can, for example, be designed to implement the method described below.

[0055] The method includes an acquisition E100 of an acoustomyographic signal from the muscles of the person performing a gesture, an analysis E200 of the acquired acoustomyographic signal to extract features of the acoustomyographic signal and an input E300 of the extracted features into a classifier.

[0056] The acquisition E100 of the acoustomyographic signal can be carried out using the device 10 described above. The signal acquisition includes the emission E110 of sounds by the muscles of the person performing a movement. The acquisition E100 also includes the detection E120 of sounds by each acoustic sensor 16 to obtain an electrical signal from the acoustic sensor. The acquisition E100 further includes amplification and filtering E130 of each electrical signal, by for example, by the amplification and filtering board 36 associated with the acoustic sensor 16. The acquisition E100 further includes an analog-to-digital conversion E140 of each amplified and filtered signal to obtain a digital electrical signal. The analog-to-digital conversion can, for example, use a sampling frequency Fs of 1024s1. The acquisition E100 of the acoustomyographic signal further includes obtaining E150 a signal vector x[n], each component of the signal vector x[n] being a digital electrical signal from one of the acoustic sensors. The signal vector x[n] has a number of components corresponding to the number Nc of acoustic sensors 16. For example, the device 10 can include between 2 and 8 acoustic sensors 16, and the digital electrical signal vector x[n] can thus include between 2 and 8 digital electrical signals:

[0057] / x^] \ x[n] = I ï ] I \ xJn] /

[0058] The E200 processing of the acoustomyographic signal may include obtaining the frequency spectrum of the acquired acoustomyographic signal. Obtaining the frequency spectrum of the acoustomyographic signal may be performed by the processing and analysis system. In particular, the signal vector x[n] may be transmitted from the device 10 to the processing and analysis system, for example, by wireless communication. Obtaining the frequency spectrum may, for example, be performed by a first method ([Fig. 8]) or a second method ([Fig. 9]). The first method makes it possible to recognize the person's gestures over short temporary windows, in particular windows containing fewer than 128 samples, preferably 32 samples. The second method makes it possible, in particular, to recognize gestures over longer durations, for example, windows containing more than 128 samples.

[0059] The first method may include E210 preprocessing of the acoustomyographic signal obtained from the E100 signal acquisition. For example, the E210 preprocessing may include subsampling each signal in the vector. The E210 preprocessing may also include detecting the presence of muscle activity, for example, using a detection algorithm. The detection algorithm makes it possible, in particular, to identify the signal characteristics that indicate movement of the person. An example of such an algorithm is the Voice Activation Detection (VAD) algorithm.

[0060] The first method includes an E220 windowing to obtain a plurality of time frames of the acoustomyographic signal. The E220 windowing includes the weighting xp[n] of the signal vector x[n] by the p-th window. The window can be

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077] a Hann window or a Hamming window. The length N of a window p can, for example, be between 32 and 512 samples. The overlap between two successive windows p can, for example, be 50%. For example: =*[ (t -1) xp + w] xw[? / ] p^ {0, ...,P-1} andn e {0, ...,^-1} w[m] = aQ- ( 1-^0)004^] Where P corresponds to the number of windows, N is the length of a window, w is a Hamming window (a0=0.56) or a Hann window (a0=0.5). The first method involves applying an E230 short-term Fourier transform to each time frame to obtain the frequency spectrum of each time frame: Where K is the number of points for calculating the short-term Fourier transform. The second method for obtaining the frequency spectrum involves an E201 decomposition of the acquired acoustomyographic signal into intrinsic oscillation modes (IMFs) using the adaptive decomposition method "Empirical Mode Decomposition" (EMD). The signal resulting from the decomposition can be represented as the sum of the IMFs and the final residual rest[n]: i=0 +rest[n] where] = Re(Ai[n] 0i[n]=2jr(Fi[n}-Fi[nl])Ts Where A;[n] is the instantaneous amplitude, F;[n] is the instantaneous frequency and Nimfest the number of intrinsic modes of oscillation. The second method includes an E202 application of a nonlinear TKEO operator and a discrete energy separation algorithm (DESA) to obtain the instantaneous amplitude and frequency of the decomposed signal. The application of the nonlinear TKEO operator allows the signal energy to be estimated as follows: W[F.[n]] = / 72[n] -ri[n + l]rr[n-1] The energy separation algorithm (DESA) allows us to obtain the instantaneous amplitude Ai[n] and the instantaneous frequency Fi[n] of the acoustomyographic signal: Ar. ' ”J ~ 'FErfnfrJw-l]] +T[rjmT]-rJn]]

[0078] The second method includes an E203 windowing of the instantaneous amplitude and frequency to obtain frames of the instantaneous amplitude and frequency. The windowing may include the application of a Hann or Hamming window to obtain overlapping frames. The overlap between two successive windows p may, for example, be 50%. The instantaneous frequency and amplitude can be represented as follows:

[0079] A / / ?[n] = (y -1) x p+w]x)v[n]

[0080] Fip[n] (y -1) x p+w]xw[n]

[0081] oàpe {0, ...,P-Ï}andne {0, ...,^-1}

[0082] [ n ] _ cospF ]

[0083] Where P corresponds to the number of windows, N is the length of a window, w is a Hamming window (a0=0.56) or a Hann window (a0=0.5).

[0084] The second method comprises an E204 calculation of a Hilbert marginal spectrum (MHS) on the amplitude and the windowed instantaneous frequency to obtain the frequency spectrum of each frame. The MHS provides a measure of the total amplitude of each frame with a resolution corresponding to the ratio between the sampling frequency Fs and the number of points K (Fs / K). The number of points K corresponds to the number of points K used for applying the Fourier transform of the first method. 100851 wr-a^w)

[0086] [FmifP ^max] {0,1}

[0087] whereQ k (f) 0 otherwise

[0088]

[0089] The E200 processing of the acoustomyographic signal may also include a calculation of the MEL coefficients. The calculation of the Mel coefficients here includes a passage E240 of the frequency spectrum xJà:] of each window p through a bank of filters distributed according to the Mel scale, a taking of the logarithm E250 of the outputs of each filter and the application E260 of a discrete cosine transform (DCT) on the logarithm to calculate the Mel coefficients. Passing the E240 frequency spectrum through a filter bank allows to obtain a set of coefficients representing the signal power in different frequency bands. Each filter m is a bandpass filter whose frequency band is defined according to the Mel scale. Each filter m is placed non-linearly on the Mel frequency scale. The filter bank is distributed according to the MEL scale adapted to acoustomyographic signals. The number of The number of filters m in the filter bank can be between 7 and 13, preferably 7. The filter bank can be spread over a frequency range between 0 Hz and 300 Hz, preferably between 0 Hz and 100 Hz. Each filter m in the filter bank has a defined frequency response Hm around its center frequency fm. The output of each filter m is a power represented as follows:

[0091] Where Hm [k] is the frequency response of the m-th adapted Mel filter.

[0092] The E240 passage of the frequency spectrum through a filter bank may include the transition from the linear scale to the MEL scale:

[0093] [o £1 ; . F 1 $ / f \ f —> Sxlog 1 + 7-)

[0094] The E240 passage of the frequency spectrum through a filter bank may also include the passage from the MEL scale to the linear scale;

[0095] [FF 1 __ [ 0 — 1 , r mtr 1 maxj * Lv' 2 1

[0096] F^etF^v are the minimum and maximum frequencies of the acoustomyographic signals. The number of filters M > 7.

[0097] Here each filter has a center frequency f _ JSa response m M+l / The frequency is given by: 0 0 sif>f .

[0098] Taking the logarithm E250 of the output of each filter makes it possible to imitate the logarithmic perception of the filtered sound intensity: log s(p. m)

[0099] The E260 application of the DCT to the logarithm of the output of each filter is carried out as follows:

[0100] "0) x - 0.5}p)

[0101] The E260 application of the DCT to the logarithm allows obtaining a coefficient which represents the characteristics of the signal on the Mel frequency scale for each if f < f Ai / < f <f you are <f<f 5 J m JJ m+1 window p and each filter m in the filter bank. The DCT application can further include a summation of the coefficients for all windows p:

[0102] p)

[0103] Where C(m) is a characteristic of the signal for the m-th filter. The characteristic C(m) is a Mel coefficient (Mel-Frequency Cepstral Coefficients, MFCC) of the acoustomyographic signal. This yields a coefficient C(m) that is no longer time-dependent.

[0104] The feature input E300 in a classifier can include an E310 calculation of the first and second derivatives of each Mel coefficient. Calculating the first and second derivatives of the Mel coefficients allows for the use of additional features for gesture recognition. Indeed, the Mel coefficients represent the signal strength on the Mel scale, the first derivative represents the change in MFCC over time, and the second derivatives capture the variation of these changes. This improves the robustness of the classifier.

[0105] The Mel coefficients, and where applicable their derivatives and second derivatives, can be entered into the classifier, which can be an artificial neural network (ANN), a Support Vector Machine (SVM) model, or a k-nearest neighbors (KNN) algorithm. Such models are relatively simple and can be implemented at low cost.

[0106] The classifier is generated and trained beforehand. The classifier's performance can be verified by cross-validation to test its capabilities. In the embodiment where the method is for recognizing a person's gesture, the classifier can include pre-recorded features associated with a particular gesture. In the embodiment where the method is for identifying a person, the classifier can include pre-recorded features associated with the person. Indeed, for the same movement, the sound emitted by a person's muscles constitutes a signature unique to that person.

[0107] The E300 input of features in the classifier can include an association E320 of features, for example Mel coefficients, and where applicable their derivative and second derivative, to a specific person. In this case, the classifier is configured to associate a specific person with the coefficients. Thanks to the method described above, a hand-closing movement can be used to distinguish between two people. In this case, the method is particularly suitable for a security application. Any person can be identified from their acoustomyographic signals.

[0108] Alternatively, feature input may include associating features with a specific gesture. The specific gesture may be recognized for controlling a prosthesis, controlling an object, or playing a virtual reality game. < / f>

Claims

Demands

1. A gesture and / or person recognition device (10) comprising: - a bracelet (12) configured to be positioned around a limb of a person; - a plurality of cavities (14) formed in the bracelet (12), each cavity having an entrance (28) intended to be in contact with the skin of said limb of the person and being delimited by a bottom (30), opposite the entrance (28); and - a plurality of acoustic sensors (16), each acoustic sensor being housed within or adjacent to the bottom (30) of a cavity of the plurality of cavities (14) to receive acoustomyographic signals produced by the muscles of the limb of the person.

2. Device (10) according to claim 1, wherein each cavity is spherical, cylindrical, conical or Gaussian in shape.

3. Device according to claim 1 or 2, wherein each cavity is spherical in shape and has a radius-to-height ratio of about 1, or each cavity has a conical shape and has a radius-to-height ratio of about 2.

4. Device (10) according to any one of the preceding claims, wherein each cavity has a radius (r) between 6 mm and 9 mm and a height (h) between 6 mm and 9 mm.

5. Device (10) according to any one of the preceding claims, wherein the device (10) comprises between 2 and 8, preferably between 4 and 6 acoustic sensors (16).

6. Device (10) according to any one of the preceding claims, wherein the device (10) further comprises a plurality of amplification and filtering boards (36), and each acoustic sensor is connected to one amplification and filtering board among the plurality of amplification and filtering boards (36).

7. Device (10) according to any one of the preceding claims, further comprising a plurality of rigid elements (20), each rigid element (20) comprising one of the plurality of cavities (14), and an elastic element (22) connecting the plurality of rigid elements (20).

8. A method for recognizing gestures and / or people, the method being implemented by the device according to any one of claims 1 to 7, the method comprising: - acquiring an acoustomyographic signal from the muscles of a person performing a gesture; - analyzing the acquired acoustomyographic signal to extract features from the acquired acoustomyographic signal; - inputting the extracted features into a classifier, said classifier comprising features associated with a particular gesture or a particular person, the features of the classifier having been recorded beforehand, the input of the extracted features into the classifier allowing the extracted features to be associated with the particular gesture performed or the particular person performing the gesture.

9. A method according to claim 8, wherein the analysis of the acoustomyographic signal includes obtaining the frequency spectrum of the acquired acoustomyographic signal.

10. A method for recognizing gestures and / or people according to claim 9, wherein obtaining the frequency spectrum comprises: - windowing the acquired acoustomyographic signal to obtain time frames of the acoustomyographic signal; - applying a short-term Fourier transform to each time frame to obtain the frequency spectrum of each time frame.

11. A method for recognizing gestures and / or people according to claim 9, wherein obtaining the frequency spectrum comprises: - decomposing the acquired signal into intrinsic oscillation modes, IMF, using the adaptive Empirical Mode Decomposition, EMD method; - applying a nonlinear TKEO operator and a discrete energy separation algorithm to obtain an instantaneous amplitude and frequency of the acquired signal; - windowing the instantaneous amplitude and frequency to obtain frames of the instantaneous amplitude and frequency; - a calculation of a Hilbert marginal spectrum, MHS, on the amplitude and instantaneous windowed frequency to obtain the frequency spectrum of each frame.

12. A method for recognizing gestures and / or people according to any one of claims 8 to 11, wherein the signal analysis further includes a calculation of the MEL coefficients of the acoustomyographic signal.

13. A method for recognizing gestures and / or people according to claim 12, wherein the calculation of the Mel coefficients comprises: - passing the frequency spectrum through a bank of filters distributed according to the Mel scale to obtain a set of coefficients representing the power of the signal in different frequency bands, each frequency band being defined according to the Mel scale; - taking the logarithm of the output of each filter in the filter bank; - applying a discrete cosine transform (DCT) to the logarithm to calculate the Mel coefficients.

14. A method for recognizing gestures and / or people according to claim 13, wherein the filter bank comprises between 5 and 15 filters, preferably between 7 and 13 filters, preferably 7 filters, and the filter bank is spread over a frequency range between 0 Hz and 300 Hz, preferably between 0 Hz and 100 Hz.

15. A method for recognizing gestures and / or people according to any one of claims 8 to 14, wherein the classifier is one of: - an artificial neural network (ANN) model; - a Support Vector Machine (SVM) model; - a k-nearest neighbors (KNN) algorithm

Citation Information

Patent Citations

  • A mechanomyography apparatus and associated methods

    US20230210403A1

  • Stethoscope

    US6438238B1