A microphone-based human-computer interaction method, device, medium and system

By using a microphone-based human-computer interaction method, combined with adaptive weighted fusion of transient pulse detection and spectrum analysis, the problems of increased hardware costs and insufficient detection accuracy in existing technologies are solved. This achieves multi-dimensional interactive input with high recognition rate and low false alarm rate, thus improving the user experience.

CN122431537APending Publication Date: 2026-07-21SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
Filing Date
2026-06-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In the existing technology, the human-computer interaction method of electronic devices relies on physical buttons, which leads to increased hardware costs and structural complexity. At the same time, the microphone interaction method has insufficient detection accuracy in complex noise environments and is prone to misidentification.

Method used

By using the microphone on the electronic device as a pressure detection sensor, and combining transient pulse detection and spectral feature analysis with an adaptive weight fusion mechanism, the system can identify press and release events, and dynamically adjust the weights according to environmental noise conditions to generate interactive control commands.

Benefits of technology

It improves the detection accuracy and anti-interference capability of human-computer interaction without increasing hardware costs, provides multi-dimensional interactive input, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431537A_ABST
    Figure CN122431537A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of human-computer interaction, and particularly relates to a microphone-based human-computer interaction method, device, medium and system, which comprises the following steps: acquiring audio signals collected by at least one microphone on an electronic device; performing transient pulse detection on the audio signals, identifying short-time sound pulse signals and calculating pulse signal-to-noise ratios; performing spectrum feature analysis on the audio signals, detecting spectrum change features of environmental sound signals before and after pressing; dynamically adjusting comprehensive confidence degree weighting fusion weights according to the pulse signal-to-noise ratios, and judging whether an effective event occurs; and if the effective event occurs, generating a corresponding interaction control instruction. Compared with the prior art, the application solves the problems of increasing physical structures, improving production costs, being unable to be compatible with existing production lines, and being only dependent on microphones, which has insufficient detection accuracy and is prone to misidentification. The scheme realizes microphone pressing interaction without relying on additional hardware, and has high recognition rate and low false alarm rate under various environmental noise conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction technology, specifically relating to a microphone-based human-computer interaction method, device, medium, and system. Background Technology

[0002] Taking smartphones, headphones, and other electronic devices as examples, in existing technologies, user interaction with devices is mainly achieved through physical buttons and touchscreens. However, as smart devices become increasingly feature-rich, the number of physical buttons reserved in existing devices often falls short of actual interaction needs.

[0003] In the smartphone industry, users typically need to reuse the power button or add additional physical buttons to achieve specific functions. For example, to wake up the phone and interact with a smart voice assistant, users need to press the power or volume buttons; when taking a selfie, users usually need to press the virtual shutter button on the screen with their thumb, which is inconvenient for one-handed operation and poses a risk of slipping. Currently, some smartphone manufacturers are trying to alleviate this problem by adding dedicated physical buttons, but this increases hardware costs and structural complexity. In the field of over-ear noise-canceling headphones, operations such as adjusting volume, switching tracks, and turning noise cancellation on / off usually rely on multiple physical buttons. Although mature over-ear noise-canceling headphones generally have multiple microphones (such as call microphones and noise-canceling microphones), the human-computer interaction process still mainly relies on physical buttons or the reuse of existing buttons.

[0004] If press-to-interact functionality can be achieved using existing microphones without adding extra hardware, the number of buttons can be effectively reduced or the reuse of existing buttons decreased, thus improving the user experience. Therefore, how to provide richer interaction methods for electronic devices without increasing hardware costs is a pressing technical problem to be solved in this field.

[0005] In a prior application (Chinese Patent CN121585156B), the applicant proposed a microphone-based pressure-sensitive input button scheme. This scheme constructs a miniature variable acoustic cavity that changes with the degree of pressure by setting a ring-shaped pressure-sensitive structure around the microphone's sound-transmitting hole, and calculates the pressure level by monitoring changes in the sound signal. This scheme achieves pressure-sensitive input without occupying additional body space, but it still requires adding a physical component—the pressure-sensitive structure—to the device, necessitating certain structural modifications and making it incompatible with existing production lines. Furthermore, existing technologies have proposed a method for detecting user tapping operations by monitoring the energy level of low-frequency components in the microphone audio signal (see US Patent US8634565B2). This method extracts low-frequency components through low-pass filtering and monitors their energy changes to simplify the tapping detection process; however, this scheme only utilizes low-frequency energy monitoring, and its changes during pressing and releasing actions are not significant, easily leading to misjudgments. The detection accuracy in complex noise environments still has considerable room for improvement. Furthermore, CN107094274A also discloses a wireless earphone operation method, device, and wireless earphone, which determines whether a valid touch has been made by calculating the spectral energy within a preset time period and determines the instruction based on the number of valid touches; however, this solution only compares the detected spectral energy with a preset benchmark value, without considering the misjudgment problem caused by environmental factors, and this solution only makes a simple judgment based on a preset time period, which is prone to misidentifying the actual number of touches. Summary of the Invention

[0006] The purpose of this invention is to address at least one of the aforementioned problems by providing a microphone-based human-computer interaction method, device, medium, and system. This addresses the issues of existing technologies where adding physical structures for pressure-sensitive input increases production costs and makes the technology incompatible with existing production lines, while existing microphone-only interaction methods suffer from insufficient detection accuracy and susceptibility to false recognition. This solution utilizes the existing microphone of electronic devices as a pressure detection sensor, achieving microphone-based pressure interaction without relying on additional hardware. Testing shows that this method exhibits high recognition rate and low false alarm rate under various environmental noise conditions.

[0007] The objective of this invention is achieved through the following technical solution: The first aspect of this invention discloses a microphone-based human-computer interaction method, comprising the following steps: S1. Acquire audio signals collected by at least one microphone on the electronic device; S2. Perform transient pulse detection on the audio signal acquired in step S1, identify the short-time sound pulse signal generated when a finger presses or lifts at the microphone's sound-transmitting hole, and calculate the pulse signal-to-noise ratio. S3. Perform spectral feature analysis on the audio signal collected in step S1 to detect the spectral change characteristics of the ambient sound signal before and after the finger presses the microphone sound hole. S4. Based on the pulse signal-to-noise ratio calculated in step S2, dynamically adjust the weights of the transient event confidence obtained from transient pulse detection and the contact state change confidence obtained from spectral feature analysis in the weighted fusion of comprehensive confidence, and determine whether a valid microphone pressing event and / or release event has occurred through comprehensive confidence. S5. If a valid press event and / or release event is determined to have occurred in step S4, then a corresponding interactive control command is generated; if a valid press event and / or release event is determined not to have occurred in step S4, then return to step S1.

[0008] Preferably, in step S2, the transient pulse detection is performed using one or more of the following methods: environmental sound signal adaptive thresholding, short-time energy-threshold joint detection, matched filtering, correlation detection, and machine learning classifier. When a finger approaches and presses the microphone's sound hole, a short-duration acoustic pulse signal with the first polarity is detected. When the finger is lifted and moved away from the microphone's sound hole, a short-duration acoustic pulse signal with a second polarity is detected. Among them, the short-time acoustic pulse signal with the first polarity has the opposite polarity to the short-time acoustic pulse signal with the second polarity; The candidate pulse selected based on transient pulse detection has at least one feature, and the transient event confidence P_pulse obtained based on transient pulse detection is calculated and output through normalization.

[0009] Preferably, step S2 further includes identifying the pressure intensity and release speed through a short-duration acoustic pulse signal: Establish a positive correlation between the pulse amplitude of a short-time acoustic pulse signal with first polarity and the pressure intensity, and map the finger pressure intensity based on the peak amplitude of the detected short-time acoustic pulse signal with first polarity. A positive correlation mapping relationship is established between the pulse amplitude of the short-time acoustic pulse signal with second polarity and the release speed. The finger release speed is mapped based on the peak amplitude of the detected short-time acoustic pulse signal with second polarity.

[0010] Preferably, in step S3, the spectral feature analysis includes performing frame windowing and Fourier transform on the audio signal to obtain the spectrum of each frame. Based on the spectral difference before and after pressing or between the candidate pressing window and the adjacent non-contact window, the high-frequency energy attenuation, the low-frequency contact noise rise, and / or the change in the high-low frequency energy ratio are calculated. When the high-frequency energy attenuation, the low-frequency contact noise rise, and / or the change in the high-low frequency energy ratio meet the preset conditions, the contact state change confidence P_spectrum obtained based on the spectral feature analysis is output.

[0011] Preferably, step S3 further includes identifying the pressure intensity through spectral feature analysis: Establish a mapping relationship between high-frequency energy attenuation and pressing force, and map the pressing force based on the high-frequency energy attenuation obtained from the analysis.

[0012] Preferably, in step S4, the overall confidence level is calculated using the following formula: Score = α(SNR) × P_pulse + [1-α(SNR)] × P_spectrum; In the formula, P_pulse is the confidence level of the transient event obtained based on transient pulse detection, P_spectrum is the confidence level of the contact state change obtained based on spectral feature analysis, and α(SNR) is a weighting function with pulse signal-to-noise ratio as the independent variable; The weighting function is dynamically adjusted in the following way: When the pulse signal-to-noise ratio (SNR) is greater than the preset SNR threshold, α(SNR) > 0.5; when the pulse SNR is equal to the preset SNR threshold, α(SNR) = 0.5; when the pulse SNR is less than the preset SNR threshold, α(SNR) < 0.5. Whether a valid microphone press event and / or release event has occurred is determined by comparing the overall confidence level with a preset threshold.

[0013] Preferably, the preset signal-to-noise ratio threshold is 6dB.

[0014] Preferably, the following steps are also included: T1: Acquire motion signals from the electronic device, and / or acquire a second audio signal from the electronic device that is distinct from the target microphone (e.g., an unpressed microphone); T2: Analyze the signal fluctuation characteristics caused by pressing or releasing actions in the motion signal and output the pressing confidence P_inertial based on the inertial signal, and / or analyze the difference characteristics between the audio signal and the second audio signal and output the pressing confidence P_diff based on the multi-microphone difference; Therefore, when performing the comprehensive confidence calculation in step S4, independent weights are set for the pressing confidence P_inertial based on inertial signals and / or the pressing confidence P_diff based on multi-microphone differences, and the weights of the transient event confidence obtained based on transient pulse detection and the contact state change confidence obtained based on spectral feature analysis in the comprehensive confidence calculation are dynamically adjusted according to the remaining weights.

[0015] A second aspect of the present invention discloses a microphone-based human-computer interaction device for implementing the human-computer interaction method as described in any of the preceding claims; The human-computer interaction device includes: An audio acquisition unit is used to acquire audio signals collected by at least one microphone on an electronic device; The transient pulse detection unit is used to identify short-duration acoustic pulse signals generated when a finger is pressed or lifted, and to calculate the pulse signal-to-noise ratio; The spectrum analysis unit is used to detect the spectral changes of the ambient sound signal before and after pressing; An adaptive weight calculation unit is used to dynamically adjust the weights of the transient event confidence level obtained from transient pulse detection and the contact state change confidence level obtained from spectral feature analysis based on the pulse signal-to-noise ratio. The comprehensive judgment unit is used to determine whether a valid press event and / or release event has occurred based on the weighted and fused comprehensive confidence level. The instruction generation unit is used to generate interactive control instructions when a valid event is determined to have occurred.

[0016] Preferably, the human-computer interaction device further includes an inertial signal analysis unit and / or a multi-microphone difference analysis unit, which are used to acquire signals collected by the inertial sensor integrated in the electronic device and / or at least one reference microphone that is different from the target microphone, and to input the output P_inertial and / or P_diff into the comprehensive judgment unit for judgment.

[0017] A third aspect of the present invention discloses a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the microphone-based human-computer interaction method as described in any of the preceding claims.

[0018] The fourth aspect of the present invention discloses a microphone-based human-computer interaction system, including a processor and a memory for executable instructions, wherein the processor executes the instructions to implement the steps of the microphone-based human-computer interaction method as described in any of the preceding claims.

[0019] Compared with the prior art, the present invention has the following beneficial effects: 1. Zero increase in hardware cost: This invention directly reuses the existing microphone on electronic devices as a pressure detection sensor, without adding any additional physical buttons or sensors. It achieves rich interactive functions without increasing hardware costs. Compared to the applicant's previous solution requiring the addition of a physical pressure-sensitive structure, this invention requires no modification to the device hardware. Pressure detection is achieved solely through software algorithms that reuse the existing microphone, resulting in lower deployment costs and wider device compatibility.

[0020] 2. Strong anti-interference capability and good adaptability: Because the finger presses close to the microphone's sound-transmitting hole, the pressing and releasing signal strength is very high, which, according to actual tests, is usually equivalent to environmental noise of more than 80dB. Meanwhile, this invention employs an adaptive weighted fusion mechanism based on signal-to-noise ratio (SNR). Under high SNR conditions, it mainly relies on highly reliable transient pulse characteristics, while under low SNR conditions, it automatically enhances its reliance on spectral state characteristics, thus maintaining a high recognition rate and a low false alarm rate across all scenarios.

[0021] 3. Rich interactive dimensions: This invention can not only identify whether a press is present or absent, but also identify the pressing intensity and release speed through pulse amplitude, and identify the pressing intensity through the degree of spectral attenuation, providing electronic devices with multi-dimensional pressing interactive input.

[0022] 4. Wide range of applications: This invention can be widely applied to various electronic devices equipped with microphones, such as smartphones, headphones, true wireless earphones, smartwatches, and tablets.

[0023] 5. Enhanced Interactive Experience: Taking smartphones as an example, users can press the microphone on the back of the phone (such as a noise-canceling microphone) to take photos, activate the intelligent voice assistant, answer / hang up calls, etc., making one-handed operation more convenient and effectively reducing the risk of the device slipping. Taking headphones as an example, existing microphones can be used to achieve functions such as volume adjustment, track switching, and noise cancellation on / off, reducing the number of physical buttons or reducing button reuse.

[0024] Furthermore, unlike existing technologies, this invention does not rely solely on a single physical quantity (such as pulse energy, low-frequency energy, or preset spectral energy) for judgment. Instead, it creatively utilizes evidence from two dimensions simultaneously: the transient sound pulse generated by the pressing action and the continuous spectral state changes before and after pressing. The fusion weights of these two pieces of evidence are adaptively adjusted based on the current signal-to-noise ratio of the acoustic environment. This dual-branch adaptive fusion mechanism enables the invention to achieve high recognition rates and low false alarm rates in all noisy scenarios without adding any hardware—a feat not revealed or taught in existing technologies. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the human-computer interaction method provided in Embodiment 1 of the present invention.

[0026] Figure 2 This is a schematic diagram of the time-domain waveforms of the pressing pulse and the releasing pulse in Embodiment 1 of the present invention.

[0027] Figure 3 This is a schematic diagram of the spectral changes before and after pressing in two environments, quiet and noisy, according to Embodiment 1 of the present invention (top: quiet environment; bottom: noisy environment).

[0028] Figure 4 This is a schematic diagram of the adaptive weight fusion judgment process in Embodiment 1 of the present invention.

[0029] Figure 5 This is an architectural block diagram of each functional unit in the human-computer interaction device provided in Embodiment 2 of the present invention.

[0030] Figure 6 This is a schematic flowchart of the method for fusion inertial sensor-assisted judgment provided in Embodiment 3 of the present invention.

[0031] Figure 7 This is a schematic flowchart of the multi-microphone differential judgment method provided in Embodiment 4 of the present invention.

[0032] In the diagram: 201-Audio Acquisition Unit; 202-Transient Pulse Detection Unit; 203-Spectrum Analysis Unit; 204-Adaptive Weight Calculation Unit; 205-Comprehensive Judgment Unit; 206-Instruction Generation Unit. Detailed Implementation

[0033] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0034] Unless otherwise specified in the following description, any matters not covered herein may be handled using existing technologies.

[0035] Example 1 Please see Figure 1 This embodiment provides a human-computer interaction method for electronic devices based on microphone pressing, which is applicable to electronic devices such as smartphones, headphones, true wireless headphones, smartwatches, and tablets equipped with at least one microphone.

[0036] The method includes the following steps: Step S1: Acquire audio signals collected by the microphone. The electronic device continuously collects ambient sound signals through at least one microphone configured on it. The microphone can be any one or more of a call microphone, a noise-canceling microphone, or a microphone array. Taking a smartphone as an example, a noise-canceling microphone located on the back of the device can be preferred; taking a headset as an example, an existing noise-canceling microphone or a designated microphone from a microphone array can be selected.

[0037] Step S2: Identify the short-duration acoustic pulse signal with a first polarity generated when a contact object (such as a finger) presses the microphone's sound-transmitting hole, and / or the short-duration acoustic pulse signal with a second polarity generated when the contact object is lifted away from the microphone's sound-transmitting hole. The first and second polarity characteristics are used to distinguish between pressing and releasing actions, and they exhibit opposite polarities.

[0038] The physical mechanism of the press pulse is as follows: when a user presses the microphone's sound hole with their finger, the sound hole is instantly closed, and the air in the microphone's front cavity is instantaneously compressed, creating a positive pressure change. The microphone diaphragm responds by generating a short-duration sound pulse with the first polarity (positive). The physical mechanism of the release pulse is as follows: when the user lifts their finger away from the sound hole, the sound hole instantly opens, and the air in the microphone's front cavity is instantaneously drawn out, creating a reverse pressure change. The microphone diaphragm responds by generating a short-duration sound pulse with the second polarity (reverse). Because the finger is close to the microphone, the signal strength of both the press and release pulses is very high, typically equivalent to ambient noise of over 80dB in actual tests, exhibiting a high signal-to-noise ratio.

[0039] Please see Figure 2The press and release pulses, in the time domain, are pulse waveforms whose amplitude changes drastically within a short time window and significantly exceeds the ambient noise reference level. Transient pulse detection can be achieved using one or a combination of the following methods: setting an adaptive threshold, and determining a candidate pulse when the instantaneous amplitude of the signal exceeds the ambient noise reference value by a certain multiple (e.g., 3-5 times); identifying pulse characteristic waveforms through matched filtering or correlation detection; and using short-time energy and short-time zero-crossing rate for joint judgment. Among these, transient pulse detection preferably adopts the conventional short-time energy-threshold joint detection method in this field, specifically including: first, estimating the root mean square value of ambient noise RMS_noise and background short-time energy E_noise using audio samples within a preset time window before the pulse occurs, and setting the amplitude threshold T_A = k_A × RMS_noise and the energy threshold T_E = k_E × E_noise, where k_A and k_E are preset or calibrated coefficients; secondly, within a sliding window, the short-time energy E[n], peak amplitude A_peak[n], and short-time zero-crossing rate Z[n] of the current frame are calculated. When A_peak[n] and / or E[n] exceed the corresponding threshold and the duration is less than the preset pulse upper limit, it is determined as a candidate pulse; finally, a second confirmation is performed by combining waveform polarity, peak duration, and short-time zero-crossing rate to distinguish between press pulses and release pulses. Specifically, when the short-time energy E[n] exceeds the high threshold and the short-time zero-crossing rate Z[n] is within the preset low-frequency range, it is determined as the characteristic interval of the press pulse starting point; this dual-threshold logic can effectively filter the interference of high-frequency random thermal noise on pulse detection. The above method can also be replaced by conventional pulse recognition methods such as matched filtering, correlation detection, or machine learning classifiers according to actual needs.

[0040] Simultaneously, in step S2, the transient event confidence score P_pulse based on transient pulse detection is output. P_pulse can be calculated from one or more features among the candidate pulse's peak amplitude, short-time energy exceeding threshold, pulse width matching degree, waveform polarity consistency, and start-end edge time structure. As an example, the peak amplitude exceeding threshold ratio r_A, short-time energy exceeding threshold ratio r_E, pulse width matching degree r_W, and polarity consistency r_P can be normalized to the [0-1] interval, and calculated according to P_pulse = σ(w1·r_A + w2·r_E + w3·r_W + w4·r_P - b), where σ is a logistic function (Sigmoid / Logistic mapping, the same below), and w1-w4 and b are parameters obtained through pre-experiment calibration or training; alternatively, a lookup table or piecewise linear function can be used to map the above features to confidence scores in the [0-1] interval. When a candidate pulse does not meet the preset pulse width, polarity, or start / end edge structure conditions, P_pulse is reduced or set to 0. Further, the pulse signal-to-noise ratio (SNR) of the currently detected pulse signal is estimated. The SNR can be defined as the ratio of the pulse peak amplitude to the root mean square (RMS) value of the ambient noise for a period of time before the pulse occurs, for example, SNR = 20log10(A_peak / (RMS_noise + ε)), where ε is a small constant to prevent division by zero. Actual testing shows that in typical usage scenarios, its sound pressure level is usually equivalent to ambient noise above 80dB, and the pulse SNR is generally above 10dB; however, in high ambient noise scenarios (such as streets, subways, etc.), the pulse SNR may drop to 6dB or even lower.

[0041] Furthermore, step S2 also includes the identification of pressing pressure and release speed. Actual testing revealed a positive correlation between the peak amplitude of the pressing pulse and the finger pressing pressure, and a positive correlation between the peak amplitude (absolute value) of the release pulse and the finger release speed. Therefore, a mapping relationship between pulse amplitude and force / speed can be established in advance. An exemplary implementation is as follows: before the device leaves the factory or during user calibration, pulse amplitude data is collected when the user presses with different forces and releases at different speeds, and a linear regression model or piecewise lookup table is established. In actual use, based on the detected pulse amplitude, the corresponding pressing pressure value and release speed value are calculated through interpolation or regression, and output as interactive parameters. For example, the pressing pressure can be divided into three levels: light press, medium press, and heavy press, each mapping to different interactive functions.

[0042] Step S3: Spectral Feature Analysis. Perform spectral feature analysis on the audio signal to detect the spectral changes in the ambient sound signal before and after pressing. Please refer to [link / reference]. Figure 3 , Figure 3The changes in acoustic characteristics during microphone pressing under different acoustic environments are illustrated. In a quiet environment, both pressing and releasing actions generate short transient pulse characteristics in the audio time-frequency diagram. In a noisy environment, because the finger presses down and blocks the microphone's sound transmission aperture, the ambient sound energy received by the microphone is significantly reduced during pressing and recovers after release. Therefore, combining the pressing / releasing transient pulse characteristics with the changes in ambient sound spectrum energy before and after pressing and blocking can serve as an important basis for determining whether the microphone is being pressed.

[0043] Spectral feature analysis may specifically include: performing frame windowing and Fourier transform on the audio signal to obtain the spectrum of each frame; using the adjacent non-contact windows before and after the candidate pressing window as a reference, calculating the changes (ΔR) of high-frequency energy E_H, low-frequency energy E_L, and the high-low frequency energy ratio R = E_H / (E_L + ε). As an example, the high-frequency attenuation A_H = (E_H0 - E_H) / (E_H0 + ε) can be defined, where E_H0 is the high-frequency energy in the reference window; the low-frequency contact noise rise A_L = max((E_L - E_L0) / (E_L0 + ε), 0) can be defined, where E_L0 is the low-frequency energy in the reference window; then, A_H, A_L and the change in the high-low frequency energy ratio are normalized to the [0-1] interval, and the confidence level P_spectrum of the contact state change obtained from the spectral feature analysis is output according to P_spectrum = σ(c1·A_H + c2·A_L + c3·ΔR - b_s) or lookup table method. Here, c1-c3 and b_s are parameters obtained through pre-experiment calibration or training. Under strong ambient sound excitation conditions, the weight of the high-frequency energy attenuation A_H can be relatively high; under quiet or moderate noise conditions, the low-frequency contact noise rise A_L can be used as supplementary evidence. Thus, P_spectrum is calculated. It can be seen that P_spectrum is not simply given by a single threshold, but is given by the normalized fusion result of multiple spectral variation features.

[0044] Furthermore, this step also includes pressure intensity recognition based on the degree of spectral attenuation. Actual testing revealed that the tighter the finger presses on the microphone's sound-through hole, the more tightly the hole is sealed, and the greater the attenuation of high-frequency energy. Therefore, a mapping relationship between high-frequency attenuation and pressure intensity can be pre-established. An exemplary implementation is as follows: define the high-frequency attenuation ΔE_high as the difference (or ratio) between the high-frequency energy before and after pressing; calibrate the ΔE_high values ​​corresponding to different pressure intensities experimentally, and establish a lookup table or fitting curve. In practical use, the current pressure intensity level is deduced based on the real-time calculated high-frequency attenuation.

[0045] Furthermore, the pressure intensity recognition result obtained in step S3 based on the degree of spectral attenuation can be fused with the pressure intensity recognition result based on the pulse amplitude in step S2 to improve the accuracy of pressure intensity recognition. Specifically, the pressure intensity recognition result can be obtained by weighted fusion: the pressure intensity level obtained based on the pulse peak amplitude (step S2) is denoted as L_pulse, and the pressure intensity level obtained based on the high frequency attenuation (step S3) is denoted as L_spectrum. First, the continuous fused pressure intensity L_cont = η(SNR) × L_pulse + [1-η(SNR)] × L_spectrum is calculated. Wherein, η(SNR) is a weighting function positively correlated with the pulse signal-to-noise ratio. It can be obtained through a lookup table calibrated in pre-experimentation, or fitted as η(SNR) = η_min + (η_max - η_min) / (1 + exp(-(SNR - S0) / τ)), where η_min, η_max, S0, and τ are calibration parameters. For example, when SNR > 15dB, η = 0.8; when 6dB ≤ SNR ≤ 15dB, η = 0.5; when SNR < 6dB, η = 0.2. Using this method, the pressure intensity can be determined by referring more to the stable spectral attenuation level in high-noise environments.

[0046] Step S4: Adaptive weight fusion determination. Please refer to [link / reference]. Figure 4 Based on the pulse signal-to-noise ratio estimated in step S2, the fusion weight (weighting function α(SNR)) of the transient pulse detection result (the transient event confidence P_pulse obtained from the transient pulse detection in step S2) and the spectral feature analysis result (the contact state change confidence P_spectrum obtained from the spectral feature analysis) is dynamically adjusted. Based on the weighted fusion result (comprehensive confidence score), it is determined whether a valid pressing event and / or releasing event has occurred.

[0047] The specific strategies for weight adjustment are as follows: 1) When the environment is relatively quiet and the pulse signal-to-noise ratio is high (e.g., greater than 12dB), the press pulse and release pulse signals are very significant and the false alarm probability is extremely low. At this time, the transient pulse detection result is given a high weight (e.g., α = 0.9), and the spectral feature analysis result is only used as an auxiliary reference. 2) When the ambient noise increases and the pulse signal-to-noise ratio decreases, the weight of the spectral feature analysis results is gradually increased. Typically, when the pulse signal-to-noise ratio is about 6dB, the reliability of the two detection methods is roughly equivalent. At this time, the weighting coefficient α = 0.5 is set, that is, the two are fused with equal weight. (After a large number of experimental tests, it was found that when the pulse signal-to-noise ratio is about 6dB, the accuracy of transient pulse detection is basically equivalent to the accuracy of state judgment of spectral feature analysis. At this time, fusion of the two with equal weight can obtain the optimal comprehensive detection performance.) 3) When the ambient noise is extremely high and the pulse signal-to-noise ratio is further reduced (e.g., below 3dB), the false alarm rate and false negative rate of transient pulse detection both increase significantly. At this time, the judgment mainly relies on the results of spectral feature analysis, and the weighting coefficient α approaches 0.

[0048] The overall confidence score can be calculated using the following formula: Score = α(SNR) × P_pulse + [1-α(SNR)] × P_spectrum Wherein, P_pulse is the confidence level of the transient event obtained based on transient pulse detection, used to characterize the reliability of the short-term event of pressing or releasing; P_spectrum is the confidence level of the contact state change obtained based on spectral feature analysis, used to characterize the reliability of the microphone's sound transmission hole changing from non-contact to contact or from contact to non-contact; α(SNR) is a weighting function with signal-to-noise ratio as the independent variable. This weighting function α(SNR) can be implemented through a preset lookup table or fitting formula, preferably α(SNR) = 1 / (1 + exp(-(SNR - S0) / τ)), where S0 is the transition signal-to-noise ratio when the two pieces of evidence are equally weighted, and τ is a smooth transition parameter to ensure a smooth transition when the signal-to-noise ratio changes.

[0049] A finger press / release event is confirmed when the score exceeds a preset threshold. Specifically: a press event is confirmed when a press pulse is detected and the overall confidence score exceeds a preset threshold; a release event is confirmed when a release pulse is detected and the overall confidence score exceeds a preset threshold.

[0050] Step S5: Generate interactive control commands. If a valid press event and / or release event is determined based on the comprehensive confidence level, corresponding interactive control commands are generated according to the preset mapping relationship. This mapping relationship can be flexibly configured according to different application scenarios and can provide richer interactive semantics by combining the recognized press intensity and release speed.

[0051] Exemplary interaction mappings include: 1) Smartphone scenario: (1) Single short press (force less than threshold F1): Photo shutter; (2) Single heavy press (force greater than threshold F2): wakes up the intelligent voice assistant; (3) Press twice in quick succession: Screenshot; (4) Long press (duration exceeds threshold T1): Start the recording function; (5) Quick release (release speed greater than threshold V1): Return to desktop; (6) Slow release (release speed is less than threshold V2): Open the notification bar.

[0052] 2) Over-ear headphone scenario: (1) Single short press: Play / Pause; (2) Two consecutive short presses: Next track; (3) Three consecutive short presses: Previous track; (4) Long press: Turn noise cancellation on / off; (5) Linear control of pressure: volume adjustment (the greater the pressure, the faster the volume changes).

[0053] 3) True wireless earphone scenario: (1) Single press: Answer / hang up the phone; (2) Long press: reject incoming calls or wake up the voice assistant.

[0054] All of the above thresholds can be preset as needed.

[0055] Example 2 Please see Figure 5 This embodiment provides a microphone-based human-computer interaction device for electronic devices, including a processor and a memory. When the processor executes the computer program stored in the memory (taking the human-computer interaction method given in Embodiment 1 as an example), it implements the following functional units: The audio acquisition unit 201 is used to acquire audio signals collected by at least one microphone on the electronic device. This unit can read the digital audio stream input from the microphone in real time and perform necessary preprocessing such as noise reduction and gain adjustment.

[0056] The transient pulse detection unit 202, communicatively connected to the audio acquisition unit 201, is used to identify short-duration acoustic pulse signals generated when a finger is pressed and / or released in an audio signal, and to estimate the pulse signal-to-noise ratio. This unit can employ methods such as energy threshold detection, waveform matching, or machine learning to achieve pulse recognition and signal-to-noise ratio calculation. Furthermore, this unit also outputs the pressing force value and / or release speed value based on the pulse peak amplitude.

[0057] The spectrum analysis unit 203, which is communicatively connected to the audio acquisition unit 201, is used to detect the spectral change characteristics of the ambient sound signal before and after pressing. This unit can perform time-frequency transformation on the audio signal, extract the energy distribution information of the high-frequency and low-frequency bands, and monitor their dynamic changes; in addition, this unit also outputs the pressing force value based on the high-frequency attenuation.

[0058] The adaptive weight calculation unit 204, communicatively connected to the transient pulse detection unit 202, is used to dynamically adjust the fusion weights of the transient pulse detection results and the spectral feature analysis results based on the pulse signal-to-noise ratio (SNR). This unit maintains a mapping relationship with the pulse SNR as input and the weight coefficients as output.

[0059] The comprehensive judgment unit 205 is communicatively connected to the transient pulse detection unit 202, the spectrum analysis unit 203, and the adaptive weight calculation unit 204, and is used to comprehensively judge whether a valid pressing event and / or releasing event has occurred based on the confidence level after weighted fusion.

[0060] The instruction generation unit 206, communicatively connected to the comprehensive judgment unit 205, is used to generate corresponding interactive control instructions when a valid event is determined to have occurred, and send the instructions to the device's operating system or related applications. These interactive control instructions can carry parameters such as pressing pressure and release speed for use by upper-level applications.

[0061] It should be noted that all of the above functional units are software-implemented functional modules, and their physical carriers are the processor and memory of electronic devices. The transient pulse detection unit and the spectrum analysis unit are both implemented by the same processor executing different algorithm programs in physical hardware.

[0062] Example 3 Please see Figure 6 This embodiment, based on embodiment 1, further integrates inertial sensors to assist in judgment.

[0063] When the electronic device is also equipped with inertial sensors such as accelerometers and / or gyroscopes, the method may further include: Step S31: Acquire motion signals sensed by the inertial sensor, including triaxial acceleration data and / or triaxial angular velocity data.

[0064] Step S32: Analyze the signal fluctuation characteristics caused by the pressing action in the motion signal and output the pressing confidence score P_inertial based on the inertial signal. Specifically, the three-axis acceleration and / or three-axis angular velocity signals can be first processed by DC removal, high-pass filtering, and window framing, and then the magnitude change, peak-to-peak value, energy mutation value, or attitude disturbance amount in each window can be calculated. P_inertial is not directly set high by any single feature exceeding the threshold, but is obtained by the normalized weighted result of multiple features. As an example, the normalized value of acceleration magnitude change a, the normalized value of angular velocity peak-to-peak value g, the normalized value of short-term impact energy j, and the local pressing pattern matching degree m can be input into the function P_inertial = σ(λ1·a + λ2·g + λ3·j + λ4·m - λ5·M_global - b_i), where M_global represents the long-duration motion intensity related to the overall shaking, grip adjustment, or walking, and λ1-λ5 and b_i are calibration parameters. When the motion characteristics meet the local short-term disturbance pattern and there is no violent movement of the whole device, P_inertial is high; when the whole device moves, shakes continuously, or the motion pattern is inconsistent with the local pressing, P_inertial is reduced or set to 0, so as to help suppress false triggering.

[0065] Using P_inertial as an additional confidence input, it participates in the adaptive weight fusion in step S4. The overall confidence score can then be expanded as follows: Score = α × P_pulse + β × P_spectrum + γ × P_inertial; Wherein, α, β, and γ are weighting coefficients dynamically adjusted according to the pulse signal-to-noise ratio and the quality of the inertial signal, satisfying α + β + γ = 1. Preferably, γ can be obtained by mapping the stability of the inertial signal and the impact amplitude separately. When the inertial signal is weak or the device is in a stationary state, γ can approach 0; when obvious hand-held shaking or inertial disturbance caused by pressing is detected, γ is increased accordingly; and α and β are redistributed according to the remaining confidence ratio (1 - γ) (dynamically allocated in the remaining confidence ratio in the aforementioned manner).

[0066] Example 4 Please see Figure 7 This embodiment, based on embodiment 1 or embodiment 3, further utilizes multi-microphone differential judgment.

[0067] When the electronic device is equipped with multiple microphones, the method further includes: Step S31': Acquire the first audio signal captured by the pressed microphone, and at least one second audio signal not captured by the pressed microphone.

[0068] Step S32': Analyze the difference features between the first audio signal and the second audio signal, and output the press confidence score P_diff based on multi-microphone differences. Specifically, the two signals can be time-synchronized and framed, and the energy ratio, spectral similarity, correlation coefficient or pulse arrival time difference within the same time window can be calculated. These difference features are then fused into a difference score D. When D exceeds a preset threshold, a high confidence score P_diff is output (the specific method is the same as in Example 3).

[0069] Using P_diff as an additional confidence input, it participates in the adaptive weight fusion in step S4 (which can adopt the same fusion method as in Example 3) to further improve detection performance.

[0070] Test case This embodiment provides a test example for verifying the detection performance of the present invention. This test example is based on the method of Embodiment 1, and further incorporates the inertial signal-assisted judgment and multi-microphone differential judgment from Embodiments 3 and / or 4. During the test, the area of ​​the smartphone facing away from the microphone is used as the pressing position. Data from the target microphone, reference microphone, and auxiliary sensors such as accelerometers and gyroscopes are collected. The audio signal is continuously analyzed according to a preset time window, and candidate press detection, event merging, and press type identification are completed.

[0071] In one test case, several testers were selected to conduct continuous pressing tests for several hours, including approximately a thousand pressing operations, covering typical contact conditions such as dry hands, wet hands, and wearing gloves. During the test, interference was introduced including ambient noise, occasional touches, changes in grip, changes in posture, and overall device movement. Test results showed that after adopting the signal-to-noise ratio adaptive dual-branch fusion and multimodal consistency filtering scheme of this invention, the event-level F1 score could reach approximately 96%-98%, and the false alarm rate under interference scenarios could be reduced to less than approximately 0.1 times / minute. In the semantic recognition after detection, the accuracy rate for single / double pressing and short / long press recognition could both reach over 97%.

[0072] Further comparative tests show that when using only the transient pulse branch (based solely on the transient pulse detection results from step S2 of this scheme) or only the spectral state change branch (based solely on the spectral feature analysis results from step S3 of this scheme), the detection results are easily affected by ambient noise, non-target touch, or grip changes, and the event-level F1 score is typically lower than that of the dual-branch fusion scheme. After adopting fixed-weight fusion, the detection stability is improved; further, by employing the signal-to-noise ratio adaptive weight fusion of this invention, the contribution ratio of transient pulse evidence and spectral state change evidence can be dynamically adjusted according to the current acoustic environment, thereby further reducing missed detections and false detections.

[0073] Further multimodal comparative tests showed that when using only the target microphone, the system was quite sensitive to ambient sound and overall device movement; with the addition of a reference microphone (Example 4), some shared ambient sound interference could be suppressed; and by combining inertial signal gating or motion pattern matching (Example 3), false triggering caused by non-pressing factors such as grip adjustment, walking, and posture changes could be further eliminated. Therefore, the reference microphone and inertial sensing signals can complement the localized press acoustic evidence from the target microphone, thereby improving the system's robustness in real-world usage scenarios.

[0074] Regarding the acoustic mechanism, the test results also show that when a finger touches the microphone's sound-transmitting hole, it alters the local acoustic channel characteristics: in a relatively quiet environment, the contact action easily introduces low-frequency contact noise; when the ambient noise is strong, pressing to block or altering the sound transmission path causes high-frequency energy attenuation. Therefore, this invention comprehensively considers high-frequency energy attenuation, low-frequency contact noise changes, and high-low frequency energy ratio changes in the spectral branches, which is beneficial for maintaining stable recognition under different noise conditions.

[0075] It should be noted that the technical features in the above embodiments can be combined to form more implementation schemes. For example, inertial sensor and multi-microphone differential judgment can be fused simultaneously to form an adaptive weighted fusion scheme of four signals (transient pulse, spectral features, inertial signal, and multi-microphone difference).

[0076] This invention reuses existing microphones in electronic devices as pressure detection sensors, achieving rich pressure-based interactive functions without increasing hardware costs. It employs an adaptive weighted fusion mechanism based on signal-to-noise ratio to ensure excellent robustness and accuracy under varying noise conditions. Furthermore, this invention can also identify pressure intensity and release speed, providing more dimensions for interaction. This solution is applicable to various electronic devices equipped with microphones, such as smartphones, headphones, true wireless earbuds, smartwatches, and tablets, and has good industrial applicability and widespread application value.

[0077] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0078] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0079] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0081] The above description of the embodiments is provided to enable those skilled in the art to understand and use the invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without inventive effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the invention should be within the protection scope of the present invention.

Claims

1. A microphone-based human-computer interaction method, characterized in that, Includes the following steps: S1. Acquire audio signals collected by at least one microphone on the electronic device; S2. Perform transient pulse detection on the audio signal acquired in step S1, identify the short-time sound pulse signal generated when a finger presses or lifts at the microphone's sound-transmitting hole, and calculate the pulse signal-to-noise ratio. S3. Perform spectral feature analysis on the audio signal collected in step S1 to detect the spectral change characteristics of the ambient sound signal before and after the finger presses the microphone sound hole. S4. Based on the pulse signal-to-noise ratio calculated in step S2, dynamically adjust the weights of the transient event confidence obtained from transient pulse detection and the contact state change confidence obtained from spectral feature analysis in the weighted fusion of comprehensive confidence, and determine whether a valid microphone pressing event and / or release event has occurred through comprehensive confidence. S5. If a valid press event and / or release event is determined to have occurred in step S4, then a corresponding interactive control command is generated; if a valid press event and / or release event is determined not to have occurred in step S4, then return to step S1.

2. The microphone-based human-computer interaction method according to claim 1, characterized in that, In step S2, the transient pulse detection is performed using one or more of the following methods: environmental sound signal adaptive thresholding, short-time energy-threshold joint detection, matched filtering, correlation detection, and machine learning classifier. When a finger approaches and presses the microphone's sound hole, a short-duration acoustic pulse signal with the first polarity is detected. When the finger is lifted and moved away from the microphone's sound hole, a short-duration acoustic pulse signal with a second polarity is detected. Among them, the short-time acoustic pulse signal with the first polarity has the opposite polarity to the short-time acoustic pulse signal with the second polarity; Based on the transient pulse detection results, at least one feature of the candidate pulse is selected, and the transient event confidence P_pulse obtained from the transient pulse detection is calculated and output through normalization.

3. The microphone-based human-computer interaction method according to claim 1, characterized in that, Step S2 also includes identifying the pressure intensity and release speed through short-duration acoustic pulse signals: Establish a positive correlation between the pulse amplitude of a short-time acoustic pulse signal with first polarity and the pressure intensity, and map the finger pressure intensity based on the peak amplitude of the detected short-time acoustic pulse signal with first polarity. A positive correlation mapping relationship is established between the pulse amplitude of the short-time acoustic pulse signal with second polarity and the release speed. The finger release speed is mapped based on the peak amplitude of the detected short-time acoustic pulse signal with second polarity.

4. The microphone-based human-computer interaction method according to claim 1, characterized in that, In step S3, the spectral feature analysis includes performing frame windowing and Fourier transform on the audio signal to obtain the spectrum of each frame. Based on the spectral difference before and after pressing or between the candidate pressing window and the adjacent non-contact window, the high-frequency energy attenuation, the low-frequency contact noise rise, and / or the change in the high-low frequency energy ratio are calculated. When the high-frequency energy attenuation, the low-frequency contact noise rise, and / or the change in the high-low frequency energy ratio meet the preset conditions, the contact state change confidence P_spectrum obtained based on the spectral feature analysis is output.

5. The microphone-based human-computer interaction method according to claim 1, characterized in that, Step S3 also includes identifying the pressure intensity through spectral feature analysis: Establish a mapping relationship between high-frequency energy attenuation and pressing force, and map the pressing force based on the high-frequency energy attenuation obtained from the analysis.

6. The microphone-based human-computer interaction method according to claim 1, characterized in that, In step S4, the overall confidence level is calculated using the following formula: Score = α(SNR) × P_pulse + [1-α(SNR)] × P_spectrum; In the formula, P_pulse is the confidence level of the transient event obtained based on transient pulse detection, P_spectrum is the confidence level of the contact state change obtained based on spectral feature analysis, and α(SNR) is a weighting function with pulse signal-to-noise ratio as the independent variable; The weighting function is dynamically adjusted in the following way: When the pulse signal-to-noise ratio (SNR) is greater than the preset SNR threshold, α(SNR) > 0.5; when the pulse SNR is equal to the preset SNR threshold, α(SNR) = 0.5; when the pulse SNR is less than the preset SNR threshold, α(SNR) < 0.

5. Whether a valid microphone press event and / or release event has occurred is determined by comparing the overall confidence level with a preset threshold.

7. The microphone-based human-computer interaction method according to claim 1, characterized in that, It also includes the following steps: T1: Acquire motion signals from the electronic device, and / or acquire a second audio signal from at least one reference microphone on the electronic device that is distinct from the target microphone; T2: Analyze the signal fluctuation characteristics caused by pressing or releasing actions in the motion signal and output the pressing confidence P_inertial based on the inertial signal, and / or analyze the difference characteristics between the audio signal and the second audio signal and output the pressing confidence P_diff based on the multi-microphone difference; Therefore, when performing the comprehensive confidence calculation in step S4, independent weights are set for the pressing confidence P_inertial based on inertial signals and / or the pressing confidence P_diff based on multi-microphone differences, and the weights of the transient event confidence obtained based on transient pulse detection and the contact state change confidence obtained based on spectral feature analysis in the comprehensive confidence calculation are dynamically adjusted according to the remaining weights.

8. A microphone-based human-computer interaction device, characterized in that, Used to implement the human-computer interaction method as described in any one of claims 1-7; The human-computer interaction device includes: An audio acquisition unit (201) is used to acquire audio signals acquired by at least one microphone on an electronic device; The transient pulse detection unit (202) is used to identify short-duration acoustic pulse signals generated when a finger is pressed or lifted, and to calculate the pulse signal-to-noise ratio; The spectrum analysis unit (203) is used to detect the spectral change characteristics of the ambient sound signal before and after pressing; An adaptive weight calculation unit (204) is used to dynamically adjust the weights of the transient event confidence obtained based on transient pulse detection and the contact state change confidence obtained based on spectral feature analysis according to the pulse signal-to-noise ratio. The comprehensive judgment unit (205) is used to determine whether a valid pressing event and / or releasing event has occurred based on the comprehensive confidence level after weighted fusion. The instruction generation unit (206) is used to generate interactive control instructions when a valid event is determined to have occurred.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the microphone-based human-computer interaction method as described in any one of claims 1-7.

10. A microphone-based human-computer interaction system, characterized in that, The device includes a processor and a memory for executable instructions, wherein the processor, when executing the instructions, implements the steps of the microphone-based human-computer interaction method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Operation method and device of wireless headset, and wireless headset

    CN107094274A

  • Microphone-based pressure input key and method

    CN121585156B

  • Audio signal processing apparatus, audio signal processing method, and program

    US8634565B2