Microphone-based pressure input key and method
By constructing a miniature variable acoustic cavity and a ring-shaped pressure-sensitive structure inside the microphone, pressure-sensitive input is achieved by utilizing changes in the microphone's acoustic signal. This solves the problems of waterproofing, dustproofing, and cost associated with traditional buttons, and enables multi-level pressure-sensitive recognition and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-10
AI Technical Summary
Existing electronic device buttons have shortcomings in terms of waterproofing, dustproofing, human-computer interaction, and cost control. Traditional mechanical buttons and capacitive buttons require additional hardware sensors, resulting in limited device space and unstable sensitivity.
By constructing a miniature variable acoustic cavity inside the microphone, the pressure level is determined by monitoring changes in the sound signal. Combined with a ring-shaped pressure-sensitive structure and processing unit, pressure-sensitive input is achieved using the microphone's existing hardware.
It achieves multi-level pressure sensitivity recognition without the need for additional sensors, maintains a simple device appearance, reduces hardware costs, and improves recognition rate and stability.
Smart Images

Figure CN121585156B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, in particular to a microphone-based pressure input key and method. BACKGROUND
[0002] There are some defects in the existing technology of electronic device keys, such as the defects of traditional external mechanical keys in pressure recognition, overall appearance integration, waterproof and dustproof, and reliability. Therefore, considering waterproof and dustproof, human-computer interaction, cost control and other aspects, electronic terminal devices are gradually reducing the use of function keys, leading to a gradual reduction in the application of electronic devices such as mobile terminals, earphones / charging boxes, and wearable devices.
[0003] For example, CN105302373A discloses a method, system and mobile terminal for realizing mobile terminal operation according to touch signals, the method comprising the steps of: a fingerprint detection device of the mobile terminal receiving a touch signal and detecting a pressure level of the touch signal; and the mobile terminal performing corresponding operations according to the pressure level of the touch signal. However, although the pressure sensing scheme can replace traditional keys, it usually requires a pressure sensor with high cost.
[0004] For example, CN109189212A discloses a touch feedback method and device, including a finger pressing area acquisition device for acquiring a finger pressing area value when a finger is pressed; a control device for setting a pressure level interval according to the relationship between the finger pressing area value and the pressure value when the finger is pressed, the pressure level interval including a light pressing pressure interval, a pressing pressure interval and a third pressure interval between the two; when the acquired finger pressing area value is in the light pressing pressure interval, output a light pressing feedback signal to the feedback device; if not, determine whether the finger pressing area value is in the pressing pressure interval, if yes, output a pressing feedback signal to the feedback device, if not, determine whether the finger pressing area value is in the third pressure interval, if yes, output a third feedback signal to the feedback device. This scheme judges by pressing area, but it still needs to increase the corresponding sensor in the tight device space to capture the change of the pressing area; at the same time, the moisture-proof effect of the capacitive key is not ideal, leading to unstable touch sensitivity and difficulty in pressure recognition.
[0005] In the case of one or more microphones commonly installed in various types of intelligent emerging electronic terminal devices, if a key with certain pressure recognition function can be realized based on the microphone without adding hardware inside, it can meet the design requirements of electronic devices, especially intelligent terminal devices, in terms of simple appearance, low cost and human-computer interaction. SUMMARY
[0006] The present application aims to solve at least one of the above problems, and provides a microphone-based pressure-sensitive input button and method to solve the problem in the prior art that a pressure sensor / capacitance sensor needs to be additionally added to measure the pressing level. The present application establishes an association between the acoustic signal and the pressing form by constructing a miniature variable acoustic cavity that can change with the degree of finger contact and pressing at the microphone, so that the pressing level can be calculated by monitoring the acoustic signal collected by the microphone.
[0007] The present application aims to solve at least one of the above problems, and provides a microphone-based pressure-sensitive input button and method to solve the problem in the prior art that a pressure sensor / capacitance sensor needs to be additionally added to measure the pressing level. The present application establishes an association between the acoustic signal and the pressing form by constructing a miniature variable acoustic cavity that can change with the degree of finger contact and pressing at the microphone, so that the pressing level can be calculated by monitoring the acoustic signal collected by the microphone.
[0008] The present application aims to solve at least one of the above problems, and provides a microphone-based pressure-sensitive input button and method to solve the problem in the prior art that a pressure sensor / capacitance sensor needs to be additionally added to measure the pressing level. The present application establishes an association between the acoustic signal and the pressing form by constructing a miniature variable acoustic cavity that can change with the degree of finger contact and pressing at the microphone, so that the pressing level can be calculated by monitoring the acoustic signal collected by the microphone.
[0009] The microphone is integrated in the electronic device and is used to collect acoustic signals entering the interior of the electronic device through the shell thereof;
[0010] The sound-permeable hole is provided on the shell of the electronic device and is in communication with the microphone to form an acoustic channel;
[0011] The pressure-sensitive structure is a ring-shaped structure surrounding the sound-permeable hole and has at least one gap to form a neck connecting the sound-permeable hole and the atmosphere;
[0012] The processing unit is integrated in the electronic device and is in communication with the microphone, and is used to acquire and process the acoustic signals collected by the microphone, and then determine the pressing level of the pressure-sensitive input button according to the characteristics of the acoustic signals;
[0013] The pressure-sensitive structure, the user's finger and the shell together form a miniature acoustic cavity with variable acoustic resonance characteristics, and the gap constitutes an equivalent neck connecting the miniature acoustic cavity and the atmosphere, so that the change in the pressing level causes the equivalent resonance characteristics of the acoustic channel to change; the processing unit grades the pressing level based on the change in the equivalent resonance characteristics.
[0014] Preferably, the pressure-sensitive structure is a single-stage ring or a multi-stage ring arranged coaxially;
[0015] When the pressure-sensitive structure is a multi-stage ring, the height of the innermost stage ring to the outermost stage ring increases successively, and the height ratio between adjacent rings satisfies wherein, Hin is the height of the ring located on the inner side of the adjacent rings, Hout is the height of the ring located on the outer side of the adjacent rings.
[0016] Preferably, the cross-sectional shape of the ring-shaped structure is circular, quasi-circular or polygonal;
[0017] The gap has 1-6 gaps.
[0018] The second aspect of the present application discloses a microphone-based pressure input method, which adopts any of the above pressure input buttons;
[0019] The pressure input method comprises the following steps:
[0020] S10, signal acquisition and candidate event detection:
[0021] Obtain the signal collected by the microphone; determine the pressing event interval based on the signal according to the preset event detection criterion, and merge adjacent or overlapping pressing event intervals to obtain a candidate event segment; S20, preprocessing of the candidate event segment:
[0022] The candidate event segment obtained in step S10 is subjected to normalization processing and / or alignment processing;
[0023] S30, spectral estimation of the candidate event segment:
[0024] The candidate event segment obtained in step S20 is subjected to power spectral density estimation method to obtain power spectral density , and the preset frequency band of interest is subjected to smoothing processing and / or interpolation processing;
[0025] S40, feature extraction:
[0026] The power spectral density obtained in step S30 and the time-domain signal corresponding to the candidate event segment are extracted to obtain multi-dimensional features, and a feature vector for pressing force level classification determination is constructed ; wherein the multi-dimensional features at least include one of the center frequency , the quality factor , the frequency band energy , the energy ratio and the transient rise time
[0027] S60, fusion determination:
[0028] For each pressing level , a respective feature interval is constructed, and a threshold rule score is calculated ; a lightweight classification model is used to map the feature vector to a probability value of each level ; a fusion score is calculated , and a pressing force level classification determination is performed to obtain the pressing level of the current candidate event segment; the pressing level at least includes two or more levels of non-touch, light touch, medium press and deep press;
[0029] S70, de-bouncing output:
[0030] In a preset time window, majority voting or consistency judgment is performed on the determination results in the candidate event segment, and a final pressing level is output.
[0031] Preferably, in step S10, the time-domain signal collected by the microphone is preprocessed to calculate the normalized increment of short-time energy and passband energy in a sliding window; when and / or the background energy exceeds a preset multiple, the sliding window is marked as a pressing event interval, and adjacent frames are merged to obtain a candidate event segment.
[0032] Preferably, in step S10, when the ambient sound is insufficient, the device can be enabled to play a preset incentive sound to ensure the signal-to-noise ratio and repeatability.
[0033] Preferably, in step S10, the preprocessing includes pre-emphasis and band-pass filtering, and the signal collected by the microphone is processed by frame in a sliding window.
[0034] Wherein, the frame length of the sliding window is .
[0035] Preferably, in step S20, after normalization processing, windowing and / or framing are further performed to reduce spectral leakage.
[0036] Preferably, in step S30, the power spectrum density estimation method is one or a combination of multiple of non-parametric method, parametric method and time-frequency analysis method.
[0037] Preferably, in step S40, the energy ratio is the ratio of the frequency band energy to the full-band energy .
[0038] The quality factor is the ratio of the center frequency to the bandwidth , wherein the bandwidth is estimated by the half-height width method.
[0039] The transient rise time is the length of time when the energy exceeds the background threshold to reach the steady state.
[0040] Preferably, in step S60, the lightweight classification model is a small-scale neural network, a logistic regression or a decision tree.
[0041] The fusion score is calculated by the following formula: , wherein, The weight coefficient is between 0 and 1.
[0042] Preferably, the electronic device is provided with a plurality of microphones.
[0043] Step S20 further comprises clock synchronization and gain calibration of the time domain signals collected by different microphones and preprocessed.
[0044] Step S40 further comprises calculating the double-microphone differential spectrum and / or cross-spectrum coherence of the time domain signals collected by different microphones and preprocessed as one of the elements of the feature vector.
[0045] Between steps S40 and S60, step S50, double-microphone differential / coherent enhancement, is further included, which performs enhancement or weighting processing on the double-microphone features such as double-microphone differential spectrum and / or cross-spectrum coherence extracted in step S40, to suppress far-field noise and highlight near-field pressing features.
[0046] The working principle of the present application is as follows:
[0047] The pressure input button is provided with a notched annular pressure sensing structure around the microphone and sound transmission hole; when the finger contacts / presses the area, the finger surface and the annular structure and the device shell together form a variable volume micro acoustic cavity, according to the different pressing force, the notch and the gap will form different sound transmission paths. The structure is equivalent to an acoustic resonance network that changes with the pressing force, which can produce different responses to incident sound, thereby mapping different pressing forces into distinguishable acoustic features.
[0048] Compared with the prior art, the present application has the following beneficial effects:
[0049] The pressure input button does not need to add additional sensors and will not damage the hydrophobic membrane and the compact appearance design; by setting a first-order ring, it can stably achieve not less than two levels of pressure sensing recognition, and when further made into an n-order coaxial ring, the classification will naturally expand to n+1 levels.
[0050] If further double-microphone cooperation is added, the reference microphone can provide an environmental baseline and adaptive calibration, and the resolution and recognition rate of the classification are higher.
[0051] The pressure input button reuses the existing microphone hardware of the device, has small system changes to the device, and has very low hardware cost. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 The structure diagram of the first-order ring pressure sensing structure.
[0053] Figure 2 The structure diagram of the electronic device (a) and the second-order ring pressure sensing structure (b).
[0054] Figure 3 Fig. 2 is a schematic diagram of the structure side view of the second-order ring pressure sensing structure, wherein (a) is not pressed, (b) is slightly pressed, (c) is moderately pressed, and (d) is deeply pressed.
[0055] Figure 4 Fig. 3 is a schematic diagram of the structure top view of the second-order ring pressure sensing structure.
[0056] Figure 5 Fig. 4 is an effect diagram of the step and high-frequency suppression of the pressure sensing structure in the application example 1 at the natural frequency in the pressing process.
[0057] In the figure, 100 is a sound transmission hole, 101 is a first-order ring structure, 200 is an electronic device, 201 is a second-order pressure sensing button, 201a is a second-order inner ring structure, 201b is a second-order outer ring structure, and 203 is an electronic device shell. DETAILED DESCRIPTION
[0058] The application will be described in detail below with reference to the accompanying drawings and specific examples.
[0059] In the following description, if not specifically stated, the methods used are the means known in the art, and the remaining matters are the prior art.
[0060] Example 1
[0061] A microphone-based pressure sensing input button can refer to Figures 1-4 , which comprises a microphone, a sound transmission hole, a pressure sensing structure, and a processing unit.
[0062] The microphone is integrated in the internal part of the electronic device and is used to collect sound signals entering the internal part through the shell of the electronic device.
[0063] The sound transmission hole is opened on the shell of the electronic device, and the sound transmission hole and the microphone are in communication to form an acoustic channel.
[0064] The pressure sensing structure is a ring structure arranged around the sound transmission hole, and the pressure sensing structure is provided with at least one gap to form a neck part connecting the sound transmission hole and the atmosphere.
[0065] The processing unit is integrated in the internal part of the electronic device and is in communication connection with the microphone, and is used to acquire and process the sound signals collected by the microphone, and then grade the pressing force of the pressure sensing input button according to the characteristics of the sound signals.
[0066] A microphone-based pressure sensing input method is adopted by the pressure sensing input button as described above.
[0067] The pressure sensing input method comprises the following steps:
[0068] S10, signal acquisition and candidate event detection:
[0069] Obtaining a signal collected by a microphone; determining a pressing event interval according to a preset event detection criterion based on the signal, and merging adjacent or overlapping pressing event intervals to obtain a candidate event segment; specifically, pre-processing a time-domain signal collected by the microphone, calculating a normalized increment of short-time energy and passband energy in a sliding window; when and / or the background energy exceeds a preset multiple, marking the sliding window as a pressing event interval and merging adjacent frames to obtain a candidate event segment;
[0070] S20, pre-processing of the candidate event segment:
[0071] normalizing and / or aligning the candidate event segment obtained in step S10;
[0072] S30, spectral estimation of the candidate event segment:
[0073] obtaining a power spectral density of the candidate event segment obtained in step S20 by a power spectral density estimation method, and performing smoothing and / or interpolation processing on a preset frequency band of interest;
[0074] S40, feature extraction:
[0075] extracting multi-dimensional features from the power spectral density obtained in step S30 and the time-domain signal corresponding to the candidate event segment, and constructing a feature vector for pressing force level classification determination; wherein the multi-dimensional features at least include one of center frequency , quality factor , frequency band energy , energy ratio and transient rise time ;
[0076] S60, fusion determination:
[0077] constructing a respective feature interval for each pressing level , and calculating a threshold rule score ; mapping the feature vector to a probability value of each level using a lightweight classification model ; calculating a fusion score and performing pressing force level classification determination to obtain a pressing level of the current candidate event segment; the pressing level at least includes two or more levels of non-touch, light touch, medium press and deep press;
[0078] S70, de-bouncing output:
[0079] In a preset time window, majority voting or consistency judgment is performed on the determination results in the candidate event segment, and a final pressing level is output.
[0080] More specifically, in the present embodiment:
[0081] The present scheme designs a ring structure with notches around the microphone and sound transmission hole as a pressure sensing structure, and uses the passive acoustic channel changes caused by finger touch to realize multi-level (at least two levels) pressure sensing input. The pressure sensing input key and method are particularly suitable for electronic devices that have already installed noise reduction microphones or microphone arrays.
[0082] The scheme proposed in the present scheme uses passive acoustic multi-level pressure sensing input coupled with a finger to determine the touch (pressing) level based on frequency band energy, energy ratio, spectral peak / spectral notch amplitude and bandwidth, and cross-spectral coherence, without adding pressure / capacitance sensors, combining microphone passive acquisition and feature processing methods; Through coaxial multi-stage ring structure and optional differential / coherent noise reduction of double microphones, the noise and false touch suppression capability is improved to meet the design requirements of electronic devices in terms of simple appearance, low cost and interaction.
[0083] The present scheme mainly sets a ring pressure sensing structure with notches around the microphone sound transmission hole of the device shell, and the microphone and signal acquisition and processing unit are still arranged inside the shell. When the device needs to set multiple keys or multiple key functions, different key positions can be formed by differentiating the design of the cross-sectional shape, height, number of stages or notch layout of the ring structure; The user presses to different degrees to cause changes in the volume of the micro sound cavity and the geometry of the acoustic channel, thereby causing distinguishable changes in the acoustic signal features collected by the microphone. To improve the recognition stability in a complex noise environment, the target microphone at the key and another microphone (such as a noise reduction microphone or a microphone array unit) inside the device can work together to suppress far-field noise through differential spectrum or cross-spectral coherence methods. The key does not need to actively sound, uses environmental sound as passive excitation, extracts features such as frequency band energy, energy ratio, spectral peak / spectral notch amplitude and bandwidth, and realizes at least two levels of pressing force grading combined with threshold rules or lightweight models; The pressing levels can correspond to states such as no pressing, light pressing, medium pressing, and deep pressing.
[0084] The pressure sensing input key specifically includes a microphone, a sound transmission hole, a pressure sensing structure, and a processing unit; wherein the microphone, the sound transmission hole, and the processing unit can be directly applied to the original structure of the electronic device. The pressure sensing structure is a ring structure arranged around the sound transmission hole on the outer surface of the electronic device, and at least one notch is formed on the ring structure to form a neck connecting the sound transmission hole and the atmosphere. According to the number of functional levels required, the ring structure can be designed as a single-stage ring (with two levels of pressure sensing) or a multi-stage ring (with more than two levels of pressure sensing). When there are more than two levels of pressure sensing, The height of the innermost ring is h1, the height of the outermost ring is hN, and the height ratio between adjacent rings is h1 / h2 / h3 / … / hN-1 / hN, where h1 / hN>1, h1 / h2>1, h2 / h3>1, …, hN-1 / hN>1. wherein, h1 is the height of the ring located on the inner side of the adjacent rings, hN is the height of the ring located on the outer side of the adjacent rings. The cross-sectional shape of the ring structure can be designed as a circle, a circle-like shape, or a polygon to form differentiation between the keys. The notch on the ring structure can be provided with 1-6 notches, and each level of the multi-level ring is provided with a notch. If there are multiple notches, they can be distributed at equal angles or unequal angles. The pressure sensing structure is configured to significantly change the effective opening state and / or the cavity state of the acoustic channel when the user presses the microphone sound transmission hole peripheral area with different pressing forces, so that the sound signal reaching the microphone produces a distinguishable change with the pressing force. The processing unit is configured to collect the sound signal output by the microphone during the key operation, and to grade the pressing force of the key operation according to the sound signal features for characterizing the acoustic channel state, so as to distinguish at least two different pressing forces.
[0085] The specific technical principle of the pressure sensing input key is that the pressure sensing input key is provided with a notched ring pressure sensing structure around the microphone and the sound transmission hole. When the finger contacts / presses the area, the finger surface, the ring structure, and the device shell together form a micro acoustic cavity with a variable volume. According to the different pressing forces, the notches and gaps form different sound transmission paths. The structure is equivalent to an acoustic resonance network that changes with the pressing force, which can produce different responses such as "full-frequency shielding" or "band-pass filtering" to the incident sound, thereby mapping different pressing forces into distinguishable acoustic features.
[0086] 1) Full plugging (corresponding to deep pressing):
[0087] When the finger presses the microphone sound transmission hole completely, the external sound wave incident path is cut off, and the external sound signal received by the microphone is significantly attenuated in the full frequency band, which is in an approximate "mute / shielding" state.
[0088] 2) Non-full plugging (formed by notches):
[0089] When the finger contacts but does not completely plug the sound transmission hole, the cavity surrounded by the finger-ring structure-housing and the notches / gaps form a "cavity-neck" system, which produces a band-pass filtering effect on the sound wave; its equivalent center frequency approximately satisfies:
[0090] ;
[0091] wherein, c is the speed of sound, A is the equivalent neck (notch / gap) cross-sectional area, V is the equivalent cavity volume, Equivalent neck length. The pressing force degree change will cause the linkage change of 、 、 , resulting in predictable displacement and change of the passband center frequency and bandwidth (or quality factor ).
[0092] 3) Mapping relationship of force degree-acoustic characteristics
[0093] Different pressing force degrees correspond to different equivalent cavities and channel geometries, which are manifested as: monotonic upward shift (or multi-peak merging / migration) of the bandpass peak position (center frequency) , repeatable change of bandwidth / quality factor , and step difference of passband energy relative ratio. By threshold / template judgment on the spectral characteristics of the microphone output, the states of light touch, step-by-step pressing, and full occlusion can be stably distinguished, realizing multi-level pressure input without adding independent mechanical buttons.
[0094] The pressure input button is set as a first-order ring structure (single ring with a notch), as shown in Figure 1 :
[0095] When lightly touching, the finger closes the single ring surface to form a closed cavity, and the sound wave is coupled to the microphone through the notch on the ring, showing a bandpass response: the frequency components within the passband pass through, and the frequency bands outside the passband are suppressed; if the pressure is continuously increased to full occlusion, the "shielding" state is entered. Therefore, the first-order structure corresponds to two levels of pressure: light touch (bandpass state) and heavy pressing (shielding state).
[0096] The pressure input button is set as a -order ring structure (concentric multi-rings, the height of the inner ring is gradually lower than that of the outer ring), as shown in Figure 2 、 3 :
[0097] When using a 2-order or more concentric multi-ring structure and the inner ring is gradually lower: light touch: the outer ring is first closed by the finger, and the sound wave is directly coupled to the microphone through the notch of the outer ring, forming the first bandpass state; continue to press: the inner ring is sequentially closed, and the sound wave needs to be transmitted through the notch of the outer / inner ring and the annular gap between the concentric rings, the equivalent is reduced but is significantly lengthened, making and evolve in a stepwise manner, forming distinguishable acoustic fingerprints of different bandpass center frequencies / bandwidths; until the heaviest pressing, the sound transmission hole is completely occluded, entering the shielding state. Therefore, -order structure can realize -level pressure (including the final full occlusion state).
[0098] The microphone in the pressure-sensitive input button can be composed of multiple microphones, such as two microphones, one of which is a main microphone with a pressure-sensitive structure around its periphery, and the other of which is a reference microphone for collecting ambient sound. The multiple microphones can all directly use the original microphones of the device, such as noise reduction microphones or microphone array units.
[0099] The pressure-sensitive input method of the pressure-sensitive input button includes the following steps:
[0100] S10, signal acquisition and candidate event detection
[0101] The passive acoustic signal of the target microphone arranged in the device shell near the button area is collected, and the sampling rate is preferably 16 kHz to 48 kHz. The collected time-domain signal is pre-emphasized and band-pass filtered, and the band-pass frequency band is, for example, 500 Hz to 8 kHz, to suppress direct current drift and ultrahigh frequency noise.
[0102] In a sliding window with a frame length (for example, it can be selected within 10 to 100 ms, preferably within 10 to 30 ms) and a step length (for example, it can be selected within 5 to 15 ms, preferably within 5 to 10 ms), the normalized increment of short-time energy and passband energy is calculated, and when and / or exceeds a preset multiple of the background energy, it is marked as a candidate pressing event interval, and adjacent frames are merged to obtain an event segment.
[0103] Among them, the "background energy" can be estimated by the output of the reference microphone or the long-time average energy.
[0104] In addition, when the ambient sound is insufficient, the device can be enabled to play a preset excitation sound to ensure the signal-to-noise ratio and repeatability.
[0105] S20, pre-processing
[0106] The candidate event segment is normalized to remove the direct current component and the overall gain difference; if necessary, it can be framed with a Hanning window or other window functions to reduce spectral leakage. For the case where there is a second microphone (reference microphone), the two signals are clock-synchronized and gain-calibrated to make them dimensionally consistent.
[0107] S30, spectral estimation
[0108] The Welch method or other power spectral density estimation method is used to obtain the power spectral density of the pre-processed candidate event segment. Within the preset frequency band of interest (for example, 500 Hz to 8 kHz), the power spectral density Smoothing or interpolation is performed to improve the stability of the features.
[0109] S40, feature extraction
[0110] In this solution, the feature extraction step is based on the physical mechanism of the acoustic resonance network, and multi-dimensional features are extracted from the power spectrum and time-domain signal to form the feature vector for the pressing level determination. Specifically, but not limited to, the following features are included:
[0111] 1. Band energy and energy ratio features
[0112] In the power spectrum , one or more target frequency bands (e.g. narrow bands covering the natural frequency of the ring structure) are selected, and the energy of each frequency band is calculated:
[0113] .
[0114] The total band energy is calculated:
[0115] .
[0116] The energy ratio is obtained:
[0117] .
[0118] Different pressing forces change the equivalent cavity volume and channel cross-section, resulting in and monotonic or step changes between levels.
[0119] 2. Resonance center frequency and quality factor features
[0120] Search for the power spectrum peak in the target frequency band to obtain the corresponding resonance center frequency ;
[0121] Estimate the bandwidth in the half-height full-width (FWHM) way , determine the upper and lower limit frequencies , at which the power spectrum decreases to a certain proportion (e.g. -3dB~-6 dB) of the peak ;
[0122] Thus, the quality factor is obtained.
[0123] As the pressing force changes from no pressing, light pressing to deep pressing, the position of and all show a repeatable change trend.
[0124] 3. Time-domain transient features
[0125] Compute the rise time of the energy (transient) at the start of the press event (the time from when the energy exceeds the background threshold to when it reaches steady state);
[0126] Select time-domain features such as the slope of the rising edge around the event peak, the duration of the peak, etc.
[0127] Under different press habits and structural combinations, The energy peak shape can assist in distinguishing between light presses, medium presses, and fast deep presses.
[0128] 4. Dual-microphone features (optional when equipped with dual microphones)
[0129] In the case of a device equipped with a second microphone, the target microphone signal and the reference microphone signal Compute the difference spectrum and / or cross-spectrum coherence to suppress the influence of far-field noise:
[0130] Difference spectrum: Or calculate the amplitude difference after aligning the gains;
[0131] Cross-spectrum , self-spectrum , , calculate the cross-spectrum coherence:
[0132] .
[0133] Select the average coherence within a certain frequency band as one of the criteria for whether the far-field noise is sufficiently canceled out and whether the press event is a near-field disturbance.
[0134] Integrate the above features to construct a feature vector:
[0135] ;
[0136] At least including one or more of the center frequency , quality factor , band energy , energy ratio and transient rise time .
[0137] S50, dual-microphone difference / coherence enhancement (optional when equipped with dual microphones)
[0138] For embodiments equipped with a second microphone, using the second microphone as a reference microphone, the dual-microphone features extracted in step S40, such as the dual-microphone differential spectrum and / or cross-spectral coherence, are enhanced or weighted based on the differential relationship and / or cross-spectral coherence relationship between the reference microphone signal and the target microphone signal. This is done to suppress the influence of far-field noise on the features and highlight near-field features related to pressure. Specifically, spectral enhancement of the power spectral density of the target microphone can be performed in the spectral domain.
[0139] ;
[0140] in, The power spectral density of the target microphone signal. The power spectral density of the reference microphone signal or the far-field noise power spectral density estimated from it, The weighting coefficients are used; or the spectrum or features are weighted according to the cross-spectral coherence to highlight the near-field pressing features.
[0141] S60, Fusion Decision
[0142] When insufficient ambient sound leads to a low signal-to-noise ratio, the device can be enabled to play a preset excitation sound to improve feature stability and repeatability.
[0143] The naming system for pressure levels in this manual is as follows: no touch (level 0), light touch (level 1), medium press (level 2), and deep press (level 3); the "level 1 / level 2 / level 3" mentioned below correspond to light touch / medium press / deep press, respectively.
[0144] In this scheme, the fusion decision step adopts a parallel fusion approach of "threshold rules + lightweight model", and is based on a multi-level threshold set. The current pressure level is determined by the maximum fusion score. (Based on the set of pressure levels.) For example, a set of feature templates or feature intervals is set for each level to obtain the score for each level based on threshold rules. and with feature vectors Input a lightweight classification model and output a probability vector .in:
[0145] 1) Threshold-side score :
[0146] For each pressure level According to its corresponding feature interval (e.g. , , (etc.), transforming the cases where each feature falls within the interval into local scores and performing weighted summation or normalized distance mapping to obtain a threshold score within the range of 0 to 1. . The larger, the more the current feature matches the physical template of the pressing level.
[0147] 2) Model-side probability :
[0148] A lightweight classification model (e.g., small-scale neural network, logistic regression, or decision tree, etc.) is used to map the feature vector to the probability value of each level . The model can be trained by the prototype calibration data, automatically learning , , , and the nonlinear relationship between the double-microphone features.
[0149] Considering both kinds of information, the fusion score of each level is calculated:
[0150] ;
[0151] where is a weight coefficient between 0 and 1, preferably between 0.3 and 0.7. In the single-microphone or far-field noise is small scene, a larger value can be taken to highlight the stability based on the physical template; in the double-microphone cooperation, when the reference microphone mutual spectrum coherence decreases, and the environmental noise is unstable, appropriately reduce to enhance the adaptability of the model output in complex noise.
[0152] According to the size of , the pressing level is determined, and the following three groups of judgment conditions are processed:
[0153] 1. Multi-level threshold and "no press" threshold (first group of conditions)
[0154] Let the global optimal level be:
[0155] .
[0156] If is lower than the preset overall threshold , or the corresponding threshold score is lower than the minimum physical credible threshold , and the overall energy feature , does not exceed the trigger threshold, it is determined that the current state is "no press", and the pressing level is output, so as to avoid that the environmental noise or non-contact disturbance is misjudged as an effective press.
[0157] wherein the overall threshold and the physical credible threshold The preset threshold is preferably obtained by prototype calibration: collect samples without pressing under various environmental noise conditions, and statistically fuse the scores The threshold score is The distribution is selected according to the target touch rate on the verification set, and the quantile threshold or the ROC (receiver operating characteristic curve) optimal point is selected as the threshold; during operation, the threshold can be adaptively corrected in a small range based on the background energy estimated by the reference microphone.
[0158] 2, hysteresis retention and de-bouncing processing (second group of conditions)
[0159] If the current optimal level is different from the output level of the previous moment , but the difference between their fusion scores is: less than the preset hysteresis threshold , then enter hysteresis retention: continue to maintain the previous level within the current stable time window (for example, 80~150 ms) Only when is greater than and the consistent determination of most frames in the window supports , update to the new pressing level.
[0160] Through the above hysteresis and majority voting mechanism, de-bouncing output is realized to avoid frequent level jumps caused by transient noise or slight finger shaking.
[0161] 3, adaptive weighting based on reference microphone coherence correlation (third group of conditions, optional when dual microphones are available)
[0162] In embodiments with a second microphone, if the cross-spectral coherence average is higher than the coherence threshold , it means that the observations of the main microphone and the reference microphone on the environmental noise are consistent, and the dual-microphone difference / coherence filtering effect is reliable, at this time can be appropriately increased, so that the fusion is more dependent on the on the physical template side to obtain stable multi-level determination;
[0163] On the contrary, when is lower than or the reference microphone signal is detected to be abnormal, it means that the reliability of the dual-microphone cooperative information is reduced, at this time can be reduced or temporarily switched to single-microphone / model determination dominated by to avoid false difference or coherence estimates masking the real pressing features.
[0164] This condition ensures that the S60 fusion determination has a clear corresponding relationship with the dual-microphone cooperation, cross-spectral coherence noise suppression mechanism used in this scheme.
[0165] S70, de-bounce output and function mapping
[0166] In a stable time window (e.g. 80~150 ms), majority voting, hysteresis or consistency judgment is performed on the press level determination results of consecutive frames. Only when a certain level appears no less than a preset proportion (e.g. more than 2 / 3 of the total number of frames) in the window, the level is output as the final press level. The total system latency is no more than 150 ms. The final press level can be mapped to volume adjustment, play / pause, mode switching, etc.
[0167] Application Example 1: mobile phone, single microphone, first-order ring
[0168] Structure: a microphone sound transmission hole 100 is arranged on the outer surface of the electronic device 200, the hole diameter =0.6~0.9 mm (preferably 0.8±0.05 mm). The hole is formed with a first-order ring structure 101 with a notch: outer diameter =4.0~5.5 mm, ring top width =0.6~1.0 mm, ring height =0.35~0.55 mm (preferably 0.45 mm); the number of coupling slit openings =1~2, single opening angle =20°~40°, slit width =0.25~0.45 mm. The sound transmission hole 100 can be covered with a hydrophobic acoustic film. To compensate for tolerances, the outer hole diameter of the electronic device 200 (coaxial with the first-order ring structure 101) is enlarged by 0.10~0.20 mm compared to the PCB through hole. The above-mentioned shell outer diameter refers to the diameter (or equivalent diameter) of the sound transmission hole 100 on the surface of the electronic device shell 203, which is directly opened on the outer shell and located inside the boundary of the ring structure, which can be referred to Figures 2-4 ; the value is used to define the coupling interface with the finger / hydrophobic film. The above-mentioned PCB through hole refers to the equivalent diameter of the through hole or acoustic guide opening on the main board (PCB) that communicates with the microphone acoustic inlet; together with the shell outer hole, it constitutes a series / parallel impedance element of the sound transmission channel.
[0169] Determination:
[0170] ① Light touch 0.5~1 N (can be referred to in Figure 3 (b), the finger pulp locally closes the first-order ring structure 101): the system jumps from the release state, and a step and a resonance peak appear (the performance test results are as shown in Figure 5 ); record , and passband energy ratio.
[0171] ② Heavy press ≥2~3 N (can be referred to in Figure 3Middle (d): Near full enclosure, significant attenuation of broadband into shielded state. Stable window 80~150 ms, total latency <150 ms.
[0172] I. Key detection method configuration
[0173] Based on the above structure, the key detection method described in S10~S70 is used to determine the key operation, and the specific parameter configuration is as follows:
[0174] 1. Sampling and preprocessing (corresponding to S10~S20):
[0175] The microphone sampling rate is 48 kHz, pre-emphasis and 500 Hz~8 kHz band-pass filtering are performed on the collected passive acoustic signals; the short-time energy and passband energy increment are calculated in the sliding window with frame length 20 ms and step length 10 ms, which are used for candidate event detection.
[0176] 2. Power spectrum estimation and feature extraction (corresponding to S30~S40):
[0177] Welch method is used for power spectrum estimation on the candidate event interval, and the FFT point number is 512, and the power spectrum density in the range of 0~24 kHz is obtained . In the target frequency band of 500 Hz~8 kHz, at least one of the center frequency , the quality factor (counted by half-height full width FWHM), the target frequency band energy , , the energy ratio and the transient rise time is extracted to form a feature vector .
[0178] 3. Fusion determination and debouncing (corresponding to S60~S70):
[0179] Taking the feature vector as input, the threshold rule score and the lightweight classification model probability of each pressing level (not touch, light touch, medium press, deep press) are calculated, and the fusion score is calculated according to , . The determination results of continuous multiple frames in the stable window of 80~150 ms are majority voted or hysteresis maintained, and the final pressing level is output.
[0180] II. Spectral test and determination example of first-order structure light touch state (based on Figure 5 )
[0181] The first-order ring structure described above was tested in the "unpressed" and "lightly pressed" states, and the test conditions are as follows:
[0182] 1. Test environment and posture:
[0183] The terminal is in a quiet indoor environment, and the device is played at a constant volume.
[0184] Non-pressed state: the user's finger is away from the key area, and only the environmental and system background noise is collected;
[0185] Light touch state: the user lightly touches the top of the ring surface structure with normal operating force, forming a closed microcavity, but no obvious downward deformation is generated.
[0186] 2. Test results (see Figure 5 ):
[0187] Figure 5 The solid line B corresponds to the non-pressed state, and the dashed line A corresponds to the lightly pressed state. It can be seen that near about 4.00 kHz, the dashed line A appears a narrow-band amplitude lift, which produces a step difference of about 6~10 dB relative to the solid line B, forming a characteristic resonance peak in the lightly pressed state of the first-order structure.
[0188] 3. Feature extraction and numerical examples (corresponding to S40):
[0189] The 3.5~4.5 kHz band is targeted for feature extraction of the spectral shapes of the two states shown in Figure 5 , and the typical results are as follows (the numerical values are a representative sample):
[0190] Light touch state (curve A):
[0191] Center frequency ≈4.00 kHz;
[0192] Quality factor ≈6~8, sharp spectral peak;
[0193] Energy ratio in the band ≈0.45~0.55;
[0194] The corresponding transient rise time is moderate.
[0195] Non-pressed state (curve B): no narrow-band peak appears in the 3.5~4.5 kHz interval, not obviously defined; energy ratio ≈0.20~0.30, significantly lower than the lightly pressed state.
[0196] Therefore, the 4 kHz step peak generated by the first-order ring structure in the tapping state can be used as the core acoustic feature to distinguish between the "unpressed" and "tapping" levels.
[0197] Fusion determination process example (corresponding to S60):
[0198] The above feature vector The input fusion determination module calculates the threshold score for the "unpressed" and "tapping" levels respectively And the model probability A set of typical results are obtained, for example:
[0199] Tapping: ≈0.80, ≈0.78, ≈0.79;
[0200] Unpressed: ≈0.20, ≈0.18, ≈0.19.
[0201] Therefore, the maximum fusion score corresponds to the "tapping" level, which is significantly higher than other levels such as "unpressed".
[0202] 4. De-bouncing and final output (corresponding to S70):
[0203] In the stable time window of about 100 ms, the fusion determination results of most frames are "tapping", and the energy feature continuously meets the tapping threshold. Therefore, the final output of this embodiment is the tapping level, realizing stable two-level pressure sensing (unpressed / tapping) under the first-order ring structure.
[0204] Application example 2: TWS / neck hanging type, single or double microphone, second-order ring
[0205] Structure: a sound transmission hole 100 is formed on the outer surface of the electronic device 200, =0.4~0.7 mm (preferably 0.6±0.05 mm). The second-order pressure sensing button 201 is arranged around the hole, which is composed of a second-order outer ring structure 201b and a second-order inner ring structure 201a: outer / inner ring outer diameter =3.2~3.8 mm / 2.2~2.8 mm, ring top width =0.5~0.8 mm / 0.4~0.6 mm, ring height =0.28~0.40 mm / 0.15~0.30 mm; number of coupling seams =2 (central line angle 180°) or 3 (equiangular 120°), single opening angle =15°~30°, seam =0.20~0.35 mm; preferably low flow resistance thin acoustic membrane to compensate for the acoustic impedance of small aperture.
[0206] Decision: only the second-order outer ring structure 201b is closed first (a) Figure 3 (b) → step, resonance established, judge level (light touch);
[0207] Continue to press to close the second-order inner ring structure 201a (c) Figure 3 (b) → cavity coupling changes, , Step + displacement, judge level two (middle press);
[0208] Almost full (d) Figure 3 (b) → wideband significant attenuation, judge level three (deep press).
[0209] Application Example 3: Watch / bracelet, side-mounted, two / three-order ring
[0210] Structure: side-mounted sound-transparent hole 100, =0.5~0.8 mm; surrounding two / three-order notched ring (corresponding to the coaxial sequence of the second-order outer ring structure 201b → the second-order inner ring structure 201a → the third ring), outer / inner ( / third) ring outer diameter =3.5~4.5 / 2.5~3.2( / 1.8~2.4) mm; ring height =0.30~0.45 / 0.20~0.35( / 0.12–0.25)mm; =2, =20°~35°, =0.25~0.40 mm. Side hole with membrane frame, ring top and membrane frame gap 0.10~0.20mm; outer protective electronic device shell 203. Decision: lateral pressing successively closes the outer / inner ( / third) ring to form multiple levels; the heaviest press nearly full to enter the shielding state.
[0211] Application Example 4: Multi-key and blind touch, same machine multi-key
[0212] Configure multiple keys on the same electronic device 200: single ring uses first-order ring structure 101 (see Figure 1 ), double ring uses two-order pressure-sensitive key 201, including second-order inner ring structure 201a, second-order outer ring structure 201b (see Figure 2 / Figure 3 (a). Form structural differences (can refer to , , , and cross-sectional shape among the keys Figure 4), the different textures and height differences can be perceived by scanning the fingerpads, realizing blind touch positioning.
[0213] Function mapping: mapping the levels ( Figure 3 b), Figure 3 c), Figure 3 d) with key combinations, such as corresponding volume increase / decrease, play / pause, Bluetooth switching, etc.; the appearance surface type can remain the same, and only the ring height difference / opening layout is used to express the difference. The single ring step response is referred to Figure 5 .
[0214] Application Example 5: Dual-microphone cooperation, main microphone + reference microphone
[0215] The main microphone adopts a ring-shaped microphone with a notch (first-order ring structure 101 or second-order pressure-sensitive button 201, see Figure 1 , Figure 2 ); the reference microphone is the original noise reduction microphone or array unit of the whole machine. The two are first synchronized in time and calibrated in gain to ensure dimensional alignment.
[0216] When running, the reference microphone continuously gives the baseline of the "environmental noise floor", and the main microphone side is sensitive to the peak position and bandwidth changes caused by the fingertip coupling. In the target frequency band, a short-time spectrum is made, the main reference transfer is estimated and canceled, and the residual spectrum is obtained; in the residual, the , and passband energy ratio are tracked; the main / reference features are subtracted or compared and weighted by coherence; first triggered by change point detection, then divided into multiple levels (light touch / medium press / deep press) by multi-level threshold + hysteresis; when the coherence is insufficient or there is abnormal noise, the weight is reduced or converted to a single channel. On the one hand, it directly improves the estimation accuracy of , and passband energy ratio, and on the other hand, it subtracts or compares the two features. Finally, it is divided on the multi-level threshold. In this way, whether it is subway entry, wind hole or music playback, it is more difficult to mix light touch, medium press and deep press together, the overall recognition rate is higher, and there is no need to ask the user to increase the pressing force.
[0217] Application Example 6: Unequal Angle Distribution, Two / Three Order Rings
[0218] Taking the second order of Figure 3 (a) as the base type, the second order outer ring structure 201b / second order inner ring structure 201a (and the third ring) is offset by 30°~45° on the circumference, or the central angle of each coupling gap is allocated unequally to meet the avoidance of antennas, battery ribs, and waterproof channels, and to preferentially trigger the first closure in the "main operation direction".
[0219] The structure design can produce the following beneficial effects: firstly, the closing order of the outer-to-inner (to the third) ring is more controllable in the target direction, and the finger rolling can "step by step" obtain the first (light touch), second (medium press), and third (deep press) pressing degrees; secondly, after breaking the complete symmetry, the easily superimposed resonance is broken, the peak position is more stable, and the energy ratio is more neat.
[0220] The above description of the embodiments is to facilitate those skilled in the art to understand and use the invention. Those skilled in the art can easily make various modifications to these embodiments, and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and improvements and modifications made by those skilled in the art without departing from the scope of the present invention should be within the scope of protection of the present invention.
Claims
1. A microphone-based pressure-sensitive input button, characterized by The microphone, the sound transmission hole, the pressure-sensitive structure and the processing unit are included. The microphone is integrated in the electronic device and is used to collect sound signals entering the inside of the electronic device through the shell of the electronic device. The sound transmission hole is arranged on the shell of the electronic device and is in communication with the microphone to form an acoustic channel. The pressure-sensitive structure is a ring structure surrounding the sound transmission hole and is provided with at least one gap to form a neck part connecting the sound transmission hole and the atmosphere. The processing unit is integrated in the electronic device and is in communication connection with the microphone, is used to acquire and process the sound signals collected by the microphone, and then performs a graded judgment on the pressing force of the pressure-sensitive input button according to the characteristics of the sound signals. The pressure-sensitive structure, the user's finger and the shell form a micro acoustic cavity with variable acoustic resonance characteristics together, the gap constitutes an equivalent neck part connecting the micro acoustic cavity and the atmosphere, so that the change of the pressing force causes the change of the equivalent resonance characteristics of the acoustic channel, and the processing unit grades the pressing force based on the change of the equivalent resonance characteristics.
2. The pressure input key based on a microphone according to claim 1, wherein, The pressure-sensitive structure is a single-stage ring or a multi-stage ring arranged coaxially. When the pressure sensing structure is a multi-stage ring, the height of the innermost stage ring to the outermost stage ring increases in turn and the height ratio between adjacent rings satisfies wherein, is the height of the ring located on the inner side in the adjacent rings, is the height of the ring located on the outer side in the adjacent rings.
3. The pressure input key based on a microphone according to claim 1, wherein The cross-sectional shape of the ring structure is circular, quasi-circular or polygonal. The gap is provided with 1-6 gaps.
4. A microphone-based pressure input method, comprising: The pressure-sensitive input button as claimed in any one of claims 1-3 is adopted. The pressure-sensitive input method comprises the following steps. S10, signal acquisition and candidate event detection: The signal collected by the microphone is acquired, and the pressing event interval is determined based on the signal according to the preset event detection criterion, and adjacent or overlapping pressing event intervals are merged to obtain a candidate event segment. S20, preprocessing of the candidate event segment: The candidate event segment obtained in step S10 is normalized and / or aligned. S30, spectral estimation of the candidate event segment: The power spectrum density of the candidate event segment obtained in step S20 is acquired by a power spectrum density estimation method and performs smoothing and / or interpolation on the preset frequency band of interest S40, feature extraction: The power spectrum density obtained from step S30 and the time domain signal corresponding to the candidate event segment are extracted to obtain multi-dimensional features, and a feature vector for the pressing force grading determination is constructed ; wherein the multi-dimensional features at least include one of a center frequency , a quality factor , a frequency band energy , an energy ratio , and a transient rise time S60, fusion judgment: for each pressing level builds respective feature intervals, and calculates threshold rule scores ; uses a lightweight classification model to map the feature vector into probability values for each level ; calculates a fusion score and makes a grading determination of the pressing force, obtaining the pressing level of the current candidate event segment S70, de-bounce output: In a preset time window, the judgment results in the candidate event segment are executed by majority voting or consistency judgment, and the final pressing level is output.
5. The method according to claim 4, wherein, In step S10, the pre-processing includes pre-emphasis and band-pass filtering, and the signal collected by the microphone is processed by a sliding window for frame processing. Wherein, the frame length of the sliding window , the step length .
6. The method of claim 4, wherein the method further comprises: In step S20, after normalization processing, windowing and / or framing are further performed to reduce spectral leakage.
7. The method of claim 4, wherein the method further comprises: In step S30, the power spectral density estimation method is one or a combination of non-parametric method, parametric method and time-frequency analysis method.
8. The method of claim 4, wherein, In step S40, the energy ratio is the ratio of the band energy to the full-band energy . the quality factor the center frequency the ratio of the bandwidth , wherein the bandwidth is estimated by the half-height width method The transient rise time is the length of time for the energy to exceed the background threshold to reach steady state.
9. The method of claim 4, wherein the method further comprises: In step S60, the lightweight classification model is a small-scale neural network, logistic regression or decision tree. the fusion score is calculated by the following formula: wherein, is a weighting factor between 0 and 1.
10. The method of claim 4, wherein the method further comprises: When the electronic device is provided with a plurality of microphones: Step S20 further comprises clock synchronization and gain calibration of the time domain signals collected by different microphones and pre-processed; Step S40 also comprises computing the bispectrum and / or the cross-spectral coherence on the time-domain signals captured by the different microphones and pre-processed as feature vectors one of the elements of the set Between steps S40 and S60, step S50, differential / coherent enhancement of the microphones, is further included, that is, the double-microphone differential spectrum and / or the double-microphone feature of the inter-spectrum coherence extracted in step S40 are executed by enhancement or weighting processing to suppress far-field noise and highlight near-field pressing features.
Citation Information
Patent Citations
Method and system for achieving operation of mobile terminal according to touch signals and mobile terminal
CN105302373A
Touch feedback method and device and electronic device
CN109189212A
Box body and signal transmission method
CN121334539A
Combined microphone seal and air pressure relief valve that maintains a water tight seal
US9408009B1
Cited By
A microphone-based human-computer interaction method, device, medium and system
CN122431537A