Intelligent earphone multi-mode interaction system and method
By constructing a multimodal interaction system and utilizing multi-dimensional collection and analysis of earphone surface texture, hand micro-movements, and voice breath, the problem of high false trigger rate of smart earphones in complex environments has been solved, achieving high-precision user operation intent determination and dynamic tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BOSHI (HEILONGJIANG) ARTIFICIAL INTELLIGENCE IND DEVELOPMENT CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing smart headphones have a high rate of false triggering due to misalignment of the wearing position in complex environments, making it difficult to accurately identify the user's true operating intentions, especially due to the lack of joint constraints between multimodal signals.
By collecting texture clusters on the earphone surface from multiple dimensions, and combining hand micro-movements and voice breath, a multimodal interaction system is constructed. Using technologies such as multi-channel touch sensors, differential motion analysis, time-weighted functions, and short-time Fourier transform, an intent establishment path is generated, and conflict nodes between the main touch pulse and the short vibration accompanying pulse are identified.
It significantly improves the response accuracy and real-time performance of smart headphones in complex environments, reduces the error rate, enhances the adaptability to multimodal input, and improves the user experience.
Smart Images

Figure CN122054034A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile smart terminals, specifically to a smart earphone multimodal interaction system and method. Background Technology
[0002] In the collaborative use of smart headphones and mobile smart terminals, users typically issue input commands to the terminal via voice wake-up, the headphone touch panel, or slight head movements. However, in real-world applications, when users lightly touch the outside of the headphones to perform pause or play operations in subway or office settings, the touch area may be misinterpreted as a swipe gesture due to slight shifts in wearing position, thus triggering incorrect functions. Existing smart headphones mainly rely on a single touch pressure threshold or a fixed-direction swipe template to recognize input events, but they struggle to distinguish between genuine user commands and unintentional touches caused by wearing posture within milliseconds when there are sudden changes in touch area, insufficient touch path length, or simultaneous slight head movements at the moment of triggering. The lack of joint constraints between different modal signals leads to a significant increase in false trigger rates in crowded environments, mobile scenarios, and biased wearing states. Therefore, it is essential to design a multimodal interaction system and method for smart headphones that reliably determines user intent. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a multimodal interaction system and method for smart headphones, which has the advantage of reliably determining user intent and solves the problems mentioned in the background technology.
[0004] To achieve the aforementioned objective of reliably determining user intent, the present invention provides the following technical solution: a multimodal interaction method for smart headphones, comprising the following steps: Multi-dimensional acquisition of the contact texture clusters formed on the outer contact surface of the earphone during the pressing, sliding and light squeezing process is performed. The degree of contraction of the contact spots, the discontinuous path of micro-pressure advancement and the sudden characteristics of contact surface stress jump are extracted to form contact surface source trace fragments for depicting the initial interaction tendency. Interlacing and splicing the trace fragments on the touch surface within a specified time period, and incorporating the residual texture of the fingertips formed by the user under similar action conditions, the habitual rhythm of the slight adjustment of the hand, and the tail dragging of the voice breath, a multimodal coarse guidance layer that presents intention bias is constructed. Based on the multimodal coarse guidance layer, the mixed area of contact surface fragmentation path and short earthquake scatter points is sequentially split, the disturbance level difference and trend reversal point of each mode in the local window are extracted, and the mode sequence is rearranged according to the reversal relationship. The intention concentration of the screened modal sequences is re-compressed, and the discrete distribution of the contact surface stagnation amplitude difference, short vibration rebound and the attenuation factor of the speech tail are jointly folded to construct a modal condensation surface with multi-level intention depth. Under the action of modal condensation, the conflict node between the main touch pulse and the short vibration accompanying pulse is identified, and combined with the short-cycle trigger offset signal that appears in the input event receipt of the mobile terminal, the intent establishment path is generated.
[0005] Preferably, the process of forming a touch surface source trace fragment for depicting the initial interaction tendency is as follows: The displacement changes and contact area fluctuations of the contact clusters under different pressure intensities are collected by a multi-channel touch sensor. The differential motion analysis method is used to identify the short-term continuous sliding path and sudden pressure point of the contact point, and the continuous signal is divided into initial segment units by the time weighting function. The initial segment unit is encoded with feature vectors, including contraction amplitude, path curvature and stress peak, to form a contact source trace segment for depicting the initial interaction tendency.
[0006] Preferably, the process of interleaving and splicing contact surface source traces within a specified time period is as follows: The contact surface source trace fragments are segmented and indexed based on the collection timestamp to form a modal fragment set of time-division windows; Within each time-sharing window, the contact surface source trace fragments are arranged in a staggered sequence, and the temporal characteristics of different fragments overlap to form an interleaved combination; By applying sliding time frames and phase adjustment methods, minor deviations between segments are corrected and overlapping interference is eliminated, generating continuous multimodal fusion segments.
[0007] Preferably, the process of constructing a multimodal coarse-coarse guiding layer that presents an intentional bias is as follows: The residual texture at the fingertips is time-aligned with the multimodal fusion fragment to form preliminary tactile fusion features; Analyze the short-vibration patterns of the hand's micro-adjustment trajectory and the output of the inertial unit to extract the micro-amplitude rhythm and inertial wave characteristics; Short-time Fourier transform is performed on the lingering of the speech tail segment, and envelope features are extracted and mapped to the tactile and inertial feature space. By combining tactile, inertial, and speech features through a weighted fusion strategy, a multimodal coarse guidance layer with preliminary intention bias indications is formed.
[0008] Preferably, the process of sequentially separating the mixed region of contact surface fragmentation path and short-range earthquake points is as follows: The overlapping area between the contact surface and the inertial signal in the multimodal coarse guidance layer is locally windowed, and the perturbation gradient and peak value change within each window are calculated. Identify interference segments between different modes and remove abnormal or random signals using dynamic thresholds; The modal segments are reordered according to the perturbation level and local trend reversal points to form a modal sequence that has undergone interference removal and windowing.
[0009] Preferably, the process of rearranging the modal sequence according to the foldback relationship is as follows: Within each local window, calculate the instantaneous amplitude difference and peak-valley reversal point between the touch and inertial signals; Using a multidimensional sorting algorithm, perturbation differences and trend reversal points are mapped to the index space of the modal sequence; Time alignment and deviation compensation are performed on the turnaround nodes between different modes to generate a continuous, identifiable and time-aligned optimized mode sequence.
[0010] Preferably, the process of re-compressing the intended concentration of the screened modal sequences is as follows: Calculate the frequency and intensity weight of each modal segment in the modal sequence; Based on the local accumulated energy of modal segments, the modal sequence is hierarchically compressed to enhance the intention weight of high-intensity segments; A nonlinear normalization function is applied to the rectification results to map the intensity of different modes to a unified intention quantization scale, and the final mode sequence with weight labels and time indexes is output.
[0011] Preferably, the process of constructing a modal condensed finger surface with multi-level intentional depth is as follows: The final modal sequence is superimposed with local perturbation features through multiple channels to form an intention matrix; By using a weighted folding algorithm, the features of touch stagnation amplitude difference, short vibration rebound, and voice attenuation are fused into the intention matrix; Multi-layer convolutional mapping and time-series compression are applied to the intention matrix to form a modal condensation surface with multi-layer intention depth.
[0012] Preferably, the process of generating the intent to establish the path is as follows: By analyzing the peak values and time intervals of pulses in each layer of the modal condensation surface, conflict nodes between the main pulse and the accompanying pulse are identified. By combining the short-period trigger offset signal received by the mobile terminal, the time stamp of the conflicting node is dynamically corrected. Based on the conflict nodes and the corrected pulse sequence, the intended path is generated.
[0013] A multimodal interaction system for smart headphones, comprising: Contact surface acquisition module: performs multi-dimensional acquisition of the contact point texture formed during the pressing, sliding and light squeezing processes, and extracts sudden features to form contact surface source trace fragments; Modal mapping module: interleaving and splicing contact surface source trace fragments within a specified time period to construct a multimodal coarse guidance layer with intention bias; Sequence decomposition module: It breaks down the mixed area of contact surface fragmentation path and short-shock scattered points, extracts the disturbance level difference and trend reversal point, and rearranges the modal sequence; Intent Condensation Module: The intent concentration of the selected modal sequences is compressed, and the touch, short vibration and voice features are folded together to form a multi-level intent depth modal condensation surface; Path determination module: Under the action of modal condensation, it identifies the conflict node between the main touch pulse and the short vibration accompanying pulse, and combines the short period trigger offset signal to generate the intended path.
[0014] Compared with the prior art, the present invention provides a multimodal interaction system and method for smart headphones, which has the following beneficial effects: This invention achieves high-precision determination and dynamic tracking of the user's true operational intent through multimodal perception, fusion, and hierarchical processing of the earphone touch surface, hand micro-movements, and voice tail segments. By capturing the contraction degree of touch texture clusters, micro-pressure propulsion paths, and sudden stress changes, a fine segment of the initial interaction tendency is formed. Combined with temporal interleaving splicing technology, continuous multimodal fusion information is generated, enabling the system to take into account the temporal relationship and local perturbation characteristics of touch, inertia, and voice signals, thereby effectively reducing the probability of misjudgment caused by random jitter or operational interference. Using local windowing splitting, perturbation level extraction, and foldback sorting techniques, the modal information is structured and optimized, strengthening the intent weight of key operation pulses while maintaining the temporal continuity and recognizability of the sequence. Through multi-channel folding and multi-layer convolutional mapping, a multi-level intention depth modal condensation surface is formed, achieving accurate identification of conflict nodes between the main pulse and accompanying pulses. Combined with the short-cycle trigger offset signal of the mobile terminal for dynamic correction, a stable and reliable intent establishment path is generated. It not only significantly improves the accuracy and real-time performance of smart headphones in responding to complex and varied user interactions, but also enhances the system's adaptability to multimodal inputs such as micro-motions, light touches, swipes, and voice operations, thereby improving the user experience, reducing the error rate, and enhancing the intelligence and usability of mobile smart terminals in complex interactive scenarios. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the method of the present invention; Figure 2 This is a schematic diagram of the structure of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1: Please refer to Figure 1 As shown in the figure, a multimodal interaction method for smart headphones in an embodiment of the present invention includes the following steps: S1: Multi-dimensional acquisition of the contact texture clusters formed on the outer contact surface of the earphone during the pressing, sliding and light squeezing processes, extracting the degree of contraction of the contact spots, the discontinuous path of micro-pressure advancement and the sudden characteristics of contact surface stress jumps, forming contact surface source trace fragments used to depict the initial interaction tendency.
[0018] The process in S1 for forming the touch source trace fragment used to depict the initial interaction tendency is as follows: The displacement changes and contact area fluctuations of the touch clusters under different pressure intensities are collected by a multi-channel touch sensor. A multi-channel touch sensor composed of capacitive and micro-pressure arrays is deployed in the touch area on the outside of the earphone. It can simultaneously detect the spatial position signals and local pressure distribution changes of multiple touch points within a millisecond sampling period. When the user applies pressure, swiping or squeezing action to the surface of the earphone, the touch clusters will show micro-displacement trajectories and expansion or contraction of the contact area as the pressure intensity changes. These raw touch data are collected in real time and multi-dimensional parameters including touch point coordinates, contact area estimates, instantaneous pressure amplitude and pressure duration are recorded. In order to avoid instantaneous noise from interfering with the data, a sliding window filtering strategy is used to smooth the output of each channel and the upper and lower edges of the pressure change curve are detected to ensure that the segment generation process can accurately reflect the real user touch behavior, rather than occasional capacitance fluctuations or environmental disturbances. Differential motion analysis is used to identify short-term continuous sliding paths and sudden pressure points of touch points, and the continuous signal is divided into initial segment units by a time weighting function. Differential motion analysis is performed on the continuous displacement and pressure changes of touch point clusters on the time axis to extract minute sliding behaviors and sudden pressure point events. The analysis method calculates displacement data to identify acceleration change points, sliding direction change points, and micro-segment trajectories after stagnation and then advancement in the path. Local peak points of pressure signals are extracted to determine the start and end positions of sudden pressure actions. The detected feature events are segmented and sliced according to a time weighting function, which is set according to the natural rhythm of touch actions. This results in long-term slow sliding being divided into multiple effective sub-segments, while short pressure point events are retained as independent segment units. This slicing method avoids the problem of blurred action boundaries caused by traditional fixed window methods, making each initial segment unit more accurately correspond to the actual touch action structure. The initial segment units are encoded with feature vectors, including contraction amplitude, path curvature, and stress peak, to form touch surface source trace segments that depict the initial interaction tendency. For the initial segment units that have been sliced, their representative features are extracted one by one and encoded into feature vectors to form touch surface source trace segments for further multimodal fusion. In this process, the spatial deformation characteristics of the segment on the touch surface are calculated, including the contraction amplitude of the contact area, the peak position of the pressure waveform, and the local minimum pressure range, to reflect the changes in the user's finger force application. The sliding path is modeled using curvature analysis methods to calculate the trajectory's turning point, rate of curvature change, and directional offset to reveal the subtle motion trend of the segment. The stress characteristics of the segment are evaluated by the difference between the pressure peak and the baseline pressure, and a feature vector of a unified dimension is generated by combining the time span. All vectors are normalized and then organized into touch surface source trace segments in a specific format.
[0019] S2: Interweave and splice the trace fragments on the touch surface within a specified time period, and introduce the residual texture of the fingertips formed by the user under similar action conditions, the habitual rhythm of the slight adjustment of the hand, and the trailing sluggishness of the voice breath to construct a multimodal coarse guidance layer that presents the intention bias.
[0020] The process of interleaving and splicing contact surface source traces within a specified time period in S2 is as follows: The touch surface trace segments are segmented and indexed based on the acquisition timestamps to form a modal segment set for time-division windows. The acquisition timestamps attached to the obtained touch surface trace segments are uniformly calibrated so that different sensing channels (touch, micro-pressure, micro-displacement) can be compared under the same time base. By setting the time window width (e.g., a micro-segment scale of 10-50ms), all touch surface trace segments are divided into the corresponding time-division windows according to the timestamps, thus forming a modal segment set arranged in chronological order. Each time-division window not only contains basic touch features such as displacement, contact area, and local pressure peaks, but also retains the local change trend inside the segment, ensuring that the segments can still reflect fine-grained motion structures before splicing. A buffer is added at the window boundary to allow the tail of the segment to be extended across the window, thereby reducing the segment breakage caused by window division. Within each time-sharing window, touch source trace segments are arranged in a staggered sequence, with the temporal characteristics of different segments overlapping each other to form an interleaved combination. After forming the modal segment set of the time-sharing window, the segments located in the same window are arranged in a staggered temporal sequence to reveal the correspondence between segments at a small time scale. According to the order of occurrence of the internal micro-events of the segment (such as the sliding start point, short pressure peak, and contact area contraction segment), the segments are arranged in a staggered manner so that the key events of different segments can partially overlap or cover each other. Through this interleaved arrangement, the system can construct richer local behavioral patterns than a single segment, so that touch actions of different durations and intensities can form a complementary relationship in the micro-temporal structure. For example, when one segment contains the beginning of a sliding action, and another segment captures the end of the same action with a light pressure, the interleaved arrangement can connect the two into a continuous action trajectory. By applying a sliding time frame and phase adjustment method, minor deviations between segments are corrected and overlapping interference is eliminated to generate continuous multimodal fusion segments. After completing the segment interleaving combination, a sliding time frame mechanism is introduced to detect the offset between adjacent segments on the time axis by scanning frame by frame. For the detected minor deviations (such as delay differences of 1 to 5 ms or pressure peak alignment errors), the phase adjustment method is used to perform correction, so that the key nodes of the segments (pressure peaks, direction reversal points, contact area abrupt change points) are aligned under a unified time reference. The overlapping interference that may occur in the segment interleaving area is processed. The signal strength, noise level and signal stability of multiple segments in the conflict area are calculated respectively. The most representative features are retained and random disturbances are suppressed according to the weighting strategy. By reconstructing the segments after sliding frame correction, a multimodal fusion segment with continuous temporal structure, highly consistent internal features and synchronous cross-modal information fusion is generated.
[0021] The process of constructing a multimodal coarse-grained guiding layer with intentional bias in S2 is as follows: The residual texture of the fingertip is time-aligned with the multimodal fusion fragment to form a preliminary tactile fusion feature. The residual texture of the fingertip is extracted from the high sampling rate capacitance grid of the headphone touch area. It includes stable features such as touch area, local capacitance gradient, and micro-pressure diffusion trace, which reflect the user's touch surface inertia in the previous action cycle. The residual texture is time-aligned with the multimodal fusion fragment obtained in the previous step. To achieve high-precision alignment, a micro-scale sliding comparison window is set. In each window, the time deviation between the texture peak position, texture diffusion direction and touch sub-event in the fusion fragment is calculated. During the alignment process, the local inertial feature weight of the texture is retained to ensure that the tactile fusion information is not completely flattened. The generated tactile fusion feature contains both the inertial residue of the previous action and the consistency with the fine-grained touch structure of the current action. We analyze the micro-adjustment trajectory of the hand and the short vibration pattern output by the inertial measurement unit to extract micro-amplitude rhythm and inertial wave characteristics. We acquire three-axis acceleration and gyroscope data from the inertial measurement unit inside the earphone, and use high-pass filtering to remove large-scale gesture changes, retaining the micro-adjustment trajectory generated by the user's hand during touch actions. These micro-adjustment trajectories are usually manifested as extremely low amplitude directional deviations, slight rotations, or compensation actions during the stabilization of fingertip posture. Based on the time-domain window analysis method, we extract the peak-valley changes, frequency reproducibility, and interval stability of adjacent micro-events in these micro-trajectories as micro-amplitude rhythm features. We perform wavelet packet decomposition on the short vibration pattern in the inertial measurement unit to capture the high-frequency micro-vibration signal accompanying the user's light touch or swipe actions, and calculate the amplitude attenuation rate, source position change trend, and inertial coupling direction. After noise reduction and normalization, these inertial wave characteristics are organized into a structured feature group together with the micro-amplitude rhythm, which can reflect the stability of the user's actions, the directional changes of touch intention, and the inherent rhythmic characteristics of the actions. The system performs a short-time Fourier transform on the speech tail drag and extracts envelope features, which are then mapped to the tactile and inertial feature spaces. A short-time Fourier transform is performed on the speech signal acquired by the headphone microphone, focusing on the energy decay process at the end of the user's speech, i.e., tail drag. In the Fourier transform results, the system extracts envelope features such as the low-frequency delay peak, the length of the energy descent interval, and the rate of change of the sound pressure gradient at the end of the speech. These features reflect whether the user is in a state where the voice command is about to end, the tone is gentle, or the intention is still continuing. These speech envelope features are projected into the tactile and inertial feature spaces through a feature space mapper. The mapper uses a combination of multidimensional linear projection and local feature preservation to establish a comparable relationship between the information of the speech tail and features such as tactile pauses and inertial micro-vibration rhythms. This allows the system to determine whether the user is transitioning from voice interaction to touch interaction or whether there is a composite intention. The mapping process also includes feature scale unification to ensure that the three modalities are comparable and have common interpretability during the fusion stage. By combining tactile, inertial, and speech features through a weighted fusion strategy, a multimodal coarse guidance layer with preliminary intention bias is formed. After completing the mapping of tactile fusion features, inertial fluctuation features, and speech tail envelope, the three types of features are combined using a weighted fusion strategy. The weights are determined based on the current action context (such as touch intensity, inertial stability, and whether the speech is still active) and the signal quality and confidence assessment of each type of feature. For example, when the touch features are stable and the lag is low, the weight of the tactile features is increased; when a significant micro-vibration rhythm is detected and the touch action is in a transition phase, the contribution of the inertial features is increased. A local attention window is used during the fusion process so that features from different time periods can contribute independently to the coarse guidance layer without overlapping or canceling each other out. The resulting multimodal coarse guidance layer is a feature structure that presents a preliminary intention bias. It contains consistent temporal organization, local trends, and preliminary directionality across modalities, which can quickly lock in the user's interaction intent and avoid accidental touches or confusion between multiple intents.
[0022] S3: Based on the multimodal coarse guidance layer, the mixed area of contact surface fragmentation path and short-shock scatter points is sequentially split, the disturbance level difference and trend reversal point of each mode in the local window are extracted, and the mode sequence is rearranged according to the reversal relationship.
[0023] The process of sequentially separating the mixed region of contact surface fragmentation path and short-vibration scatter points in S3 is as follows: The overlapping areas of the touch surface and inertial signal in the multimodal coarse guidance layer are locally windowed, and the perturbation gradient and peak value change within each window are calculated. On the time axis, fine-grained local windowing is performed on the segments where tactile and inertial information overlap. The window length can be adaptively selected according to the action category (e.g., 20–80ms for walking scenarios and 5–30ms for static interaction scenarios). A buffer of 1 / 4 window length is reserved at both ends of the window to cover cross-window boundary events. For the touch surface capacitance / pressure waveform and the inertial measurement unit acceleration signal that exist simultaneously in each window, the perturbation gradient (using first-order difference and smoothing) and peak value change (detecting local extrema and measuring peak height, peak width, and peak position changes) are calculated respectively. Second-order statistics such as energy density, signal slope, and mutation rate are extracted to characterize the transient behavior within the window. Interference segments between different modalities are identified, and abnormal or random signals are eliminated through dynamic thresholding. After obtaining the perturbation measurement of each window, a multi-dimensional discrimination process is used to identify possible interference segments. A two-dimensional feature vector space is constructed for tactile and inertial features, and density clustering (e.g., clustering based on Gaussian mixture model) is used to separate frequently occurring, patterned segments from sparse and isolated outliers. Thresholds are dynamically set according to the historical baseline noise level within the window, the current signal-to-noise ratio, and the cross-window consistency index. Signal segments below the threshold and without sustained support are judged as random or abnormal and eliminated. Segments judged as suspected interference but highly correlated with the end of the speech or the preceding and following tactile events are retained and marked as candidate correction segments. The modal segments are reordered according to the perturbation level difference and local trend reversal point to form a modal sequence after interference removal and windowing. After interference removal, the remaining modal segments are sorted according to their perturbation level difference (e.g., peak amplitude difference, energy gradient) and local trend reversal point (i.e., the location where the amplitude or directionality inflection point occurs). The perturbation intensity index (based on peak normalized amplitude and energy integral) and the trend reversal point position (represented by time index) are calculated for each segment. A weighted multi-attribute sorting algorithm is used to map the segments to the index space of the modal sequence, so that segments with high perturbation intensity and continuity in trend are prioritized, while segments with low perturbation or isolation are placed later or as secondary labels. During the sorting process, time consistency checks and shortest interval constraints are performed (e.g., the time interval between adjacent segments must not exceed a preset threshold, otherwise it is considered a sequence break and triggers a merging or re-windowing strategy). Local time compensation (based on the maximum cross-correlation value within the window to estimate and correct the time delay) can also be applied to the time deviation caused by sensing delay.
[0024] The process of rearranging the mode sequence according to the foldback relationship in S3 is as follows: Within each local window, the instantaneous amplitude difference and peak-valley reversal points of the touch and inertial signals are calculated. Within each local window, baseline correction and bandpass filtering are performed on the touch channel and the inertial channel, respectively. The instantaneous envelope value at each time point is obtained by using Hilbert transform or short-time energy estimation, thus forming the instantaneous amplitude curves of the two channels. The two curves are subtracted point by point under the same time reference, and the absolute value is taken to obtain the instantaneous amplitude difference sequence. The maximum value, mean, variance, and kurtosis coefficient of the instantaneous amplitude difference sequence within the window are statistically analyzed as amplitude difference features. Adaptive threshold peak-valley detection (the threshold is determined by a linear combination of the mean and standard deviation within the window) is applied to the original and envelope signals to extract the peak-valley positions. The second-order difference method is used to identify the reversal point (i.e., the time index of the reverse amplitude slope), and the amplitude difference before and after the reversal, the reversal duration, and the reversal rate are recorded. Using a multidimensional ranking algorithm, perturbation level differences and trend reversal points are mapped to the index space of the modal sequence. The perturbation level difference features (such as maximum amplitude difference, energy integral, instantaneous slope) and trend reversal point features (such as reversal time index, reversal amplitude, reversal rate) obtained for each modal segment within the window are merged into a multidimensional feature vector. After standardization of each dimension, feature weights (tactile weight, inertial weight, time-sensitive weight) are assigned based on historical data or online evaluation results. A multidimensional ranking strategy is used to map these vectors into index values of the modal sequence, including weighted linear scoring and trained sorters. The ranking algorithm also adds a time consistency penalty term during mapping and applies a discount factor to low-confidence features, thereby outputting a set of preliminary indices to characterize the relative priority and positional candidates of the segments in the sequence. Time alignment and bias compensation are performed on the foldback nodes between different modes to generate a continuous, identifiable and time-aligned optimized mode sequence. The time delay estimate between adjacent segments is calculated by cross-correlation function or short-time cross-spectral analysis (time delay indicated by the position of cross-correlation peak). If the cross-correlation is unstable, the alternative phase difference estimate (time delay is solved by inversely solving the phase spectrum in the frequency domain) is used to handle nonlinear distortion. The obtained delay estimate is used to perform time resampling or interpolation (e.g., cubic spline or cubic interpolation) on the lagging segments and adjust their foldback time index accordingly. A sensor delay compensation table (generated by factory calibration or online self-calibration during operation) is introduced to correct the fixed channel delay difference. Significant inconsistencies that occur during the alignment process (such as energy non-conservation or event sequence conflict) are checked for consistency. If the check fails, a soft merging or backoff strategy is triggered (e.g., merging adjacent segments and recalculating features or reordering).
[0025] S4: The intention concentration of the selected modal sequence is re-compressed, and the discrete distribution of the contact stagnation amplitude difference, short vibration rebound and the attenuation factor of the speech tail are jointly folded to construct a modal condensation surface with multi-level intention depth.
[0026] The process of re-compressing the intended concentration of the screened modal sequences in S4 is as follows: Calculate the frequency and intensity weight of each modal segment in the modal sequence; perform a traversal scan of the optimized modal sequence, and determine the frequency of each modal segment by repeat detection within the time window. Repeat detection is based on the feature similarity of the segments, and the feature vector similarity comparison is combined with the temporal proximity to identify multiple occurrences of the same action segment. The temporal proximity can be set to the range of tens of milliseconds to hundreds of milliseconds to adapt to different interaction rhythms. Calculate the intensity weight of each identified segment. The intensity assessment adopts a weighted synthesis of three empirical measures: energy approximation, peak amplitude, and duration within the segment (the relative importance of different measures is set according to experience during implementation). Combine the signal-to-noise ratio of the channel in which the segment is located as a confidence correction factor to suppress the influence of low-quality sensing data. Energy approximation is estimated by envelope energy or short-time energy integral, peak amplitude is based on the local maximum envelope value, and duration is measured by the segment time span. Based on the local cumulative energy of modal fragments, the modal sequence is hierarchically compressed to enhance the intention weight of high-intensity fragments. A sliding time window is used to calculate the local cumulative energy distribution on the modal sequence, dividing the time axis into high-energy, medium-energy, and low-energy regions. The time window length and sliding step size are adjustable within the specified range (e.g., the window width can be adjusted between 100–500 ms, and the step size between 10%–50% of the window width). Within each region, the fragment intensity is hierarchically processed according to preset rules. For modal fragments falling into the high-energy region, an intensity amplification and frequency weighting strategy is implemented to highlight their relative importance; for the medium-energy region, a moderate retention strategy is implemented. To reduce the possibility of noise misjudgment, suppression or backgrounding is implemented in low-energy regions. The implementation of hierarchical compression also needs to consider the reward mechanism of occurrence frequency. Segments that appear repeatedly and have high similarity in the same high-energy region should receive additional boosts to their final intention contribution. Upper and lower limit pruning and confidence gating are introduced to avoid judgment imbalance caused by individual extreme values. For example, in walking scenarios, a shorter window width is used and the inertial mode is given a higher initial weight; in static interaction scenarios, the weight of the tactile mode is increased. It is also explained how to dynamically adjust the amplification or suppression intensity based on real-time signal-to-noise ratio and historical samples, so as to stably strengthen high-intensity, repeatable segments as intentional targets in different usage scenarios. A nonlinear normalization function is applied to the compression results to map the intensity of different modes to a unified intent quantization scale, outputting a final mode sequence containing weight labels and time indices. The processing results of each mode segment are scaled, and the weights of different modes after compression are converted into a unified intent quantization scale through nonlinear mapping. The nonlinear method used for mapping is described in the specification (e.g., local normalization combined with an overall compression strategy, local normalization based on a sliding window to maintain local contrast, and then suppressing extreme values through smoothing mapping). Confidence correction is also applied to ensure that the final quantization result reflects signal quality. The output format is a structured final mode sequence entry, each containing start and end time indices, mode type identifier, original compression weights, normalized intent values, confidence labels, and source channel information. At the same time, time smoothing and jump suppression processing are applied to the sequence to avoid decision instability caused by short-term jitter.
[0027] The process of constructing a modal condensed finger surface with multi-level intentional depth in S4 is as follows: The final modal sequence is superimposed with local perturbation features in multiple channels to form an intention matrix. The final modal sequence, which has been compressed and normalized, is read line by line to extract the intention weight, time index, and channel information of each modal segment. At the same time, perturbation features of tactile, inertial, and speech signals, including short-term peak values, amplitude changes, and envelope energy, are obtained from local windows. The modal sequence and local perturbation features are superimposed with local perturbation features in multiple channels according to time index and modal channel to form a two-dimensional or three-dimensional intention matrix. The time axis of the matrix corresponds to the continuous acquisition time, the modal axis corresponds to different signal types, and the superimposed value represents the intention intensity and local perturbation contribution of each modal segment. The contribution of each modality in the matrix is processed through different weighting strategies. For example, tactile modality is given a higher weight in manual interaction scenarios, and inertial modality is more sensitive in dynamic motion scenarios. The original amplitude and frequency information of local perturbation features are preserved. Through a weighted folding algorithm, touch stagnation amplitude difference, short vibration rebound, and speech attenuation features are fused into the intention matrix. The local cumulative value of touch stagnation amplitude difference is calculated, and the touch stagnation time and amplitude information are used as weights to map to the corresponding time period in the matrix. The instantaneous amplitude and rebound frequency of short vibration rebound are mapped to the inertial modal channel to reflect the micro-fluctuations of local actions. The speech tail attenuation features are converted into energy attenuation ratio or envelope attenuation metric and mapped to the speech modal channel. In this weighted folding process, the contribution coefficients of each feature should be appropriately adjusted to avoid a single modality dominating the matrix, while preserving the temporal relationship and local peak information of the original perturbation, so as to ensure that the fused intention matrix accurately presents the comprehensive intention indication of multimodal signals in the temporal and modal dimensions.
[0028] The intention matrix is subjected to multi-layer convolutional mapping and time series compression to form a modal condensation surface with multi-level intention depth. The convolutional operation extracts the trend of continuous action patterns on the time axis and the coupling relationship between different signal channels on the modal axis, generating higher-order intention feature layers layer by layer. At the same time, windowing pooling or compression processing is applied to downsample the time series to reduce redundant information and highlight the core intention signal. After multi-layer mapping and compression, a modal condensation surface with multi-level intention depth is obtained. Each layer corresponds to an intention representation at a different scale, which can reflect both local action bias and present the comprehensive intention trend across modalities.
[0029] S5: Under the action of modal condensation, identify the conflict node between the main touch pulse and the short vibration accompanying pulse, and combine it with the short-cycle trigger offset signal that appears in the input event receipt of the mobile terminal to generate the intent establishment path.
[0030] The process of generating intent and establishing a path in S5 is as follows: By analyzing the pulse peaks and time intervals of each layer in the modal condensation surface, conflict nodes between the main pulse and the accompanying pulse are identified. The modal condensation surface with multi-level intention depth is read, and the pulse peaks and their corresponding time intervals of each layer are statistically analyzed and sorted. By comparing the pulse amplitudes and occurrence times in different modal layers, the overlapping or similar occurrence phenomena between the main pulse and the accompanying pulse can be identified. The conflict node is the position of the main pulse and the accompanying pulse with similar or overlapping amplitudes within the same time window. By setting thresholds to determine the peak difference and the minimum time interval, the conflict node can accurately capture the key time points that may cause confusion in intention determination. By combining the short-cycle trigger offset signal received by the mobile terminal, the time stamp of the conflict node is dynamically corrected. The offset signal received by the terminal can reflect the time drift and delay of the actual touch event or inertial micro-vibration response. By analyzing the amplitude and direction of the offset signal, the time stamp of the conflict node is finely adjusted to eliminate the error caused by sensor delay, sampling time difference or signal jitter. This process ensures that the time position of the conflict node in the modal condensation surface is synchronized with the actual user operation time, while retaining the relative intensity information of each pulse. Based on the conflict nodes and the corrected pulse sequence, an intent-establishing path is generated. The main pulse, the accompanying pulse, and their corrected time information are correlated in chronological order. The continuity, amplitude variation trend, and multimodal coupling characteristics between pulses are analyzed to generate a path trajectory that reflects the user's true interaction intent. The generated intent path is used to drive the headphone interaction commands and prompts, while retaining the pulse hierarchy and modal channel information.
[0031] Example 2: Figure 2 As shown, a multimodal interaction system for smart headphones includes: Contact surface acquisition module: performs multi-dimensional acquisition of the contact point texture formed during the pressing, sliding and light squeezing processes, and extracts sudden features to form contact surface source trace fragments; Modal mapping module: interleaving and splicing contact surface source trace fragments within a specified time period to construct a multimodal coarse guidance layer with intention bias; Sequence decomposition module: It breaks down the mixed area of contact surface fragmentation path and short-shock scattered points, extracts the disturbance level difference and trend reversal point, and rearranges the modal sequence; Intent Condensation Module: Performs intent concentration compression on modal sequences, and combines touch, short vibration and voice features to form a multi-layered intent depth modal condensation surface; Path determination module: Under the action of modal condensation, it identifies the conflict node between the main touch pulse and the short vibration accompanying pulse, and combines the short period trigger offset signal to generate the intended path.
[0032] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0033] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal interaction method for smart headphones, characterized in that, Includes the following steps: Multi-dimensional acquisition of the contact texture clusters formed on the outer contact surface of the earphone during the pressing, sliding and light squeezing process is performed. The degree of contraction of the contact spots, the discontinuous path of micro-pressure advancement and the sudden characteristics of contact surface stress jump are extracted to form contact surface source trace fragments for depicting the initial interaction tendency. Interlacing and splicing the trace fragments on the touch surface within a specified time period, and incorporating the residual texture of the fingertips formed by the user under similar action conditions, the habitual rhythm of the slight adjustment of the hand, and the tail dragging of the voice breath, a multimodal coarse guidance layer that presents intention bias is constructed. Based on the multimodal coarse guidance layer, the mixed area of contact surface fragmentation path and short earthquake scatter points is sequentially split, the disturbance level difference and trend reversal point of each mode in the local window are extracted, and the mode sequence is rearranged according to the reversal relationship. The intention concentration of the modal sequence is re-compressed, and the discrete distribution of the contact stagnation amplitude difference, short vibration rebound and the attenuation factor of the speech tail are jointly folded to construct a modal condensation surface with multi-level intention depth. Under the action of modal condensation, the conflict node between the main touch pulse and the short vibration accompanying pulse is identified, and combined with the short-cycle trigger offset signal that appears in the input event receipt of the mobile terminal, the intent establishment path is generated.
2. The multimodal interaction method for smart headphones according to claim 1, characterized in that, The process of forming a touch source trace fragment used to depict the initial interaction tendency is as follows: The displacement changes and contact area fluctuations of the contact clusters under different pressure intensities are collected by a multi-channel touch sensor. The differential motion analysis method is used to identify the short-term continuous sliding path and sudden pressure point of the contact point, and the continuous signal is divided into initial segment units by the time weighting function. The initial segment unit is encoded with feature vectors, including contraction amplitude, path curvature and stress peak, to form a contact source trace segment for depicting the initial interaction tendency.
3. The multimodal interaction method for smart headphones according to claim 2, characterized in that, The process of interleaving and splicing contact surface source traces within a specified time period is as follows: The contact surface source trace fragments are segmented and indexed based on the collection timestamp to form a modal fragment set of time-division windows; Within each time-sharing window, the contact surface source trace fragments are arranged in a staggered sequence, and the temporal characteristics of different fragments overlap to form an interleaved combination; By applying sliding time frames and phase adjustment methods, minor deviations between segments are corrected and overlapping interference is eliminated, generating continuous multimodal fusion segments.
4. The multimodal interaction method for smart headphones according to claim 3, characterized in that, The process of constructing a multimodal coarse-grained guiding layer that presents intentional bias is as follows: The residual texture at the fingertips is time-aligned with the multimodal fusion fragment to form preliminary tactile fusion features; Analyze the short-vibration patterns of the hand's micro-adjustment trajectory and the output of the inertial unit to extract the micro-amplitude rhythm and inertial wave characteristics; Short-time Fourier transform is performed on the lingering of the speech tail segment, and envelope features are extracted and mapped to the tactile and inertial feature space. By combining tactile, inertial, and speech features through a weighted fusion strategy, a multimodal coarse guidance layer with preliminary intention bias indications is formed.
5. The multimodal interaction method for smart headphones according to claim 4, characterized in that, The process of sequentially decomposing the mixed region of contact surface fragmentation path and short-range earthquake scatter points is as follows: The overlapping area between the contact surface and the inertial signal in the multimodal coarse guidance layer is locally windowed, and the perturbation gradient and peak value change within each window are calculated. Identify interference segments between different modes and remove abnormal or random signals using dynamic thresholds; The modal segments are reordered according to the perturbation level and local trend reversal points to form a modal sequence that has undergone interference removal and windowing.
6. The multimodal interaction method for smart headphones according to claim 5, characterized in that, The process of rearranging the modal sequence according to the foldback relationship is as follows: Within each local window, calculate the instantaneous amplitude difference and peak-valley reversal point between the touch and inertial signals; Using a multidimensional sorting algorithm, perturbation differences and trend reversal points are mapped to the index space of the modal sequence; Time alignment and deviation compensation are performed on the turnaround nodes between different modes to generate a continuous, identifiable and time-aligned optimized mode sequence.
7. A multimodal interaction method for smart headphones according to claim 6, characterized in that, The process of re-compressing the intended concentration of the screened modal sequences is as follows: Calculate the frequency and intensity weight of each modal segment in the modal sequence; Based on the local accumulated energy of modal segments, the modal sequence is hierarchically compressed to enhance the intention weight of high-intensity segments; A nonlinear normalization function is applied to the rectification results to map the intensity of different modes to a unified intention quantization scale, and the final mode sequence with weight labels and time indexes is output.
8. A multimodal interaction method for smart headphones according to claim 7, characterized in that, The process of constructing a modal condensed finger surface with multi-level intentional depth is as follows: The final modal sequence is superimposed with local perturbation features through multiple channels to form an intention matrix; By using a weighted folding algorithm, the features of touch stagnation amplitude difference, short vibration rebound, and voice attenuation are fused into the intention matrix; Multi-layer convolutional mapping and time-series compression are applied to the intention matrix to form a modal condensation surface with multi-layer intention depth.
9. A multimodal interaction method for smart headphones according to claim 8, characterized in that, The process of generating intent and establishing a path is as follows: By analyzing the peak values and time intervals of pulses in each layer of the modal condensation surface, conflict nodes between the main pulse and the accompanying pulse are identified. By combining the short-period trigger offset signal received by the mobile terminal, the time stamp of the conflicting node is dynamically corrected. Based on the conflict nodes and the corrected pulse sequence, the intended path is generated.
10. A smart earphone multimodal interaction system, applied to the method described in any one of claims 1-9, characterized in that, include: Contact surface acquisition module: performs multi-dimensional acquisition of the contact point texture formed during the pressing, sliding and light squeezing processes, and extracts sudden features to form contact surface source trace fragments; Modal mapping module: interleaving and splicing contact surface source trace fragments within a specified time period to construct a multimodal coarse guidance layer with intention bias; Sequence decomposition module: It breaks down the mixed area of contact surface fragmentation path and short-shock scattered points, extracts the disturbance level difference and trend reversal point, and rearranges the modal sequence; Intent Condensation Module: Performs intent concentration compression on modal sequences, and combines touch, short vibration and voice features to form a multi-layered intent depth modal condensation surface; Path determination module: Under the action of modal condensation, it identifies the conflict node between the main touch pulse and the short vibration accompanying pulse, and combines the short period trigger offset signal to generate the intended path.