Earphone wearing state identification method based on environmental acoustic characteristics and related equipment
Through a method based on environmental acoustic features, the audio stream is dynamically captured and physical vibration features are extracted, voice features are suppressed, and a lightweight temporal inference engine is built. This solves the anti-interference and privacy compliance issues of headphone wearing status detection, achieving high-precision detection and reducing hardware costs.
Patent Information
- Application Number
- CN202510938418.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing headphone wearing status detection technology has poor anti-interference ability, strong hardware dependence, and poses privacy compliance risks.
By dynamically capturing the ambient audio stream of the headphone microphone, generating a time-correlated acoustic matrix, extracting the physical vibration feature set related to the wearing state, and suppressing the voice features, a lightweight time-series inference engine is constructed, and the decision boundary is established using multi-scale time domain convolution and attention weighting mechanism.
It achieves high-precision, anti-interference headphone wearing status detection, reduces hardware customization costs, meets data protection regulations, and adapts to different headphone types.
Smart Images

Figure CN120708656A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for identifying earphone wearing status based on environmental acoustic features and related equipment. Background Art
[0002] With the rapid development of smart headset technology, users' demand for intelligent interaction with devices is increasing. As a core basic function for realizing intelligent control (such as automatic play / pause), the detection accuracy and reliability of headset wearing status directly affect the user experience. The current mainstream detection technology relies on physical sensors to achieve status perception. This technology route has been widely deployed in consumer electronic products such as Bluetooth headsets and sports headsets. Existing technologies mainly detect the wearing status of headsets through contact sensors (such as capacitive sensors) or motion sensors (such as accelerometers). However, existing technologies have key defects that affect practicality, such as poor anti-interference ability, strong hardware dependence, and privacy compliance risks. Summary of the Invention
[0003] The present application provides a headphone wearing status recognition method and related equipment based on environmental acoustic characteristics, which can achieve high-precision and interference-resistant headphone wearing status detection.
[0004] In one aspect, the present application provides a method for identifying earphone wearing status based on ambient acoustic characteristics, the method comprising:
[0005] Dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with temporal correlation in real time;
[0006] Implementing feature decoupling in acoustic matrix processing to extract a set of physical vibration features that are strongly correlated with the wearing state, while suppressing semantic features related to the user's voice;
[0007] Constructing a lightweight temporal reasoning engine that receives the physical vibration feature set as input, and establishing a wearing state decision boundary through multi-scale temporal convolution and attention weighting mechanism;
[0008] Output a classification result corresponding to the decision boundary, wherein the classification result includes identification labels of wearing state, not wearing state, and wearing direction.
[0009] On the other hand, the present application provides a device for identifying earphone wearing status based on environmental acoustic characteristics, the device comprising:
[0010] A generation module is used to dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with time-series correlation in real time;
[0011] A decoupling module, configured to implement feature decoupling operations in acoustic matrix processing to extract a set of physical vibration features that are strongly correlated with the wearing state while suppressing semantic features related to the user's voice;
[0012] An inference module, configured to construct a lightweight temporal inference engine that receives the physical vibration feature set as input and establishes a wearing state decision boundary through a multi-scale temporal convolution and attention weighting mechanism;
[0013] An output module is used to output a classification result corresponding to the decision boundary, wherein the classification result includes identification labels of wearing, not wearing state and wearing direction.
[0014] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the technical solution of the method for identifying the wearing status of headphones based on environmental acoustic features as described above are implemented.
[0015] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-mentioned method for identifying the wearing status of an earphone based on environmental acoustic features.
[0016] From the technical solutions provided in this application, it can be seen that, on the one hand, by dynamically capturing the ambient audio stream and generating an acoustic matrix, a set of physical vibration features that are strongly correlated with the wearing state is extracted, the vibration signal caused by wearing is effectively separated from the ambient noise, and the misjudgment problem caused by sweat contamination and continuous movement interference is overcome; on the other hand, by suppressing the semantic features related to the user's voice, the physical vibration feature set is extracted and a decision boundary is established to ensure that only the physical acoustic features related to vibration are analyzed, and the voice content collection behavior is avoided from the source, meeting the requirements of data protection regulations; thirdly, by constructing a lightweight temporal inference engine, state recognition is achieved based on acoustic matrix processing, that is, a pure software algorithm solution is used to replace the dependence on hardware sensors, and a unified audio processing pipeline is used to adapt to different types of headphones (for example, Bluetooth headphones, wired headphones, on-ear headphones or in-ear structures, etc.), thereby reducing the cost of hardware customization development; fourthly, a decision boundary is established through multi-scale time domain convolution and attention weighting mechanism, and the attention mechanism is combined to focus on key time segments (such as the contact response at the moment of wearing), significantly improving the ability to distinguish complex scenes. In summary, the technical solution of the present application can achieve high-precision, interference-resistant headphone wearing status detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flow chart of a method for identifying earphone wearing status based on environmental acoustic features provided by an embodiment of the present application;
[0019] Figure 2 1 is a schematic structural diagram of a headphone wearing status recognition device based on environmental acoustic features provided in an embodiment of the present application;
[0020] Figure 3 It is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] In this specification, adjectives such as first and second may be used only to distinguish one element or action from another element or action, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.
[0023] In this specification, for the convenience of description, the sizes of various parts shown in the drawings are not drawn according to the actual proportions.
[0024] With the rapid development of smart headphone technology, users are increasingly demanding intelligent device interaction. Headphone wear status detection, a core fundamental function for enabling intelligent controls (such as automatic play / pause), requires accurate and reliable detection that directly impacts the user experience. Current mainstream detection technologies rely on physical sensors for state awareness, a technology widely deployed in consumer electronics such as Bluetooth headsets and sports headphones. Existing technologies primarily detect headphone wear status using contact sensors (such as capacitive sensors) or motion sensors (such as accelerometers). Typical implementations include the following three approaches: 1) Contact sensor solutions, which employ electrode arrays placed on the inner wall of the earcup to determine wear status by detecting changes in skin contact capacitance; 2) Sensor solutions, which utilize accelerometers to capture characteristic motion patterns (such as changes in gravity direction) during the donning and removal process; and 3) Hybrid sensor solutions, which fuse accelerometer and gyroscope data and combine them with simple threshold rules for state classification. However, the above-mentioned existing technologies have the following key defects that affect their practicality: 1) Poor anti-interference ability, among which contact sensors are easily affected by sweat and oil, resulting in false triggering, and motion sensors have a high misjudgment rate due to continuous vibration when users exercise (such as running and cycling); 2) Strong hardware dependence, specifically, different types of headphones require customized sensor modules, and two sets of detection systems need to be independently developed for on-ear and in-ear headphones, increasing R&D costs; 3) Privacy compliance risks, that is, some solutions that use auxiliary microphones need to collect voice content for voiceprint analysis, which violates the provisions of data protection regulations on the processing of biometric data.
[0025] In view of the above problems in the prior art, this application proposes a method for identifying the wearing status of headphones based on environmental acoustic characteristics, the flow chart of which is shown in the attached figure. Figure 1 As shown, it mainly includes steps S101 to S104, which are detailed as follows:
[0026] Step S101: dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with time-series correlation in real time.
[0027] Considering that, on the one hand, the headphone microphone is a universal acoustic sensor, and the original audio stream it obtains contains time-domain phase information and frequency-domain energy distribution, and the construction of the time-correlated acoustic matrix is essentially to convert the one-dimensional waveform signal into a two-dimensional time-frequency representation through frame segmentation and time-frequency conversion. This conversion retains the causal relationship of the mechanical vibration signal on the time axis. For example, the contact vibration signal during the wearing of the headphone will form a continuous energy diffusion band, while the ambient noise is distributed in discrete points. This means that if it is replaced by single-frame spectrum analysis, the time evolution characteristics of the vibration event will be lost, and it will be impossible to distinguish between instantaneous impact and continuous vibration states; on the other hand, signal processing that is separated from the time dimension will cause the system to be unable to recognize progressive actions (such as the slow sliding process of the headphone). At the same time, due to the lack of time context association, sudden environmental noise can be easily misjudged as a wearing state switching action. Therefore, in an embodiment of the present application, the ambient audio stream of the headphone microphone can be dynamically captured, and an acoustic matrix with time correlation can be generated in real time.
[0028] Specifically, as an embodiment of the present application, dynamically capturing the ambient audio stream of the headphone microphone and generating an acoustic matrix with temporal correlation in real time can be achieved through steps S1011 to S1015, as detailed below:
[0029] Step S1011: Use dual microphone signals to collaboratively collect ambient audio streams.
[0030] Specifically, the main microphone directionally collects vibration signals close to the ear canal, and the auxiliary microphone synchronously captures the air-conducted noise in the environment. The hardware filter pre-eliminates high-frequency interference signals and retains the acoustic characteristics within the effective frequency band.
[0031] Step S1012: Frame and time-domain window the ambient audio stream.
[0032] Specifically, the ambient audio stream can be divided into analysis units of fixed length, with adjacent units partially overlapping. Then, a special window function is used to perform amplitude weighting on each audio stream to reduce the frequency domain distortion caused by segmentation and obtain a windowed signal.
[0033] Step S1013: performing frequency domain conversion and noise separation on each windowed signal segment.
[0034] Specifically, each windowed signal is subjected to spectrum transformation to obtain frequency-energy distribution characteristics. At the same time, similar noise components in the main signal are dynamically estimated and subtracted based on the characteristic pattern of the auxiliary microphone signal to obtain the denoised frequency stream.
[0035] Step S1014: reconstruct the noise frequency stream.
[0036] Specifically, the linear frequency scale can be converted into a non-uniform frequency band division that conforms to human auditory perception. Then, the spectral energy is integrated within each auditory band to form a compressed representation, and finally the auditory scale spectrum is obtained.
[0037] Step S1015: constructing a spatiotemporal acoustic matrix based on the auditory scale spectrum.
[0038] Specifically, the auditory scale spectra of multiple consecutive analysis units can be stacked and arranged in time sequence to form a three-dimensional structure. Then, the auditory frequency band of each three-dimensional structure can be independently calibrated by sliding energy to eliminate the influence of slow changes in the environment and standardize the dynamic characteristics.
[0039] Step S102: Implementing a feature decoupling operation in the acoustic matrix processing to extract a set of physical vibration features that are strongly correlated with the wearing state, while suppressing semantic features related to the user's voice.
[0040] It should be noted that the acoustic wave frequencies generated by the vibration of the earphone structure range from 200 to 1500 Hz, overlapping with the air-conducted speech frequencies of 300 to 3400 Hz in the frequency domain. However, the vibration signal has a higher mechanical impedance and can be separated using the energy gradient characteristics of the critical frequency band. Semantic feature suppression, on the other hand, eliminates the fundamental frequency component of human speech through bandpass filtering, making it impossible for the system to reconstruct speech content at the signal processing level. This frequency-domain excision-based technique avoids the risk of caching raw speech data compared to feature masking schemes commonly used in speech recognition. This fact means that if only mixed features are extracted without suppressing the semantic components, the system must retain the complete speech frequency band information during data processing. This would trigger regulatory requirements for "biometric data," making the product unsuitable for legal deployment in major markets. To avoid legal risks, after generating an acoustic matrix with temporal correlation, feature decoupling can be performed during acoustic matrix processing to extract a set of physical vibration features strongly correlated with the wearing state while suppressing semantic features related to the user's speech.
[0041] Furthermore, since static energy features cannot describe the continuity of motion, a time-varying energy gradient spectrum is needed to capture the instantaneous rate of change. However, a two-dimensional matrix loses frequency band correlation. Therefore, in order to avoid confusing wind noise with wearing vibration and effectively detect slowly varying motions, the physical vibration feature set that is strongly correlated with the wearing state can be extracted in this embodiment through steps S1021 to S1023, as detailed below:
[0042] Step S1021: Divide the acoustic matrix into critical frequency sub-bands, and calculate the time-varying energy gradient spectrum of each sub-band.
[0043] Specifically, according to the differences in the sensitivity of the human ear to vibrations in different frequency bands, the acoustic signal can be divided into key response frequency bands, namely the bone conduction dominant area suitable for on-ear headphones and the air vibration sensitive area suitable for in-ear headphones. The former focuses on the 200-800Hz range, which can effectively transmit the structural vibration generated by the contact between the headphones and the head; the latter extends to the 400-1500Hz range to capture the acoustic characteristics caused by changes in air pressure in the ear canal; then, the energy intensity of all frequency components in the area is aggregated within each divided frequency band to form an energy distribution reflecting a specific vibration mode. At the same time, overlapping analysis window processing is used to ensure a smooth transition of energy characteristics of adjacent time segments to avoid the interruption of motion characteristics caused by segmentation; finally, by calculating the difference rate of energy values in adjacent time periods, the instantaneous change trend of vibration intensity is captured, that is, a positive change indicates an increase in contact pressure, such as a wearing action, and a negative change indicates a decrease in contact, such as a falling-off process. At the same time, short-term (for example, 30ms level) and long-term (for example, 100ms level) energy gradients are analyzed, taking into account the detection requirements of transient actions and continuous state changes, and completing the analysis of the time-varying energy gradient spectrum of each sub-band.
[0044] Step S1022: applying environmental noise baseline calibration to the time-varying energy gradient spectrum to generate an interference-resistant vibration intensity representation.
[0045] Specifically, the environmental noise baseline calibration is applied to the time-varying energy gradient spectrum to generate an anti-interference vibration intensity representation, which includes a noise background learning link, a dynamic noise suppression link, and a vibration intensity normalization link. The noise background learning link can be collected during the initial 5-second silence period when the headphones are placed in a vibration-free environment (for example, a desktop), and then the energy gradient baseline value of each sub-band is recorded. Dynamic noise suppression can eliminate environmental background noise from gradient spectrum in real time. in, is the rate of change of the time-varying energy gradient spectrum, α is the safety factor to prevent over-suppression of weak vibrations, ReLU(x)=max(0,x), which is used to eliminate negative interference contributions. The vibration intensity normalization step is calibrated according to the maximum sensitivity of the sub-band Among them, is the human perception threshold of the kth sub-band, where the low-frequency β value of the on-ear headphone is smaller than the high-frequency value.
[0046] Step S1023: stacking the anti-interference vibration intensity representation along the time dimension to construct a three-dimensional time-varying feature tensor.
[0047] Specifically, the anti-interference vibration intensity representation is stacked along the time dimension to construct a three-dimensional time-varying feature tensor. This can be achieved through three steps: time segment interception, multi-channel tensor encapsulation, and motion interference compensation. In the time segment interception step, the vibration intensity of L consecutive time points can be intercepted with the current moment as the center, that is:
[0048] P(t)=[V1(t-L+1),V1(t-L+2),...,V1(t)]...[V K (t-L+1),V K (t-L+2),...,V K (t)]
[0049] In the multi-channel tensor encapsulation stage, K sub-band data can be stacked along the new dimension, that is:
[0050]
[0051] In motion interference compensation, when violent motion is detected, the time dimension is automatically compressed to , and the inter-subband correlation constraint is strengthened, that is:
[0052]
[0053] Where, is the neighborhood subband weighting coefficient, for example, the center subband has a weight of 0.6, and the adjacent subbands have weights of 0.2 each.
[0054] From step S1021 to step S1023 of the above embodiment, it can be seen that by accurately separating the target vibration frequency band and quantifying the instantaneous changes, the interference of environmental noise is effectively suppressed. Since the three-dimensional tensor retains the spatiotemporal coupling relationship, the system can track the evolution of continuous actions, thereby significantly improving the reliability of identifying slow-changing processes such as earphones slipping off.
[0055] Since reverse wearing samples are rare in real scenarios, it is necessary to generate adversarial negative samples through physical field simulation. Fixed-length pooling will truncate short-term effective signals, so it is necessary to deploy deformable pooling to dynamically compress non-significant segments. Therefore, as an embodiment of the present application, suppressing semantic features related to user speech can be: constructing a speech privacy barrier layer in the time-frequency feature space; in the speech privacy barrier layer, using a bandpass filter bank to filter out the fundamental frequency component of human speech; applying a nonlinear projection transformation to the filtered time-varying feature tensor to map the feature vector potentially containing speech information to a subspace decoupled from the vibration feature. From the above embodiment, it can be seen that due to the introduction of direction-flipped adversarial samples, the system has abnormal state immunity, and dynamic compression of non-significant periods allows computing resources to focus on key action segments, thereby significantly enhancing the robustness of conventional wear detection. The above embodiment of suppressing semantic features related to user speech can also include a privacy leakage self-destruction mechanism, that is, when it is detected that the suppression rate of the fundamental frequency component of speech is lower than a preset threshold ratio (e.g., 80%) or the feature decoupling effectiveness index drops by more than a preset threshold ratio (e.g., 30%), the feature cache of the headset is automatically cleared and the model is reset. After the self-destruction mechanism is triggered, a virtual noise training set can be generated locally for emergency retraining of the model, and encrypted exception logs can be uploaded to the security audit server. During the retraining process, the sparsity constraint coefficient of the bottleneck layer of the autoencoder network is forcibly enhanced.
[0056] Furthermore, considering the significant differences in confidence levels of feature frames (high signal-to-noise ratio / low signal-to-noise ratio), it is necessary to set a dual threshold control path switching. Transient events require redundancy compression, and continuous processes require associated context. The use of maximum pooling and recursive fusion dual paths can effectively avoid losing transient event details or amplifying long-term noise. Therefore, the above embodiment applies a nonlinear projection transformation to the filtered time-varying feature tensor, specifically by: constructing an autoencoder network with sparse activation characteristics, whose bottleneck layer dimension is lower than the input layer dimension by more than a preset threshold ratio; passing the filtered time-varying feature tensor through the autoencoder, and actively removing orthogonal bases in the encoding vector that are highly correlated with the speech content. In the above embodiment, since high-confidence frames trigger pooling to compress redundant information to avoid false triggering, and low-confidence frames trigger recursive fusion to retain associations to prevent missed detection, instantaneous contact events such as earphones touching the collar can be accurately captured.
[0057] Step S103: Construct a lightweight temporal reasoning engine that receives a physical vibration feature set as input, and establish a wearing state decision boundary through multi-scale time domain convolution and attention weighting mechanism.
[0058] On the one hand, single-scale convolution cannot capture both slow-changing processes and transient events simultaneously, which inevitably leads to misclassification of motion patterns. The design of multi-scale convolution kernels is a mathematical simulation of the physical propagation laws of acoustic vibrations. Specifically, large-scale convolution kernels can match the inertial delay characteristics of human movements, such as the slow-changing process of skin contact separation when taking off headphones (time constant of approximately 200ms), and small-scale convolution kernels correspond to the transient response caused by the elastic deformation of the material, such as the micro-pressure fluctuations at the moment the earplugs are inserted into the ears (duration of approximately 5 to 15ms); on the other hand, the attention mechanism is essentially a selective amplifier that constructs time-domain features. It originates from the masking effect of acoustic signals - weak vibration signals in a strong noise environment will be ignored by the auditory system. By calculating the energy entropy confidence of each time frame, it can dynamically suppress time periods dominated by environmental noise (for example, noisy streets) while enhancing segments with clear vibration signals (for example, adjusting headphones when the user is stationary), which can adapt to dynamic acoustic scenes. Therefore, after extracting the physical vibration feature set that is strongly correlated with the wearing state, a lightweight temporal inference engine can be constructed that receives the physical vibration feature set as input, and establishes the wearing state decision boundary through multi-scale time domain convolution and attention weighting mechanism.
[0059] It should be noted that the lightweight temporal reasoning engine in the above embodiment is a pre-trained reasoning model. Furthermore, considering that it is difficult to collect extreme cases (for example, reverse wearing) in real scenarios, simulation generation breaks through the data bottleneck, while fixed-length pooling will discard key fragments (such as short signals at the moment of wearing), and dynamic compression focuses on the effective period; in other words, the lack of negative samples of direction reversal will cause the model to be unable to recognize unconventional states, making fixed pooling ineffective in high-motion scenarios. Therefore, as an embodiment of the present application, the training method of the above lightweight temporal reasoning engine can be: generating negative samples of wearing direction reversal through acoustic physics simulation; deploying a deformable time pooling module after the convolution layer, which integrates an attention weight generation unit for dynamically analyzing the decision confidence of the feature frame; based on the analysis results of the decision confidence, performing dynamic compression of non-significant time segments. The above training method can give the lightweight temporal reasoning engine a strong generalization ability to resist unconventional interference, and optimize computing efficiency through an intelligent time segment screening mechanism to avoid long-term noise interference.
[0060] Since a single threshold cannot adapt to the difference in feature quality, that is, high signal-to-noise ratio needs to suppress redundancy, and low signal-to-noise ratio needs to retain association, and transient events (for example, headphones touching the collar) require pooling compression, and continuous processes (for example, motion shaking) require recursive association, therefore, in order to achieve adaptive optimization of fine-grained temporal features, maintain high recognition robustness in noise fluctuation scenarios, and reduce the risk of missed detection of key actions, the deformable time pooling module of the above embodiment dynamically analyzes the decision confidence of the feature frame, specifically outputs the confidence score of each feature frame through the attention weight generation unit. When the confidence score is higher than the first threshold, the maximum pooling path is activated to compress the feature dimension. When the confidence score is lower than the second threshold, it switches to the gated recursive fusion path to retain the temporal association information.
[0061] It is understandable that when an object moves at high speed, the frequency of physical vibration doubles (such as running vibration), and it is generally necessary to shorten the window to capture transient details, while the low-confidence marker features are inaccurate, and the original gradient spectrum needs to be combined for physical authenticity verification. In order to automatically correct detection errors in scenarios of intense exercise and signal attenuation, and ensure the physical credibility of the results through multi-source signal fusion, the above embodiment also includes an adaptive pooling mechanism, specifically: real-time monitoring of the motion sensor data of the headset, when the motion acceleration exceeds the set acceleration threshold, shortening the pooling window width to improve the transient response capability; when the proportion of low-confidence feature frames exceeds the set proportion threshold, performing cross-frame coherence check based on the time-varying energy gradient spectrum, and correcting the current pooling output.
[0062] Furthermore, the above embodiment may also include an online optimization mechanism for a lightweight temporal reasoning engine, that is, periodically calculating the confidence variance of the state prediction results; when the confidence variance continues to exceed the stability threshold, starting an incremental model update process, wherein the incremental model update process includes scene noise fingerprint analysis and local optimization of model parameters. Specifically, the incremental model update process may be: real-time collection of background noise in the environment of the headset, extraction of frequency domain fingerprint features through Fourier transform, and generation of an adversarial sample set containing perturbation factors based on the fingerprint features; freezing the network weights of the convolutional layer and the recursive layer in the temporal reasoning engine, and only unlocking the fully connected parameters of the output layer; the adversarial sample set is input into the unlocked model layer, and adversarial regularization constraints related to the noise fingerprint are injected during the back propagation process. The optimization formula is: Among them, L ce is the cross entropy loss, adv (θ) is the adversarial loss λ is the fingerprint correlation adjustment coefficient, is the gradient operator.
[0063] Step S104: Outputting the classification result corresponding to the decision boundary, wherein the classification result includes identification labels of wearing state, not wearing state, and wearing direction.
[0064] Furthermore, the method of the above embodiment may also include, during the system verification stage, generating a micro-vibration signal of the earphone in a semi-detached state through a piezoelectric vibrator, and synchronously collecting accelerometer data to verify the authenticity of the vibration pattern; inputting the micro-vibration signal into a pre-trained generative adversarial network to synthesize an edge scene sample set in the fuzzy area of the decision boundary; using the edge scene sample set to perform adversarial training on the temporal reasoning engine to force the model to establish a clear wearing state decision boundary, wherein using the edge scene sample set to perform adversarial training on the temporal reasoning engine to force the model to establish a clear wearing state decision boundary can be: introducing a topological continuity constraint term in the model loss function, which constraint term penalizes the discrete clustering distribution of the feature space of the earphone in the same wearing state, and drives similar features to aggregate toward the topological manifold.
[0065] From the above attached Figure 1 The example of a headphone wearing state recognition method based on environmental acoustic features shows that, on the one hand, by dynamically capturing the ambient audio stream and generating an acoustic matrix, a set of physical vibration features strongly correlated with the wearing state is extracted, effectively separating the vibration signal caused by wearing from the environmental noise, and overcoming the problem of misjudgment caused by sweat contamination and continuous movement interference; on the other hand, by suppressing the semantic features related to the user's voice, the physical vibration feature set is extracted and a decision boundary is established, ensuring that only the physical acoustic features related to vibration are analyzed, avoiding the collection of voice content at the source and meeting the requirements of data protection regulations; thirdly, by constructing a lightweight temporal inference engine, state recognition is achieved based on acoustic matrix processing, that is, a pure software algorithm solution is used to replace the dependence on hardware sensors, and a unified audio processing pipeline is used to adapt to different headphone types (for example, Bluetooth headphones, wired headphones, on-ear headphones or in-ear structures, etc.), reducing the cost of hardware customization development; fourthly, a decision boundary is established through multi-scale time domain convolution and attention weighting mechanism, combined with the attention mechanism to focus on key time segments (such as the contact response at the moment of wearing), significantly improving the ability to distinguish complex scenes. In summary, the technical solution of the present application can achieve high-precision, interference-resistant headphone wearing status detection.
[0066] Please see the attached Figure 2 , is a headphone wearing status recognition device based on environmental acoustic features provided by an embodiment of the present application. The device may include a generation module 20, a decoupling module 202, an inference module 203, and an output module 204, as detailed below:
[0067] A generation module 201 is used to dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with time sequence correlation in real time;
[0068] a decoupling module 202 for performing a feature decoupling operation in acoustic matrix processing to extract a set of physical vibration features that are strongly correlated with the wearing state while suppressing semantic features related to the user's voice;
[0069] The reasoning module 203 is used to build a lightweight temporal reasoning engine that receives a physical vibration feature set as input and establishes a wearing state decision boundary through multi-scale temporal convolution and attention weighting mechanism;
[0070] The output module 204 is configured to output a classification result corresponding to the decision boundary, wherein the classification result includes identification labels of the wearing state, the non-wearing state, and the wearing direction.
[0071] From the above attached Figure 2 The example of a headphone wearing state recognition device based on environmental acoustic features shows that, on the one hand, by dynamically capturing the ambient audio stream and generating an acoustic matrix, a set of physical vibration features strongly correlated with the wearing state is extracted, effectively separating the vibration signal caused by wearing from the environmental noise, and overcoming the problem of misjudgment caused by sweat contamination and continuous movement interference; on the other hand, by suppressing the semantic features related to the user's voice, extracting the physical vibration feature set and establishing a decision boundary, ensuring that only the physical acoustic features related to vibration are analyzed, avoiding the collection of voice content from the source, and meeting the requirements of data protection regulations; thirdly, by constructing a lightweight temporal inference engine, state recognition is achieved based on acoustic matrix processing, that is, a pure software algorithm solution is used to replace the dependence on hardware sensors, and a unified audio processing pipeline is used to adapt to different headphone types (for example, Bluetooth headphones, wired headphones, on-ear headphones or in-ear structures, etc.), reducing the cost of hardware customization development; fourthly, a decision boundary is established through multi-scale time domain convolution and attention weighting mechanism, and the attention mechanism is combined to focus on key time segments (such as the contact response at the moment of wearing), significantly improving the ability to distinguish complex scenes. In summary, the technical solution of the present application can achieve high-precision, interference-resistant headphone wearing status detection.
[0072] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for a method for identifying a headphone wearing state based on environmental acoustic features. When the processor 30 executes the computer program 32, the steps of the embodiment of the method for identifying a headphone wearing state based on environmental acoustic features are implemented, such as Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of the modules / units in the above-mentioned device embodiments are realized, for example Figure 2 The functions of the generation module 20, the decoupling module 202, the reasoning module 203 and the output module 204 are shown.
[0073] Exemplarily, a computer program 32 for a method for identifying headphone wearing status based on ambient acoustic features primarily includes: dynamically capturing the ambient audio stream from the headphone microphone and generating an acoustic matrix with temporal correlation in real time; performing feature decoupling in the acoustic matrix processing to extract a set of physical vibration features strongly correlated with the wearing status while suppressing semantic features associated with the user's voice; constructing a lightweight temporal inference engine that receives the physical vibration feature set as input, establishing a wearing status decision boundary through multi-scale temporal convolution and attention weighting mechanisms; and outputting a classification result corresponding to the decision boundary, wherein the classification result includes identification labels for the wearing, not wearing state, and wearing direction. Computer program 32 can be divided into one or more modules / units, one or more of which are stored in memory 31 and executed by processor 30 to implement the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of computer program 32 in electronic device 3. For example, the computer program 32 can be divided into the functions of a generation module 20, a decoupling module 202, an inference module 203 and an output module 204 (modules in the virtual device), and the specific functions of each module are as follows: a generation module 201 is used to dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with time-series correlation in real time; a decoupling module 202 is used to implement feature decoupling operations in acoustic matrix processing to extract a physical vibration feature set that is strongly correlated with the wearing state, while suppressing semantic features related to the user's voice; an inference module 203 is used to construct a lightweight temporal inference engine that receives a physical vibration feature set as input, and establishes a wearing state decision boundary through multi-scale time domain convolution and attention weighting mechanism; an output module 204 is used to output the classification result corresponding to the decision boundary, wherein the classification result includes identification labels for wearing, not wearing state and wearing direction.
[0074] The electronic device 3 may include but is not limited to a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0075] The processor 30 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0076] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard drive or memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 3. Furthermore, the memory 31 can include both an internal storage unit of the electronic device 3 and an external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or is about to be output.
[0077] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0078] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0079] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0080] In the embodiments provided in this application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0081] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0082] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0083] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program of the headphone wearing state recognition method based on environmental acoustic features can be stored in a storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments, that is, dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with time-series correlation in real time; implement feature decoupling operations in the acoustic matrix processing to extract a set of physical vibration features that are strongly related to the wearing state, while suppressing semantic features related to the user's voice; construct a lightweight temporal reasoning engine that receives the physical vibration feature set as input, establishes a wearing state decision boundary through multi-scale time domain convolution and attention weighting mechanism; outputs the classification result corresponding to the decision boundary, wherein the classification result includes identification labels of wearing, not wearing state and wearing direction. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Storage media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content of storage media may be appropriately expanded or reduced based on the requirements of legislation and patent practice within a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, storage media do not include electric carrier signals and telecommunications signals.
[0084] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. The specific implementation methods described above further explain the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present invention.
Claims
1. A method for identifying earphone wearing status based on environmental acoustic characteristics, characterized in that: The method comprises: Dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with temporal correlation in real time; Implementing feature decoupling in acoustic matrix processing to extract a set of physical vibration features that are strongly correlated with the wearing state, while suppressing semantic features related to the user's voice; Constructing a lightweight temporal reasoning engine that receives the physical vibration feature set as input and establishes a wearing state decision boundary through multi-scale temporal convolution and attention weighting mechanism; Output a classification result corresponding to the decision boundary, wherein the classification result includes identification labels of wearing state, not wearing state, and wearing direction.
2. The method for identifying earphone wearing status based on environmental acoustic characteristics according to claim 1, characterized in that: The extraction of a physical vibration feature set that is strongly correlated with the wearing state includes: Divide the acoustic matrix into critical frequency subbands and calculate the time-varying energy gradient spectrum of each subband; Applying an environmental noise baseline calibration to the time-varying energy gradient spectrum to generate an interference-resistant vibration intensity representation; The anti-interference vibration intensity representations are stacked along the time dimension to construct a three-dimensional time-varying feature tensor.
3. The headphone wearing status recognition method based on environmental acoustic characteristics according to claim 1, characterized in that: The suppression of semantic features related to the user's speech includes: Construct a speech privacy barrier layer in the time-frequency feature space; In the voice privacy barrier layer, a bandpass filter bank is used to filter out the fundamental frequency component of human voice; A nonlinear projection transformation is applied to the filtered time-varying feature tensor to map the feature vector that potentially contains speech information to a subspace that is decoupled from the vibration features.
4. The headphone wearing status recognition method based on environmental acoustic characteristics according to claim 3, characterized in that: Applying a nonlinear projection transformation to the filtered time-varying feature tensor includes: Construct an autoencoder network with sparse activation characteristics, where the bottleneck layer dimension is lower than the input layer dimension by a preset threshold ratio; The filtered time-varying feature tensor is passed through the autoencoder to actively remove orthogonal bases in the encoding vector that are highly correlated with the speech content.
5. The method for identifying earphone wearing status based on environmental acoustic characteristics according to claim 1, characterized in that: It also includes an online optimization mechanism for the lightweight temporal reasoning engine: Periodically calculate the confidence variance of the state prediction results; When the confidence variance continuously exceeds the stability threshold, the incremental model update process is initiated; The updating process includes scene noise fingerprint analysis and local optimization of model parameters.
6. The method for identifying earphone wearing status based on environmental acoustic characteristics according to claim 5, characterized in that: The incremental model update process includes: collecting background noise of the environment in which the device is located in real time, extracting frequency domain fingerprint features through Fourier transform, and generating an adversarial sample set containing perturbation factors based on the fingerprint features; Freeze the network weights of the convolutional and recursive layers in the temporal inference engine, and only unlock the fully connected parameters of the output layer; The adversarial sample set is input into the unlocked model layer, and the adversarial regularization constraint term related to the noise fingerprint is injected during the back propagation process. The optimization formula is: Among them, L ce is the cross entropy loss, L adv (θ) is the adversarial loss, λ is the fingerprint correlation adjustment coefficient, is the gradient operator.
7. The method for identifying earphone wearing status based on environmental acoustic characteristics according to claim 1, characterized in that: The method comprises, in the system verification phase: A piezoelectric vibrator is used to generate micro-vibration signals when the earphones are half-detached, and accelerometer data is collected simultaneously to verify the authenticity of the vibration pattern. Inputting the micro-vibration signal into a pre-trained generative adversarial network to synthesize a set of edge scene samples in a fuzzy area of the decision boundary; The edge scenario sample set is used to perform adversarial training on the temporal reasoning engine, forcing the model to establish a clear wearing state decision boundary.
8. A device for identifying earphone wearing status based on environmental acoustic characteristics, characterized in that: The device comprises: A generation module is used to dynamically capture the ambient audio stream of the headphone microphone and generate an acoustic matrix with time-series correlation in real time; A decoupling module, configured to implement feature decoupling operations in acoustic matrix processing to extract a set of physical vibration features that are strongly correlated with the wearing state while suppressing semantic features related to the user's voice; An inference module, configured to construct a lightweight temporal inference engine that receives the physical vibration feature set as input and establishes a wearing state decision boundary through a multi-scale temporal convolution and attention weighting mechanism; An output module is used to output a classification result corresponding to the decision boundary, wherein the classification result includes identification labels of wearing, not wearing state and wearing direction.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.