Earphone wearing detection method and device

The method of comparing audio signal amplitudes in TWS earphones accurately determines the correct earpiece for audio pickup, addressing detection inaccuracies and improving call quality and user experience.

CN120321576APending Publication Date: 2025-07-15BEIJING SOUND PLUS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510506341.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing Bluetooth headset wear detection methods are insufficient in the accuracy of special-shaped headsets and non-fixed wear, resulting in inaccurate call sound pickup and poor user experience.

Method used

By comparing the amplitude difference of audio signal of left and right earphones, combining acoustic algorithms and multi-sensor fusion technology, we can judge the wearing status of the earphones and select the best pickup microphone to ensure the quality of the voice signal.

Benefits of technology

It improves the accuracy and call quality of headphone wear detection, adapts to different wearing methods, reduces hardware costs, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321576A_ABST
    Figure CN120321576A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of earphone control, and particularly provides an earphone wearing detection method and device.The earphone wearing detection method is applied to a wireless earphone, and the wireless earphone comprises a left earphone and a right earphone which are provided with pickup microphones; wherein the pickup microphone of the left earphone is used for picking up a first voice signal; the pickup microphone of the right earphone is used for picking up a second voice signal; the method comprises the following steps: when a first earphone is in a pickup scene, acquiring a first voice signal and a second voice signal; wherein the first earphone is one of the left earphone and the right earphone; calculating an amplitude difference value between the first voice signal and the second voice signal; and judging whether the wearing state of the first earphone is a worn state or not based on comparison between the amplitude difference value and a preset amplitude difference threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of headphone control, and in particular, to a method and device for detecting headphone wearing. Background Art

[0002] At present, the TWS (True Wireless Stereo) headphones on the market are composed of two independent headphones, which are connected to each other through Bluetooth signals. There is no physical cable connecting the headphones to the audio device, enabling the TWS headphones to provide a stereo sound effect while also achieving a completely wireless usage experience. For example, TWS headphones can be wirelessly connected to mobile phone devices separately for both ears for calls and voice pickup.

[0003] However, the bandwidth of the existing Bluetooth connection technology is limited. Therefore, during voice pickup such as calls, usually only one of the headphones is default used for voice pickup. Thus, there is a serious defect: when switching the voice pickup between the left and right headphones, if it is not known which headphone the user is wearing, it will cause problems in call voice pickup, resulting in the person on the other side not being able to hear or hearing unclearly.

[0004] Existing wearing detection methods generally adopt optical solutions or capacitance detection solutions, that is, an optical or capacitance sensor is arranged at the place where the headphone contacts the human body to detect whether the headphone is worn properly. However, both the optical solution and the capacitance solution have problems such as complex structure, high process requirements, and insufficient detection accuracy (see patents CN110114738B and CN107820410B). Currently, in-ear headphones and semi-in-ear headphones can have relatively reliable wearing positions, so the wearing detection is relatively accurate. However, for some special-shaped headphones, such as OWS (open wearable stereo) headphones, the contact positions with the human body are not fixed, and everyone's wearing methods are different. Moreover, many users like to wear only one headphone and put the other headphone aside. At this time, if the wrong headphone is selected for a call, it will result in a very poor user experience, which urgently needs to be solved. Summary of the Invention

[0005] In order to solve the above problems, embodiments of this application provide a method and device for detecting headphone wearing, which can accurately obtain the wearing state of the headphone, so as to use the headphone worn on the human ear for voice pickup and improve the quality of the picked-up voice signal.

[0006] To this end, the following technical solutions are adopted in the embodiments of this application:

[0007] In a first aspect, the present application provides a method for detecting headphone wearing, which is applied to wireless headphones. The wireless headphones include a left headphone and a right headphone each having a sound pickup microphone. Among them, the sound pickup microphone of the left headphone is used to pick up a first voice signal; the sound pickup microphone of the right headphone is used to pick up a second voice signal. The method includes: in a scenario where a first headphone is in a sound pickup state, obtaining the first voice signal and the second voice signal; where the first headphone is one of the left headphone and the right headphone; calculating an amplitude difference between the first voice signal and the second voice signal; and based on a comparison between the amplitude difference and a preset amplitude difference threshold, determining whether the wearing state of the first headphone is a worn state.

[0008] In this embodiment, in traditional TWS headphones, if the user wears the headphones incorrectly or only wears one headphone, it often leads to failure in voice pickup during a call. This solution can effectively solve this problem, ensure the correct use of the headphones, improve the call quality, and thus avoid user experience problems caused by improper wearing. In this embodiment, by judging the wearing state based on the amplitude difference of the audio signals, it can adapt to various headphone types, especially open wearable stereo headphones (OWS), where the contact position with the ear is not fixed. By comparing the audio signal intensities of the left and right headphones, the wearing state can be judged more accurately, regardless of the wearing method or the wearing position. Especially for some users with an unfixed headphone wearing position (such as those who like to wear only one headphone), this method can accurately judge the wearing state without relying on a specific wearing method and has strong universality. By judging the wearing state of the headphones in real time, adjustments can be made in a timely manner when the user wears them improperly to ensure the quality of the voice signal. For example, when the user only wears the right headphone, the system can automatically select the right headphone for voice pickup, avoiding misusing the left headphone and ensuring clear voice during a call.

[0009] In summary, this technical solution can effectively solve the limitations of the existing headphone wearing detection methods by judging the headphone wearing state based on the amplitude difference of the audio signals. Especially in the call scenario of TWS headphones, it can improve the accuracy of wearing detection and the intelligent level of the system. At the same time, this method has a simple structure, low cost, low power consumption, and can adapt to different wearing methods, and has high practicality and universality.

[0010] As an implementable embodiment, the method further includes: in a case where the amplitude difference is greater than or equal to the amplitude difference threshold, determining the first headphone with a larger amplitude as being in a worn state; and sending the first voice signal picked up by the first headphone to a terminal device.

[0011] In this embodiment, by calculating and comparing the amplitudes of the audio signals of the left and right earphones, the system can more intelligently determine which earphone is worn on the ear. This determination method is relatively accurate and does not depend on the wearing method or position, especially suitable for users with irregular wearing or those who like to wear only one earphone. This avoids misjudgments that may occur when determining the wearing state solely based on traditional sensors such as capacitive sensors or optical sensors. After determining the worn earphone, the system will select the earphone with a stronger signal intensity for voice pickup and transmit its signal to the terminal device. This can effectively avoid the situation where the voice signal is collected by the wrong earphone when the wearing state is incorrect, resulting in poor call quality and unclear sound for the other party. For example, when the user wears only the right earphone, the system will automatically select the voice signal of the right earphone for transmission to ensure that the other party can clearly hear the sound. This solution further improves the intelligence level of the earphone system. Through the dynamic detection of the wearing state and signal intensity, the system can be adjusted in real time according to the actual wearing situation of the user, and select the appropriate earphone for voice pickup. This flexibility improves the user's interaction experience and makes the earphone more efficient in complex usage environments. Since this solution does not depend on the wearing method and contact position of the earphone, it can adapt to the usage habits of different users. Whether the earphones are worn completely or single-ear worn, the wearing state can be judged by comparing the amplitude differences of the audio signals, thus maintaining a high degree of accuracy. For voice recognition or call scenarios, the correct earphone signal is crucial. By ensuring that the system selects the signal of the worn earphone for voice transmission, the accuracy of voice recognition and the clarity of the call can be effectively improved, and the situation of unclear hearing or inaccurate recognition caused by wearing the wrong earphone can be reduced. This technical solution not only accurately judges the wearing state by comparing the amplitude difference with a preset amplitude difference threshold, but also further optimizes the selection and transmission of voice signals, enhancing the user experience. It ensures the quality of voice signals by intelligently identifying the worn earphone, avoids misoperations, and simplifies the hardware design. Generally speaking, this solution not only meets the requirements of different wearing methods, but also improves the call quality, optimizes the voice recognition experience, and reduces problems caused by improper wearing.

[0012] As an implementable embodiment, the sending the first voice signal picked up by the first earphone to the terminal device includes: performing acoustic algorithm processing on the first voice signal to obtain a target voice signal, where the acoustic algorithm is used to enhance the voice of the first voice signal; and sending the target voice signal to the terminal device.

[0013] In this embodiment, by applying acoustic algorithms for speech enhancement and noise suppression, the system can significantly improve the quality of the speech signal. The user's speech will be clearer during transmission, and background noise will be effectively suppressed, enabling the other party to hear more clearly. This is particularly important in noisy environments. The speech enhancement and noise suppression algorithms can dynamically adjust the enhancement of the speech signal according to environmental changes, ensuring that the speech quality is guaranteed whether the user is in a quiet environment or a noisy place, greatly enhancing the user experience. This solution can perform well in different usage environments. Whether it is a noisy street, a meeting room, or a quiet indoor environment, the acoustic algorithms can intelligently process and adjust the speech signal to adapt to different environmental changes and ensure the clarity of the speech signal. The processed target speech signal usually has higher clarity and lower noise, which is crucial for speech recognition systems. The speech recognition system can more accurately recognize the user's speech commands or inputs with less noise interference. Optionally, the acoustic algorithms for speech enhancement can include, but are not limited to, beam processing, noise reduction processing, equalization, dynamic amplification compression, etc. For example, through technologies such as noise suppression and echo cancellation in the acoustic algorithms, the system effectively reduces the interference of background noise on the voice call, avoiding problems such as unclear conversations, echo problems, or communication interruptions caused by signal interference, and improving the quality and fluency of the call. The acoustic algorithms can not only handle basic speech enhancement and noise suppression but also make personalized adjustments according to specific usage scenarios. For example, when the user is outdoors, the system can prioritize enhancing the speech signal while reducing environmental noise; while in a quiet indoor environment, the system can automatically reduce the degree of enhancement of the speech signal, thus better meeting the actual needs. This technical solution can significantly improve the clarity and quality of speech by applying acoustic algorithms to process the speech signal, such as speech enhancement and noise suppression, making the user's speech more prominent and real, reducing the interference of background noise on calls and speech recognition. This solution can adapt to various environments, improve the accuracy of speech recognition, optimize the call quality, and reduce the dependence on the computing power of the device, thereby enhancing the overall user experience.

[0014] As an implementable embodiment, calculating the amplitude difference between the first speech signal and the second speech signal includes: performing noise reduction algorithm processing on the first speech signal and the second speech signal respectively to obtain the noise-reduced first speech signal and the second speech signal; determining the amplitude difference between the noise-reduced first speech signal and the second speech signal.

[0015] In this embodiment, through noise reduction processing, background noise and interference can be effectively removed, and the clarity of the speech signal can be improved. Calculating the amplitude difference can accurately quantify the differences in intensity and volume between two speech signals. This difference can be used to analyze the changes in the speech signal and determine whether further processing or enhancement of a certain speech signal is required. For example, in a multi-microphone system, the amplitude difference can help evaluate the signal source and select the best speech source. The noise reduction algorithm can dynamically adapt to different noise environments and automatically optimize the noise reduction effect to ensure an ideal speech signal can be obtained in various environments. After noise reduction processing, the amplitude differences of the signal are more easily calculated and analyzed accurately. The noise-reduced speech signal can eliminate the interference of most noises, making the calculation of the amplitude difference not easily affected by environmental noise, ensuring the accuracy of the calculation. This can avoid errors caused by noise in a complex background, thereby making the detection of the headphone wearing state more accurate, and further helping the system identify the best microphone input and automatically select the optimal signal source.

[0016] As an implementable embodiment, determining the amplitude difference between the first speech signal and the second speech signal after noise reduction includes: selecting speech signals in the low-frequency band and the high-frequency band from the first speech signal and the second speech signal after noise reduction to calculate the speech amplitude difference; wherein, the high-frequency band is the frequency band with a frequency value higher than the high-frequency threshold, and the low-frequency band is the frequency band with a frequency value lower than the low-frequency threshold.

[0017] In this embodiment, the low-frequency band refers to the part with a frequency lower than a certain low-frequency threshold, usually including a frequency range of about 0 Hz to 1 - 2 kHz. The low-frequency band mainly contains the fundamental pitch of speech (such as the frequency components of vowels) and has a greater impact on the clarity of speech. The high-frequency band refers to the part with a frequency higher than a certain high-frequency threshold, usually including a frequency range above 2 kHz. The high-frequency band contains the clarity information of speech, especially the high-frequency components of consonants, which can help the listener distinguish the specific content of speech clearly. The high-frequency threshold is 1800 - 2000 Hz, for example, 2000 Hz; the low-frequency threshold is 1000 - 1200 Hz, for example, 1000 Hz. By analyzing the amplitude difference between the low-frequency and high-frequency bands respectively, the speech feature differences in different frequency bands can be captured more carefully. The low-frequency and high-frequency components of speech correspond to different information. The low-frequency mainly involves the timbre and intonation of speech, while the high-frequency contains the clarity and detailed information of speech. Calculating their amplitude differences respectively can provide more accurate signal analysis. Calculating the amplitude differences for different frequency bands can help the system more accurately identify whether the headphones worn by the user are the left headphones or the right headphones, and then facilitate the selection of the best voice source. Moreover, by processing the low-frequency and high-frequency signals separately, the system can adopt different noise reduction strategies according to the characteristics of different frequency bands. For example, low-pass filtering can be used for low-frequency noise, and high-pass filtering can be used for high-frequency noise, so as to obtain a more flexible noise reduction effect.

[0018] As an implementable embodiment, a sensor is provided on the wireless headphones, and the method further includes: when the amplitude difference is less than the amplitude difference threshold, determining whether the wireless headphones are in a non-wearing state according to the detection information of the sensor.

[0019] In this embodiment, currently, the main TWS headphone wearing detection schemes include the capacitive sensor detection scheme and the optical sensor detection scheme. The wearing state detection of the headphones in the embodiments of the present application is applicable to the capacitive sensor detection scheme, the optical sensor detection scheme, or the sensor detection scheme made by combining and improving them.

[0020] As an implementable embodiment, the sensor includes any one or more of a voice pickup unit VPU, an inertial measurement unit IMU, and a six-axis accelerometer.

[0021] In this embodiment, the information of multiple sensors (VPU (Voice Pick-Up Unit, which can also be called a bone voiceprint sensor or a voice accelerometer), IMU (Inertial Measurement Unit), and six-axis accelerometer) is combined, enabling a more comprehensive and accurate perception of whether the earphone is in a worn state. Through the joint analysis of visual information and inertial data, errors caused by misjudgment of a single sensor can be effectively avoided. By combining the detection information of the sensors when the amplitude difference is less than a set threshold, misjudgments caused by environmental noise or temporary signal fluctuations can be excluded. For example, if the noise or sound changes in the environment are small, the system will not simply judge whether the earphone is not worn based on the amplitude difference, but will refer to the state of the sensors to make a more reasonable judgment. In different usage scenarios (such as intense exercise, quiet environment, etc.), the earphone may be interfered by various external factors. Through the multi-sensor fusion of the VPU, IMU, and six-axis accelerometer, the system can better adapt to these complex environments and accurately determine whether the earphone is in a worn state. For example, during exercise, the acceleration and angle changes of the earphone will help determine whether the earphone is still being worn. This technical solution can accurately determine whether the wireless earphone is in an unworn state when the amplitude difference is less than the set threshold by combining the detection information of multiple sensors such as the voice pick-up unit VPU, inertial measurement unit IMU, and six-axis accelerometer. Compared with the traditional single-sensor method, this multi-sensor fusion method improves the accuracy and reliability of the worn state detection, reduces the misjudgment rate, enhances the user experience, and has good adaptability and intelligence.

[0022] In a second aspect, an embodiment of the present application further provides a device for detecting the worn state of an earphone, including: an acquisition module, configured to acquire a first voice signal and a second voice signal in a scenario where the first earphone is in a voice pickup state; wherein, the first earphone is one of the left earphone and the right earphone; a calculation module, configured to calculate the amplitude difference between the first voice signal and the second voice signal; a worn state determination module, configured to determine whether the worn state of the first earphone is a worn state based on the comparison between the amplitude difference and a preset amplitude difference threshold.

[0023] As an implementable embodiment, when the amplitude difference is greater than or equal to the amplitude difference threshold, the calculation module determines the first earphone with a larger amplitude as being in a worn state; the detection device further includes a transceiver module, and the transceiver module is configured to send the first voice signal picked up by the first earphone to a terminal device.

[0024] As an implementable embodiment, the computing module is further configured to perform acoustic algorithm processing on the first voice signal to obtain a target voice signal, and the acoustic algorithm is used to enhance the voice of the first voice signal; the transceiver module is configured to send the target voice signal to a terminal device.

[0025] As an implementable embodiment, the computing module is further configured to perform noise reduction algorithm processing on the first voice signal and the second voice signal to respectively obtain the first voice signal and the second voice signal after noise reduction; and determine the amplitude difference between the first voice signal and the second voice signal after noise reduction.

[0026] As an implementable embodiment, the computing module is configured to determine the amplitude difference between the first voice signal and the second voice signal after noise reduction, specifically: select voice signals in the low-frequency band and the high-frequency band from the first voice signal and the second voice signal after noise reduction to calculate the voice amplitude difference; wherein, the high-frequency band is a frequency band with a frequency value higher than the high-frequency threshold, and the low-frequency band is a frequency band with a frequency value lower than the low-frequency threshold.

[0027] As an implementable embodiment, the wearing state determination module is further configured to determine whether the wireless earphone is in a non-wearing state according to the detection information of the sensor when the amplitude difference is less than the amplitude difference threshold.

[0028] In a third aspect, an embodiment of the present application further provides an earphone, including the earphone wearing state detection device as described in the second aspect, and a left earphone and a right earphone having a pickup microphone; wherein, the pickup microphone of the left earphone is used to pick up a first voice signal; the pickup microphone of the right earphone is used to pick up a second voice signal.

[0029] In a fourth aspect, an embodiment of the present application further provides an earphone, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, the memory is coupled to the processor, and when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0030] In a fifth aspect, an embodiment of the present application further provides a computing device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, the memory is coupled to the processor, and when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0031] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by a computing device, cause the computing device to execute the methods involved in the first aspect and its possible implementations.

[0032] In a seventh aspect, an embodiment of the present application further provides a computer program product, which includes computer instructions that, when executed by a computing device, cause the computing device to execute the methods involved in the first aspect and its possible implementations.

[0033] It can be understood that the beneficial effects of the above second to sixth aspects can be referred to the relevant descriptions of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 FIG. shows an application scenario architecture diagram of a headphone wearing detection method provided by an embodiment of the present application;

[0035] Figure 2 FIG. shows a flowchart of a headphone wearing detection method provided by an embodiment of the present application;

[0036] Figure 3 FIG. shows a scenario of a headphone wearing state provided by an embodiment of the present application;

[0037] Figure 4 FIG. shows a flowchart of calculating the amplitude value of the headphone wearing state provided by an embodiment of the present application;

[0038] Figure 5 FIG. shows a flowchart of a headphone wearing detection method provided by Embodiment 1 of the present application;

[0039] Figure 6 FIG. shows a principle block diagram of a headphone wearing detection method provided by Embodiment 1 of the present application;

[0040] Figure 7 FIG. shows a structural schematic diagram of a headphone wearing detection device provided by an embodiment of the present application;

[0041] Figure 8 FIG. shows a structural schematic diagram of a wireless headphone provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0043] In the description of the embodiments of this application, words such as "exemplary", "for example", or "for illustration" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration" is intended to present the relevant concepts in a specific manner.

[0044] In the description of the embodiments of this application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.

[0045] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0046] In the description of the embodiments of this application, when referring to "some embodiments", it describes a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0047] In the description of the embodiments of this application, the terms "first / second / third, etc." or module A, module B, module C, etc. are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0048] In the description of the embodiments of this application, the reference numerals representing steps, such as S100, S200... etc., do not necessarily mean that the steps will be executed in this order. Where permitted, the order of the front and back steps can be interchanged, or they can be executed simultaneously.

[0049] Relevant terms involved in the embodiments of this application:

[0050] Signal-to-noise ratio (SNR), also known as signal-to-noise ratio, is a measure of the relative strength or power ratio between a signal and noise. In any transmission or recording process, the signal is the part of the desired information, while the noise is the unwanted additional interference or background noise. The higher the SNR, the stronger the signal relative to the noise, usually meaning a clearer and more accurate signal.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0052] In related true wireless stereo (TWS) earphone solutions, optical, capacitive or infrared sensors are mostly used for wearing detection. However, for irregular earphones such as open earphones (OWS), a single sensor is difficult to be compatible with multiple wearing methods and positions; moreover, these solutions are prone to false touch problems. For example, when the earphones are placed on the table or held in the hand, they are easily misjudged as being in the worn state. Therefore, in some related technical solutions, multiple sensors are arranged at different positions for detection to improve the detection accuracy; however, this method will obviously increase the cost; for another example, the Apple AirPods series of earphones uses biometric sensors for wearing detection. Such sensors are more expensive and cannot avoid incorrect hand-held detection. When the human hand touches the sensor part of the earphone, it will also be detected as being in the worn state.

[0053] Based on this, the embodiments of this application provide a method and device for earphone wearing detection. Among them, the detection method of the embodiments of this application is applicable to capacitive sensor detection schemes, optical sensor detection schemes, or sensor detection schemes made by combining and improving them. In other words, the detection method provided by the embodiments of this application does not require additional wearing detection devices, such as capacitive, infrared, biometric sensors, etc., but adds an acoustic algorithm based on the various hardware components already configured in the wireless earphone itself, which has a lower hardware cost.

[0054] Specifically, the detection method provided by the embodiments of the present application uses an acoustic algorithm to perform the wearing detection of the left and right ears of TWS earphones, especially in functional scenarios such as calls, recordings, and voice wake-up where voice pickup is performed. The acoustic algorithm solution is simple, low-cost, and highly reliable, and can accurately complete the wearing detection of the earphones. Furthermore, in functional scenarios such as calls, recordings, and voice wake-up where voice pickup is performed, the earphones worn by the user are used for voice pickup to improve the voice pickup quality, thereby enhancing the user experience. Moreover, the detection method provided by the embodiments of the present application has accurate recognition and is less affected by wearing factors. For example, in scenarios where the user wears earphones on both ears, picks up one earphone with one hand, or wears one earphone on one ear and holds the other earphone near the sound source (e.g., puts it to the lips) for voice pickup, after recognizing the wearing state of the wireless earphones, it can be unaffected by the wearing position of the wireless earphones and still select the voice signal with the best signal for voice pickup, thereby improving the call effect.

[0055] To more fully understand the present application, the following embodiments are given. These embodiments are used to specifically illustrate the implementation schemes of the present application and should not be construed in any way as limiting the scope of the present application.

[0056] Figure 1 The application scenario architecture diagram of a headphone wearing detection method provided by the embodiments of the present application is shown. As Figure 1 shown, in this application scenario architecture, it may include: wireless earphones 100 and a terminal device 200. The wireless earphones 100 and the terminal device 200 establish a communication connection through a wireless network.

[0057] Among them, the terminal device 200 can be an entity on the user side for receiving or transmitting signals. The terminal device 200 can be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal device, industrial control terminal device, UE unit, UE station, mobile station, remote station, remote terminal device, mobile device, wireless communication device, UE agent, or UE device, etc. The terminal device can be fixed or mobile. It should be noted that the terminal device can support at least one wireless communication technology, such as long term evolution (LTE), NR, 6th-generation (6G) mobile communication system, or next-generation wireless communication technology, etc.

[0058] For example, the terminal device can be a mobile phone, a tablet computer (pad), a desktop computer, a laptop computer, an all-in-one computer, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA), a handheld device with wireless communication capabilities, a computing device, or other processing devices connected to a wireless modem, a wearable device, a terminal device in a future mobile communication network, or a terminal device in a future evolved Public Land Mobile Network (PLMN), etc. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device. A corresponding application is arranged on the terminal device 200, and the user can interact with the wireless headset 100 through this application.

[0059] Specifically, the wireless headset 100 can be a Bluetooth headset. Of course, it can also be other headsets based on wireless transmission, such as optical conduction headsets, etc. The above Bluetooth headset can also be an ordinary Bluetooth headset, that is, a headset that plays sound through a speaker. Of course, the above Bluetooth headset can also be a bone conduction headset. The present application does not limit the sound conduction method of the wireless headset.

[0060] Optionally, the wireless earphone 100 can be an over-ear earphone, an in-ear (earbud) earphone, a headphone, a neckband earphone, or a clip-on earphone. The embodiments of the present application do not limit the form of the wireless earphone 100. Exemplarily, the wireless earphone 100 includes a left earphone 101 and a right earphone 102. Among them, the left earphone 101 and the right earphone 102 can be respectively wirelessly communicatively connected to the terminal device 200 to perform data transmission with the terminal device 200 based on the wireless communication connection. Optionally, the left earphone 101 and the right earphone 102 can also be wirelessly communicatively connected to perform data transmission between the left earphone 101 and the right earphone 102 based on this wireless communication connection. When the wireless earphone 100 is in the factory settings, any one of the left earphone 101 and the right earphone 102 can be default set as the master earphone, and the other earphone is set as the slave earphone. The master earphone can be used as the communication entity of the wireless earphone 100 with the terminal device 200, receive the control instructions sent by the terminal device 200, and control the master earphone and the slave earphone according to the control instructions. In addition, the wireless earphone 100 supports the master-slave switching operation. For example, when the left earphone 101 is default set as the master earphone and the right earphone 102 is default set as the slave earphone, performing the master-slave switching operation can switch the right earphone 102 to the master earphone and the left earphone 101 to the slave earphone.

[0061] Optionally, any one of the left earphone 101 or the right earphone 102 in the wireless earphone 100, from a hardware structure perspective, can include an earphone housing, a rechargeable battery (e.g., a lithium battery) disposed within the earphone housing, a controller, a plurality of metal contacts for connecting the battery to a charging device, a sensor, a pickup microphone, a driver unit, and a speaker assembly with a directional sound port. Among them, the driver unit includes a magnet, a voice coil, and a diaphragm, and the driver unit is used to emit sound from the directional sound port. The above-mentioned plurality of metal contacts are disposed on the outer surface of the earphone housing. Although not shown, the wireless earphone 100 can also include a charging case. It can be understood that the above-mentioned hardware devices are configured when the wireless earphone 100 leaves the factory, and the hardware configurations of the left earphone 101 and the right earphone 102 are the same.

[0062] It should be noted that the controller can be the control center of the wireless earphone 100, connecting various parts of the entire terminal device 200 (such as a mobile phone, a computer, a smart watch, etc.) through various interfaces and lines, and executing various functions of the wireless earphone 100 and processing data. Optionally, the sensor can include one or more sensors such as a capacitive sensor, a light sensor, a gyroscope, a voice pickup unit (VPU), an inertial measurement unit (IMU), and a six-axis accelerometer. The sensor can be used to collect data related to wearing detection and transmit the collected data to the controller, and the controller processes the data to determine the wearing state of the wireless earphone 100. Among them, the wearing state can include a worn state and a non-worn state. The worn state refers to the state where the left earphone or the right earphone has been worn on the user's ear; the non-worn state refers to the state where the left earphone or the right earphone is not worn on the user's ear, such as the state in the scenario of being taken out of the charging case and placed on the table, held in the hand, or put in the pocket.

[0063] In one scenario, when the wireless earphone 100 assists the terminal device 200 in making a call, the pickup microphone (which can also be called a call microphone or a transmitting microphone) is generally in a relatively fixed position. Among them, the pickup microphone refers to the microphone that receives the user's voice signal when the wireless earphone 100 is in a call state. Among them, the voice signal can be sent by the terminal device 200 to the call counterpart who is making a call with the terminal device 200. The so-called call state means that the wireless earphone 100 enables the pickup microphone to receive the user's voice signal based on the call instruction sent by the terminal device 200. The wireless earphone 100 can send the user's voice signal to the terminal device 200, so that the terminal device 200 can send the user's voice signal to the call counterpart, and receive the voice signal of the call counterpart through the terminal device 200 and play the voice signal of the call counterpart through the speaker.

[0064] In some possible implementation manners, the call microphone can be any one of a capacitive microphone, a MEMS (microelectromechanical systems) microphone, and a directional microphone. The above types of call microphones all have good sound capture and transmission capabilities. It is worth mentioning that for bone conduction earphones, the pickup microphone is the voice pickup unit VPU sensor. After the voice pickup unit VPU sensor picks up the user's voice signal, it can assist in determining the wearing state of the wireless earphone 100, because such bone conduction sensors also need to have a hard connection with the human bone to pick up the user's voice signal. That is, these bone conduction sensors can also be used to vibrate the pickup signal to determine whether the user is wearing the left earphone 101 or the right earphone 102 when speaking.

[0065] In a scenario, as a TWS earphone, the wireless earphone 100 can use the earphone first worn by the user as the main earphone and the earphone later worn by the user as the secondary earphone. For example, if it is detected that the right earphone 102 is worn on the user's ear before the left earphone 101, the right earphone 102 can be used as the main earphone and the left earphone 101 as the secondary earphone. During a call, the pickup microphone on the right earphone 102 serves as the call microphone to receive voice signals. In this embodiment, when both the left earphone 101 and the right earphone 102 are in the worn state, if the user removes one of the earphones, the other earphone still in the worn state can be the main earphone. For example, if the user first removes the right earphone 102 while the left earphone 101 remains in the worn state. Generally, at this time, the left earphone 101 still in the worn state is closer to the sound source and has a better pickup effect, so the left earphone 101 can be used as the main earphone. Correspondingly, during a call, the voice signal of the user is picked up by the pickup microphone of the left earphone 101 and sent as an uplink signal to the other party of the call to ensure the call effect and improve the user experience.

[0066] Exemplarily, the wireless network through which the wireless earphone 100 and the terminal device 200 establish a communication connection can be Bluetooth, infrared, etc. For example, a first data transmission link is established between the left earphone 101 and the terminal device 200 through the wireless network, and this wireless network can be Wi-Fi technology, Bluetooth technology, visible light communication technology, invisible light communication technology (infrared, ultraviolet communication technology), etc. Based on the first data transmission link, voice data can be transmitted between the left earphone 101 and the terminal device 200; similarly, a second data transmission link is established between the right earphone 102 and the terminal device 200 through the wireless network, and based on the second data transmission link, voice data can be transmitted between the right earphone 102 and the terminal device 200.

[0067] In one embodiment, both the left earphone 101 and the right earphone 102 of the wireless earphone 100 as a TWS earphone can be communicatively connected to the terminal device 200 (such as a mobile phone, a computer, a smart watch, etc.). In this case, there is no need to distinguish between the main earphone and the secondary earphone. The left earphone 101 and the right earphone 102 can be communicatively connected, and the left earphone 101 and the right earphone 102 are respectively communicatively connected to the terminal device 200. For example, during a call, the left earphone 101 and the right earphone 102 respectively receive the voice signals of the other party of the call through the terminal device 200, so that the left earphone 101 and the right earphone 102 can play the voice signals of the other party of the call through the built-in speakers; the microphone of the left earphone 101 or the microphone of the right earphone 102 can serve as the pickup microphone to receive the voice signals of the user and send the voice signals of the user to the other party of the call through the terminal device 200, thereby realizing a call between the user and the other party of the call.

[0068] When using headphones to make a call in some noisy environments (such as a noisy street, a speeding motorcycle, a car with music on or the window open, etc.), since the wireless headphones 100 are located at a relatively fixed position (for example, near the user's ear, in the user's hand, etc.) and are at a certain distance from the sound source (for example, near the user's lips), when the pick-up microphone of the wireless headphones 100 receives the user's voice, the sound intensity of the ambient noise received may be greater than the sound intensity of the user's voice. As a result, the signal-to-noise ratio of the received voice signal is relatively small, which in turn affects the voice effect of the headphones.

[0069] Alternatively, in some environments with a relatively low sound volume of the sound source, the pick-up microphone may also be unable to effectively collect the user's voice. For example, when the user is in a meeting place or a public rest area, the user needs to lower the speaking volume. In this case, due to the certain distance between the pick-up microphone and the sound source, the pick-up microphone may be unable to effectively collect the user's voice, which in turn affects the voice effect of the wireless headphones 100.

[0070] To this end, the embodiments of the present application provide a headphone wearing detection method, which can accurately judge the wearing state of the headphones. At the same time, in the call state of the headphones, the better voice signals picked up by the pick-up microphones on the worn headphones or non-worn headphones can be uplinked to the call counterpart, improving the user experience. In this way, in an environment with high ambient noise or low sound volume of the sound source, when the user needs to use the headphones to make a call, in a scenario where one headphone is placed on the table and the other headphone is worn on the ear, the worn headphone can be used as the main headphone for sound pickup. Or, the user can move one of the headphones to a position close to the sound source (for example, put it near the lips). The headphone can detect that the headphone in the position close to the sound source is in the non-worn state. Since the pick-up microphone is closer to the sound source, it can increase the intensity of the user's voice received by the call microphone to a certain extent, thereby increasing the signal-to-noise ratio of the voice signal received by the call microphone. Therefore, at this time, the controller still selects the microphone on this headphone as the pick-up microphone to pick up the user's voice signal, thereby enhancing the voice effect of the headphones during the call.

[0071] It should be noted that the embodiments of the present application do not limit the application scenarios of the headphone wearing detection method. For example, it may include, but is not limited to, at least one of the following scenarios: indoor call scenario, outdoor call scenario, in-vehicle call scenario. The call scenarios may include quiet call scenarios, noisy call scenarios (such as streets, shopping malls, airports, stations, construction sites, in the rain, watching a game, concert, etc. scenarios), cycling call scenarios, outdoor windy call scenarios, single-ear call scenarios, double-ear call scenarios, and other scenarios where calls can be made.

[0072] Therefore, according to the different environments where the user is located, after accurately obtaining the wearing state of the earphones by using the earphone wearing detection method in the embodiments of the present application, the call microphone is used for normal sound pickup. When there are differences in the voice signals picked up by any one of the two earphones, the more robust and higher-quality voice signal can be sent upstream to the call counterpart, reducing the possibility that the other person on the call cannot hear or hears unclearly, and ensuring the user experience.

[0073] The following will exemplarily describe various possible implementation manners of the earphone wearing detection method provided by the present application in combination with specific embodiments.

[0074] Figure 2 The flowchart of an earphone wearing detection method provided by the embodiments of the present application is shown. This earphone wearing detection method is applied to Figure 1 the wireless earphones shown in the figure. The wireless earphones include a left earphone and a right earphone each having a sound pickup microphone. Among them, the sound pickup microphone of the left earphone is used to pick up the first voice signal, and the sound pickup microphone of the right earphone is used to pick up the second voice signal. As Figure 2 shown, this earphone wearing detection method may include:

[0075] S101. In a scenario where the first earphone is in a sound pickup state, obtain the first voice signal and the second voice signal. Here, the first earphone is one of the left earphone and the right earphone.

[0076] In this step, the first earphone may enter the sound pickup scenario (i.e., the call state or the recording state) based on the call instruction sent by the terminal device. The following takes the sound pickup scenario as a voice call for exemplary illustration. For example, when the terminal device starts a call and detects the first earphone connected to the terminal device, the terminal device may send a call instruction to the first earphone to instruct the first earphone to enter the call state. In the call state, the earphone will enable the sound pickup microphone to receive the user's first voice signal, and send the user's first voice signal to the terminal device connected to the first earphone. The terminal device will send the user's first voice signal to the call counterpart, and receive and play the voice signal of the call counterpart sent by the terminal device to assist the terminal device in realizing the call with the call counterpart. It should be noted that the first voice signal and the second voice signal may refer to the voice signals of the user received by the sound pickup microphones on the left earphone or the right earphone. One microphone may be provided on the first earphone, or multiple microphones may be provided. When multiple microphones are provided on the first earphone, any one of the microphones may be determined as the sound pickup microphone. Of course, at least two of the multiple microphones may also be sound pickup microphones, or the voice signals picked up by multiple sound pickup microphones on the same first earphone may be fused to obtain the first voice signal or the second voice signal.

[0077] In an embodiment of the present application, after the wireless earphone enters a sound pickup scenario such as a call scenario, the state of the first earphone can be detected. Among them, the state of the first earphone can include a worn state and a non-worn state. Figure 3 FIG. shows a scenario of the earphone wearing state provided by the embodiment of the present application. As Figure 3 shown, the so-called worn state can be understood as the state where the first earphone is worn on the user's ear. Correspondingly, the so-called non-worn state can be understood as the state where the first earphone is not worn on the user's ear. For example, the first earphone is located in the charging case (i.e., the earphone case), the earphone is on the table, the earphone is in the user's hand, the earphone is in the user's pocket, etc. are all regarded as the wearing state of the first earphone being the non-worn state.

[0078] It is worth mentioning that after the wireless earphone enters the sound pickup scenario, the state of the first earphone can be detected at regular intervals, or when any one of the sensors configured on the wireless earphone recognizes that the wearing state of the wireless earphone has changed, the state of the first earphone is detected, so as to cope with the scenario where the user changes the worn earphone during a call.

[0079] In a possible implementation manner, by setting a timing program in the earphone controller, the earphone wearing state is detected at the trigger time point of the timing program. For example, the trigger time of the timing program is 3 seconds, that is, the earphone wearing state is detected every 3 seconds. By reasonably timing the detection of the earphone wearing state, the response efficiency of the wearing detection can be improved, and at the same time, the user experience can be enhanced.

[0080] In another possible implementation manner, according to the type of the sensor configured on the wireless earphone, signals related to the earphone sensor type are collected by the sensor of the wireless earphone, and based on the collected signals, it is determined whether the wearing state of the earphone has changed. For example, when the earphone sensor type is a capacitance sensor, the capacitance value is collected by the capacitance sensor of the earphone, and based on the change of the capacitance value, it is determined whether the wearing state of the earphone has changed. Another example is that when the earphone sensor is an optical sensor, the level signal is collected by the optical sensor of the earphone, and based on the change of the level signal, it is determined whether the wearing state of the earphone has changed. Or, when the earphone sensor type is a voice pickup unit VPU sensor, an inertial measurement unit IMU, and a six-axis accelerometer, if the data of the IMU and the six-axis accelerometer show that the user drives the wireless earphone to perform an upward or downward acceleration movement, and then the voice signal detected by the voice pickup unit VPU sensor is very weak or there is no voice signal, it means that the user has removed the first earphone and made it enter the non-worn state.

[0081] When the sensor of the earphone determines that the wearing state of the earphone has switched according to the collected signal, the earphone wearing state detection method provided in this embodiment is performed again to clarify the wearing state of the earphone. For example, switching from the non-wearing state to the worn state, or from the worn state to the non-wearing state. If the sensor of the earphone determines that the wearing state of the earphone has switched according to the collected signal.

[0082] S102. Calculate the amplitude difference between the first voice signal and the second voice signal, and based on the comparison between the amplitude difference and a preset amplitude difference threshold, determine whether the wearing state of the first earphone is the worn state.

[0083] In some embodiments, step S102 can be implemented through the following sub-steps, and the schematic flow diagram is as Figure 4 shown, including:

[0084] S1021. Perform a noise reduction algorithm process on the first voice signal and the second voice signal to respectively obtain the noise-reduced first voice signal and second voice signal.

[0085] In this step, the first voice signal and the second voice signal are respectively obtained through different pickup microphones or bone conduction sensors. These signals may contain the user's voice and noise components in the environment. To reduce the influence of background noise, it is necessary to perform noise reduction processing on the original first and second voice signals. The noise reduction algorithm usually analyzes the noise characteristics in the signal and eliminates or suppresses these noises to retain as clear a voice signal as possible.

[0086] In this embodiment, the noise reduction algorithm may use various technologies, such as spectral subtraction, Wiener filtering, blind source separation, adaptive filters, etc. The processed signal will reduce the interference of environmental noise and retain the main voice components. Which noise reduction algorithm to use for noise reduction processing of the voice signal can be selected by the user in the application in the terminal device according to the scenario used. For example, a possible noise reduction processing method is differential noise reduction. Use at least one remaining microphone on the earphone other than the pickup microphone as the noise reduction microphone. Perform a differential operation on the voice signal received by the pickup microphone and the voice signal received by the noise reduction microphone, and the obtained is the noise-reduced voice signal. In addition, other noise reduction methods can also be used to perform noise reduction on the voice signal received by the pickup microphone. For example, the above noise reduction algorithm may include, but is not limited to, one or more operations such as Acoustic Echo Canceller (AEC), Ambient Noise Suppression (ANS), Automatic Gain Control (AGC), etc., which will not be elaborated here one by one.

[0087] Before comparing the amplitudes of the collected voice signals, the left earphone and the right earphone perform noise reduction processing to make the difference between the amplitude differences more accurate, so as to further improve the accuracy of the wearing detection result. Moreover, before sending the voice signal to the terminal device, the voice signal is first subjected to noise reduction processing, and then the noise-reduced voice signal is sent to the terminal device, which can further improve the quality of the voice signal and the accuracy of voice interaction.

[0088] S1022. Determine the amplitude difference between the first voice signal and the second voice signal after noise reduction.

[0089] In this step, the first voice signal and the second voice signal after noise reduction are subjected to amplitude difference calculation. The amplitude difference refers to the difference in amplitude between two voice signals, usually calculated as the difference in the signal amplitudes at corresponding positions on the time axis. Calculating the amplitude difference can quantify the difference between the two signals through some mathematical operations (such as absolute value difference, mean square error, etc.). The amplitude difference can be used to compare the intensity difference between two signals, the consistency of the voice, etc.

[0090] In a possible implementation, from the first voice signal and the second voice signal after noise reduction, the voice signals in the low-frequency band and the high-frequency band are selected to calculate the voice amplitude difference; among them, the high-frequency band is the frequency band with a frequency value higher than the high-frequency threshold, and the low-frequency band is the frequency band with a frequency value lower than the low-frequency threshold. The high-frequency threshold is 1800 - 2000 Hz, such as 2000 Hz; the low-frequency threshold is 1000 - 1200 Hz, such as 1000 Hz. In this implementation, the amplitude difference of the high-frequency signals and the amplitude difference of the low-frequency signals in each first voice signal and the second voice signal are detected. In this way, the wearing state of the earphone is comprehensively determined, making the detection of the earphone wearing state more accurate. At the same time, in a scenario, if the amplitude of the high-frequency signal a in the voice signal a is the largest and the amplitude of the low-frequency signal c in the voice signal c is the largest, then the high-frequency signal a and the low-frequency signal c are synthesized to obtain the first voice signal. Finally, the first voice signal is sent to the terminal device to improve the call quality.

[0091] Exemplarily, in this embodiment, the calculation of the amplitude difference refers to the process of performing Fourier transform on the first voice signal and the second voice signal collected by the pick-up microphones of the left earphone and the right earphone, and calculating the signals after Fourier transform. Specifically, the first voice signal and the second voice signal can obtain the amplitude spectrum and the phase spectrum of the first voice signal and the second voice signal after Fourier transform. The amplitude spectrum is a spectrum composed of the amplitudes of each frequency point, and the phase spectrum is a spectrum composed of the phases of each frequency point. The amplitude of each frequency point (for example, the frequency point is 10 Hz) is displayed in the amplitude spectrum. In the phase spectrum, the phases of each frequency point of the voice signal are displayed.

[0092] In one embodiment, after obtaining the amplitude spectra of the first voice signal and the second voice signal, the amplitude difference between the first voice signal and the second voice signal is obtained by the absolute value difference method. Specifically, the absolute value difference is the simplest calculation method, representing the difference in amplitude corresponding to each frequency point. Assuming that the amplitudes of the first voice signal and the second voice signal at the same frequency point are S1 and S2 respectively, the amplitude difference D(S) = |S1 - S2|.

[0093] In another embodiment, after obtaining the amplitude spectra of the first voice signal and the second voice signal, the amplitude difference is calculated using the RMS value (root mean square). The RMS value, also known as the mean square error (MSE), measures the difference between signals by calculating the average of the squared differences over the entire signal period. The root mean square calculation refers to the process of adding the squares of the powers of each frequency within the target frequency band range, dividing the resulting sum by the total number of frequencies, and finally taking the square root. The total number of frequencies refers to the number of frequencies within the entire target frequency band range.

[0094] The root mean square value x rms The calculation method is as follows:

[0095]

[0096] where x i represents the amplitude of the i-th frequency point in the voice signal, and N is the total number of frequency points. Then, the root mean square values between the first voice signal and the second voice signal are subtracted to obtain the amplitude difference between the two voice signals. The amplitude difference reflects the overall difference between the two signals in the frequency domain. The smaller the value, the more similar the signals are, and the larger the value, the greater the signal difference.

[0097] In this embodiment, by determining the amplitude difference between the first voice signal and the second voice signal within the target frequency band (high-frequency domain and low-frequency domain) range, and comparing this amplitude difference with the amplitude difference threshold corresponding to the target frequency band in different wearing states of the earphone, the wearing state of the earphone can be accurately determined, improving the accuracy of earphone wearing state detection.

[0098] S1023. When the amplitude difference is greater than or equal to the amplitude difference threshold, determine that the first earphone with the larger amplitude is in the worn state.

[0099] In this step, the amplitude difference threshold is used to characterize the minimum amplitude difference in the amplitude spectrum after the first voice signal is collected by the pickup microphone when the user is wearing the right earphone and the second voice signal is collected by the pickup microphone when the user places the left earphone on the table. That is, when the amplitude difference is not lower than the minimum threshold of the preset amplitude difference threshold corresponding to the worn state, the wearing state of the earphone with the larger amplitude can be determined as the worn state.

[0100] It can be understood that the minimum threshold of the preset amplitude difference threshold can be obtained by pre-measuring through experiments. The purpose of setting the minimum threshold of the preset amplitude difference threshold is to identify whether the wearing state of the earphone is the worn state by comparing the amplitude difference with the minimum threshold of the preset amplitude difference threshold. When the earphone switches to the wearing state, it is judged whether the earphone is indeed in the in-ear state by comparing the amplitude difference with the minimum threshold of the preset amplitude difference threshold, improving the accuracy of earphone wearing state detection.

[0101] For example, assume that we have two noise-reduced voice signals S1 and S2. The amplitude values of S1 at the frequency point of 40 Hz are 0.8 and 0.5 respectively. The calculated amplitude difference between the two is 0.8 - 0.5 = 0.3, and the preset amplitude difference threshold is 0.2. This indicates that the wearing state of the earphone with the larger amplitude value S1 is the worn state.

[0102] S1024. When the amplitude difference is less than the amplitude difference threshold, determine whether the first earphone is in the non-wearing state according to the detection information of the sensor.

[0103] In this step, if the amplitude difference is less than the amplitude difference threshold, it means that the difference between the first voice signal and the second voice signal is small. At this time, there are three scenarios: in one scenario, both earphones are in the worn state; in another scenario, both earphones are in the non-worn state; in yet another scenario, one earphone is in the worn state, and the other earphone is closer to the sound source in the non-worn state. For example, the user holds the earphone near the mouth for a call. In the above scenarios, the sensor configured in the earphone at the factory can be used for auxiliary detection to clarify the wearing state of the earphone.

[0104] Optionally, the sensor may include one or more sensors such as a capacitance sensor, an optical sensor, a gyroscope, a voice pickup unit (VPU), an inertial measurement unit (IMU), a six-axis accelerometer, etc. In this embodiment, according to the type of the sensor configured in the wireless earphone, signals related to the type of the earphone sensor are collected by the sensor of the wireless earphone, and based on the collected signals, it is determined whether the wearing state of the earphone has changed. For example, when the earphone sensor type is a capacitance sensor, the capacitance value is collected by the capacitance sensor of the earphone, and based on the change in the capacitance value, it is determined whether the wearing state of the earphone has changed. Another example is when the earphone sensor is an optical sensor, the level signal is collected by the optical sensor of the earphone, and based on the change in the level signal, it is determined whether the wearing state of the earphone has changed. Or, when the earphone sensor type is a voice pickup unit (VPU) sensor, an inertial measurement unit (IMU), and a six-axis accelerometer, if the data of the IMU and the six-axis accelerometer show that the user drives the wireless earphone to perform an upward or downward acceleration movement, and then the voice signal detected by the voice pickup unit (VPU) sensor is very weak or there is no voice signal, it indicates that the user has removed the first earphone and made it enter the non-wearing state.

[0105] S1025. Send the first voice signal picked up by the first earphone to the terminal device.

[0106] In this step, after detecting the correct wearing state of the earphone, it is also necessary to send the first voice signal picked up by the first earphone to the terminal device to ensure the call quality. It should be noted that regardless of whether the wearing state of the earphone is the worn state, in this embodiment, the signal with better voice signal quality (such as the voice signal with a larger amplitude value) is uplinked to the call counterpart.

[0107] In a possible embodiment, sending the first voice signal picked up by the first earphone to the terminal device includes: performing acoustic algorithm processing on the first voice signal to obtain a target voice signal, where the acoustic algorithm is used to enhance the voice of the first voice signal; and sending the target voice signal to the terminal device.

[0108] The first earphone picks up surrounding sounds through sensors such as microphones and generates a first voice signal. This signal usually contains the user's voice and surrounding environmental noise. An acoustic algorithm is applied to the first voice signal for voice enhancement processing. Voice enhancement refers to adjusting the volume and clarity of the voice signal through algorithms to make it more prominent in human voices and environmental noise. This process optimizes parameters such as the frequency and amplitude of the signal to ensure the improvement of the voice signal quality. Optionally, the acoustic algorithms for voice enhancement may include, but are not limited to, beamforming, noise reduction, equalization, dynamic amplification compression, etc. Among them, beamforming is a technique commonly used in multi-microphone arrays. Its basic idea is to control the phase and amplitude of the audio signals received by multiple microphones, thereby enhancing the sound signals from a specific direction and reducing the noise from other directions. Specifically, among the audio signals received by multiple microphones, by adjusting the phase of each microphone signal, a new signal is synthesized. This signal is more sensitive to the sounds from a specific direction (such as the direction of the speaker) and has a stronger suppression effect on the sounds from other directions (such as environmental noise). Equalization improves the sound quality of the voice signal by adjusting the gain of different frequency bands of the audio signal. Usually, the voice signal contains more mid-frequency components. Equalization can enhance the signal strength of these frequency bands, improve the intelligibility of the voice, and at the same time suppress the noise or unnecessary components in other frequency bands. The dynamic amplification compression technology is used to adjust the dynamic range of the audio signal, that is, the difference between the minimum and maximum volumes in the voice signal. When some parts of the voice signal have too low a volume, the compressor will amplify these parts; when the voice signal has too high a volume, the compressor will reduce these parts. This can avoid the audio discomfort caused by too large a volume difference, especially during calls and recordings.

[0109] In addition, the noise reduction processing in the acoustic algorithm may also include technologies such as noise suppression and echo cancellation to remove unnecessary background noise and crosstalk, making the voice clearer. The algorithm enhances the characteristics of the target voice, improves the voice quality, and enables the user's voice to be restored more clearly and realistically on the terminal device. After being processed by the acoustic algorithm, the obtained target voice signal has improved voice quality, reduced noise, increased volume, and enhanced clarity compared to the original first voice signal. Finally, the target voice signal will be sent to terminal devices (such as mobile phones, computers, voice assistants, etc.) for voice recognition, calls, or other voice-related applications.

[0110] In this step, by applying acoustic algorithms to enhance and suppress noise in the voice signal, the clarity and quality of the voice can be significantly improved, making the user's voice more prominent and real, and reducing the interference of background noise on calls and voice recognition. This solution can adapt to various environments, improve the accuracy of voice recognition, optimize call quality, and reduce the dependence on the computing power of the device, thereby enhancing the overall user experience.

[0111] The following uses a specific embodiment to illustrate the implementation scheme of the earphone wearing detection method provided by the embodiments of the present application.

[0112] Embodiment 1

[0113] Figure 5 Fig. 8 shows a schematic flowchart of an earphone wearing detection method provided by Embodiment 1 of the present application. Figure 6 Fig. 9 shows a principle block diagram of an earphone wearing detection method provided by Embodiment 1 of the present application. The execution subject of this earphone wearing detection method can be the above-mentioned wireless earphone 100. The structure of the wireless earphone 100 can be referred to the above description and will not be elaborated here. As Figure 5 and Figure 6 shown, the earphone wearing detection method provided in Embodiment 1 may include:

[0114] S201. Determine whether the earphone is in a sound pickup / call scenario.

[0115] In this step, first, it is determined whether the current working scenario of the earphone is a "sound pickup" or "call" scenario. The sound pickup scenario usually refers to the earphone collecting ambient sound, which may be used for voice assistants, ambient noise monitoring, etc. The call scenario refers to the working mode of the earphone when used for calls or voice communications. The purpose of this step is to confirm the working state of the earphone and ensure that subsequent operations are carried out in the correct scenario. For example, if the earphone is in the music playing state, it may not be necessary to start acoustic algorithms such as sound pickup noise reduction, but if it is in the call mode, the noise reduction algorithm may need to be enabled.

[0116] S202. Activate acoustic algorithms such as sound pickup / call noise reduction for both the left and right earphones.

[0117] In this step, after determining that the earphone is in the sound pickup or call scenario, the next step is to activate acoustic algorithms such as noise reduction. Noise reduction algorithm: In the sound pickup or call scenario, the earphone usually needs to use the noise reduction algorithm to remove ambient noise and improve voice clarity. This step indicates that both the left and right earphones (i.e., the left earphone and the right earphone) need to activate these acoustic processing algorithms to ensure that the signals on both earphones are subjected to the same noise reduction processing and ensure the consistency of signal quality.

[0118] S203. Compare the amplitudes of the two earphones after being processed by the acoustic algorithm.

[0119] This step is to compare the amplitudes of the audio signals processed by two earphones through an acoustic algorithm (such as a noise reduction algorithm). Amplitude comparison: At this time, we will measure the amplitudes of the output signals of the left and right earphones and compare the amplitude differences after they are processed by the acoustic algorithm. If the amplitude differences between the two earphones are very small, it indicates that the signal processing of the two earphones is consistent, which may mean that the two earphones are in the worn state. If the amplitude differences are large, it may be due to improper wearing. It is possible that one earphone is worn on the ear while the other is not worn correctly, such as being held in the hand, placed on the table or in the pocket, etc.

[0120] Based on the comparison result of the amplitudes in step S203, the type of the worn earphone can be further analyzed.

[0121] S204. Determine whether the worn earphone is the left earphone or the right earphone.

[0122] In this step, assuming that the left and right earphones have different signal amplitude differences due to different wearing positions after being processed by the same acoustic algorithm. The system can judge whether the worn earphone is the left earphone or the right earphone according to the amplitude difference. For example, if the amplitude of the left earphone signal is significantly higher than that of the right earphone, the system can infer that the worn earphone is the left earphone, and vice versa.

[0123] S205. If the worn earphone is the left earphone, upload the sound signal processed by the left earphone through the acoustic algorithm; if the worn earphone is the right earphone, upload the signal processed by the right earphone through the acoustic algorithm.

[0124] After determining whether the worn earphone is the left earphone or the right earphone, the last step is to upload the processed sound signal according to the worn earphone. Left earphone upload: If it is determined that the worn earphone is the left earphone, then upload the signal processed by the left earphone through the acoustic algorithm. This can be used for subsequent call or speech recognition processing. Right earphone upload: If it is determined that the worn earphone is the right earphone, then upload the signal processed by the right earphone. In this way, the system can ensure that the sound signal of the correct earphone is uploaded, guarantee the voice quality, and at the same time avoid misoperation or information confusion.

[0125] The core of the earphone wearing detection method provided in this embodiment is to infer the wearing position of the earphone by comparing the amplitude differences between the two earphones and combining with the acoustic algorithm, and upload the correct audio signal. This method can effectively detect the earphone wearing situation and perform corresponding audio signal processing according to the wearing situation to ensure the voice quality and accuracy in scenarios such as calls or speech recognition.

[0126] It is understandable that the magnitudes of the sequence numbers of the steps in the above various embodiments do not indicate the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to the actual situation, can be partially executed, or can be all executed, which is not limited herein. Additionally, all or part of any feature of any of the above embodiments can be freely combined arbitrarily on the premise of not being contradictory; the combined technical solution is also within the scope of the present application.

[0127] As Figure 7 shown, an embodiment of the present application further provides a headphone wearing detection device 700. The headphone wearing state detection device 700 can be specifically integrated in a wireless headphone, and includes an acquisition module 701, a calculation module 702, and a wearing state determination module 703: Among them, the acquisition module 701 is used to acquire a first voice signal and a second voice signal when the first headphone is in a voice pickup scenario; where the first headphone is one of the left headphone and the right headphone; the calculation module 702 is used to calculate the amplitude difference between the first voice signal and the second voice signal; the wearing state determination module 703 is used to determine whether the wearing state of the first headphone is a worn state based on the comparison between the amplitude difference and a preset amplitude difference threshold.

[0128] In a possible implementation manner, when the amplitude difference is greater than or equal to the amplitude difference threshold, the calculation module 702 determines that the first headphone with a larger amplitude is in a worn state; the detection device further includes a transceiver module, and the transceiver module is used to send the first voice signal picked up by the first headphone to a terminal device.

[0129] In a possible implementation manner, the calculation module 702 is further used to perform an acoustic algorithm process on the first voice signal to obtain a target voice signal, and the acoustic algorithm is used to enhance the voice of the first voice signal; the transceiver module is used to send the target voice signal to the terminal device.

[0130] In a possible implementation manner, the calculation module 702 is further used to perform a noise reduction algorithm process on the first voice signal and the second voice signal to respectively obtain a noise-reduced first voice signal and a noise-reduced second voice signal; and determine the amplitude difference between the noise-reduced first voice signal and the noise-reduced second voice signal.

[0131] In a possible implementation manner, the calculation module 702 is used to determine the amplitude difference between the noise-reduced first voice signal and the noise-reduced second voice signal, and specifically used for: selecting the voice signals in the low-frequency band and the high-frequency band from the noise-reduced first voice signal and the noise-reduced second voice signal to calculate the voice amplitude difference; where the high-frequency band is the frequency band with a frequency value higher than the high-frequency threshold, and the low-frequency band is the frequency band with a frequency value lower than the low-frequency threshold.

[0132] In a possible implementation, the wearing state determination module 703 is further configured to determine whether the wireless earphone is in a non-wearing state according to the detection information of the sensor when the amplitude difference is less than the amplitude difference threshold.

[0133] It should be understood that the acquisition module 701, the calculation module 702, and the wearing state determination module 703 can all be implemented by software or by hardware. Exemplarily, taking the acquisition module 701 as an example, the implementation manner of the acquisition module 701 will be introduced below. Similarly, the implementation manners of the calculation module 702 and the wearing state determination module 703 can refer to the implementation manner of the acquisition module 701.

[0134] As an example of a software functional unit, the acquisition module 701 may include code running on a computing instance. Wherein, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the acquisition module 701 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically close data centers. Wherein, generally one region may include multiple AZs.

[0135] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Wherein, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0136] As an example of a hardware functional unit, the acquisition module 701 may include at least one computing device, such as a server, etc. Or, the acquisition module 701 may also be a device implemented by using an application specific integrated circuit ASIC, or a programmable logic device PLD, etc. Wherein, the above PLD may be implemented by CPLD, FPGA, GAL, or any combination thereof.

[0137] The multiple computing devices included in the acquisition module 701 can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 701 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the acquisition module 701 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0138] The device can convert log data in different formats into a unified standard format for further analysis and processing. The log standardization device can interface with a variety of security devices, including but not limited to: firewalls, web application protection systems (WAFs), endpoint detection and response (EDR), network detection and response (NDR), intrusion prevention systems (IPSs), antivirus software, and other network security devices. The log data generated by these devices usually has different formats and structures. Through the headphone wearing detection device provided in the embodiments of the present application, these heterogeneous data can be converted into a unified standard format. For example, large operators have WAF devices from multiple manufacturers. Through log standardization, all WAF logs can be standardized and stored in a single table for convenient unified management and query.

[0139] Based on the same inventive concept, for the principle and beneficial effects of the headphone wearing detection device provided in the embodiments of the present application in solving problems, reference can be made to the principle and beneficial effects of the method implementation. For the sake of brevity, it will not be elaborated here.

[0140] Based on the above, referring to Figure 8 , the embodiments of the present application further provide a wireless headphone 100, including: a bus 101, a processor 102, a memory 103, and a communication interface 104. The processor 102, the memory 103, and the communication interface 104 communicate with each other through the bus 101. The wireless headphone 100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the wireless headphone 100.

[0141] The bus 101 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8It is represented by only one line in the figure, but it does not mean that there is only one bus or one type of bus. The bus 101 may include a path for transmitting information between various components of the test device 10 (for example, the memory 103, the processor 102, and the communication interface 104).

[0142] The processor 102 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0143] The memory 103 may include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0144] The memory 103 stores executable program instructions, and the processor 102 executes the executable program instructions to respectively implement the IP address configuration method involved in the above embodiments.

[0145] The communication interface 103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the wireless earphone 100 and other devices or a communication network.

[0146] In some possible implementation manners, the wireless earphone 100 and the server to be configured may be connected through a network. Among them, the network may be a wide area network or a local area network, etc.

[0147] The embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium is used to store computer program instructions. When the computer program instructions run on the test device, the test device is caused to execute the earphone wearing detection method involved in the above embodiments. The computer-readable storage medium may be a magnetic medium (for example, a floppy disk, a hard disk, or a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state drive), etc.

[0148] The embodiment of the present application further provides a computer program product including instructions. When the computer program product runs on a computing device, the computing device is caused to execute the earphone wearing detection method involved in the above embodiments.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application. Those of ordinary skill in the art should understand that although the present application has been described in detail with reference to the foregoing embodiments, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions in the various embodiments of the present application.

Claims

1. A method for detecting headphone wearing, characterized in that Applied to wireless earphones, the wireless earphones include a left earphone and a right earphone each having a sound pickup microphone; wherein, the sound pickup microphone of the left earphone is used to pick up a first voice signal; the sound pickup microphone of the right earphone is used to pick up a second voice signal; The method includes: In a scenario where the first earphone is in a sound pickup state, obtaining the first voice signal and the second voice signal; wherein, the first earphone is one of the left earphone and the right earphone; Calculating an amplitude difference between the first voice signal and the second voice signal; Based on a comparison between the amplitude difference and a preset amplitude difference threshold, determining whether the wearing state of the first earphone is a worn state.

2. The method according to claim 1, characterized in that, The method further includes: In a case where the amplitude difference is greater than or equal to the amplitude difference threshold, determining the first earphone with a larger amplitude as being in a worn state; Sending the first voice signal picked up by the first earphone to a terminal device.

3. The method according to claim 2, characterized in that, The sending the first voice signal picked up by the first earphone to a terminal device includes: Performing an acoustic algorithm process on the first voice signal to obtain a target voice signal, where the acoustic algorithm is used to enhance the voice of the first voice signal; Sending the target voice signal to the terminal device.

4. The method according to any one of claims 1 to 3, characterized in that The calculating the amplitude difference between the first voice signal and the second voice signal includes: Performing a noise reduction algorithm process on the first voice signal and the second voice signal to respectively obtain a noise-reduced first voice signal and a noise-reduced second voice signal; Determining the amplitude difference between the noise-reduced first voice signal and the noise-reduced second voice signal.

5. The method according to claim 4, characterized in that, The determining the amplitude difference between the noise-reduced first voice signal and the noise-reduced second voice signal includes: From the noise-reduced first voice signal and the noise-reduced second voice signal, selecting voice signals in a low frequency band and a high frequency band to calculate a voice amplitude difference; wherein, the high frequency band is a frequency band with a frequency value higher than a high frequency threshold, and the low frequency band is a frequency band with a frequency value lower than a low frequency threshold.

6. The method according to any one of claims 1 to 3, characterized in that, A sensor is provided on the wireless earphone, and the method further includes: In a case where the amplitude difference is less than the amplitude difference threshold, determining whether the wireless earphone is in a non-worn state according to the detection information of the sensor.

7. The method according to claim 6, characterized in that, The sensor includes any one or more of a voice pickup unit VPU, an inertial measurement unit IMU, and a six-axis accelerometer.

8. An earphone wearing state detection device, characterized in that Includes: An acquisition module, in a scenario where the first earphone is in a sound pickup state, for acquiring the first voice signal and the second voice signal; wherein, the first earphone is one of the left earphone and the right earphone; A calculation module, for calculating the amplitude difference between the first voice signal and the second voice signal; A wearing state determination module, for determining whether the wearing state of the first earphone is a worn state based on a comparison between the amplitude difference and a preset amplitude difference threshold.

9. A headphone, characterized in that, Includes the earphone wearing state detection device as described in claim 8 and a left earphone and a right earphone each having a sound pickup microphone; wherein, the sound pickup microphone of the left earphone is used to pick up a first voice signal; the sound pickup microphone of the right earphone is used to pick up a second voice signal.

10. A headset, characterized in that, Includes: At least one memory, for storing a program; At least one processor for executing a program stored in the memory; wherein the memory is coupled to the processor, and when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Wearable device wearing status detection method, device and wearable device

    CN107820410B

  • Wearable devices, wear detection methods and storage media

    CN110114738B