Earphone control method and related system, storage medium
By detecting ambient signals around the headphones and using Fourier transform and frequency domain signal analysis to recognize user gestures, the problem of inconsistent noise cancellation switching operations in true wireless stereo Bluetooth headphones has been solved. This enables simple and efficient headphone mode control without the need for additional hardware, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-12-16
- Publication Date
- 2026-04-28
AI Technical Summary
The noise cancellation switching operation of existing true wireless stereo Bluetooth earbuds is inconsistent, has a high learning cost, and the touch-based operation is inefficient, prone to misoperation, and results in a poor user experience.
By detecting ambient signals around the headphones, such as audio signals, Bluetooth signals, light signals, or ultrasonic signals, and using Fourier transform and frequency domain signal analysis, the system can identify the user's gesture information to adjust the headphone mode, such as covering the ears or listening, and thus control the noise cancellation function to be turned on or off.
It enables simple control of headphone mode without the need for additional hardware, improving user experience, reducing the risk of accidental operation, providing natural and user-friendly operation, and saving power.
Smart Images

Figure CN116266893B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic device technology, and in particular to a headphone control method and related system, and storage medium. Background Technology
[0002] With the rapid development of smartphones, people are increasingly dissatisfied with the lack of flexibility and convenience of wired headphones. Driven by technological innovation, Bluetooth has once again revolutionized how users connect, and major brands are expanding their Bluetooth headphone market and influence. Noise cancellation is a common feature in headphones. Noise-canceling headphones incorporate dedicated noise-canceling circuitry. Currently, there are two types of noise-canceling headphones: active noise-canceling headphones and passive noise-canceling headphones. Active noise-canceling headphones generate sound waves that are the opposite of external noise through a noise-canceling system, neutralizing the noise and achieving a noise-canceling effect. Passive noise-canceling headphones primarily block external noise by creating a closed space around the ears or by using sound-insulating materials such as silicone earplugs.
[0003] The noise-canceling switching operations of existing True Wireless Stereo (TWS) Bluetooth headphones on the market are inconsistent, have a high learning curve, and are inefficient and have a stethoscope effect due to their touch-based operation.
[0004] Currently, AirPods Pro and Freebuds Pro, which support noise cancellation, use a press-and-hold motion on the ear stem to activate and deactivate noise cancellation, with symmetrical operation for both earbuds. Galaxy Buds Pro / Live, on the other hand, use a long press on the earbud surface to activate and deactivate noise cancellation, also with symmetrical operation for both earbuds. Both of these methods require specialized components, such as pressure sensors or capacitors, to detect the action. Furthermore, the difference between long and short presses can easily lead to accidental activation, making the interaction less user-friendly and requiring a high learning curve for users.
[0005] Another type of headphone product on the market, such as the Freebuds 3 and OPPO Enco W51, uses double-tapping the left earbud to turn noise cancellation on and off. This operation method also requires special components, and the left and right operations are asymmetrical, requiring users to memorize them. The tapping method can also cause a stethoscope effect, resulting in a poor user experience. Summary of the Invention
[0006] This application discloses a headphone control method, related system, and storage medium, which can control the noise cancellation function of headphones without adding extra hardware, resulting in a good user experience.
[0007] In a first aspect, embodiments of this application provide a headphone control method, comprising: detecting gesture information; determining a gesture performed by a user based on the gesture information, wherein the gesture information is an environmental signal collected by the headphone, and the gesture includes at least one of covering the ear to form a cavity around the headphone and listening to the ear to form an open reflective surface around the headphone; and adjusting the headphone mode to a target mode based on the gesture performed by the user.
[0008] The aforementioned methods for detecting gesture information can include continuously acquiring environmental signals, periodically acquiring environmental signals, or acquiring and analyzing environmental signals through a trigger mechanism. Obtaining gesture information by detecting changes in environmental signals can be understood as acquiring environmental signals that conform to preset standards. The determination of the user's gesture based on this gesture information is implemented based on user interaction. That is, when the user makes different gestures around the headphones, the headphones will enter different control modes.
[0009] In this embodiment, the user's gesture is determined based on environmental signals collected by the headphones, and the headphones are then adjusted to a target mode based on that gesture. This method eliminates the need for additional hardware to control the headphones, improving the user experience. Furthermore, it is simple to operate, requiring minimal user learning, and the ear-covering gesture is natural and more user-friendly.
[0010] The environmental signals include at least one of the following: audio signals, Bluetooth signals, optical signals, and ultrasonic signals.
[0011] In the first implementation, when the environmental signal is an audio signal, determining the gesture performed by the user based on the gesture information includes: performing a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculating the ratio of the average energy of the frequency domain signal within a first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and if the ratio is greater than a first preset threshold, determining that the user has covered their ears.
[0012] This embodiment determines the user's gesture by comparing the energy information of a preset frequency band in an audio signal with a threshold. The ambient sound generated by the cavity formed by the hand and ear exhibits a relatively concentrated energy on a characteristic frequency band. Therefore, based on this audio feature, the user's action of covering their ears is detected, thereby controlling the headphones and improving the user experience.
[0013] Based on the first implementation, if the ratio is less than a second preset threshold, it is determined that the user has performed listening, wherein the second preset threshold is less than the first preset threshold.
[0014] As a second implementation, when the environmental signal is an audio signal, determining the gesture performed by the user based on the gesture information includes: performing a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculating the similarity value between the frequency energy distribution of the frequency domain signal and a preset frequency energy distribution; and determining that the user has covered their ears if the similarity value is greater than a second preset threshold.
[0015] In this embodiment, the similarity value between the frequency energy distribution of the frequency domain signal obtained from the audio signal around the headphones and the preset frequency energy distribution is used to determine the gesture performed by the user, and then the mode of the headphones is adjusted to the target mode.
[0016] As a third implementation, when the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, determining the gesture performed by the user based on the gesture information includes: performing Fourier transforms on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; calculating a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, and the second audio signal... The fourth ratio of the average energy of the frequency domain signal of the first audio signal in the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain range, wherein the frequency point of the second preset frequency domain range is lower than the frequency point of the fourth preset frequency domain range, and the frequency point of the third preset frequency domain range is higher than the frequency point of the fourth preset frequency domain range; when the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal in the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain range, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal in the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain range, it is determined that the user has covered their ears.
[0017] This embodiment determines the user's gesture based on the frequency energy distribution of the frequency domain signals obtained from two audio signals at different times around the headphones, and then adjusts the headphones to the target mode. It does not require additional parameter settings for judgment, avoiding misjudgments caused by unreasonable parameter settings and improving the reliability of headphone control.
[0018] As a fourth implementation, when the environmental signal is a second audio signal, determining the user's gesture based on the gesture information includes: sending the second audio signal to a computing unit, so that the computing unit performs Fourier transforms on the second audio signal and the first audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal, wherein the first audio signal and the second audio signal are collected by different headphones; receiving information sent by the computing unit, and determining that the user has covered their ears based on the information, wherein the information indicates that a first ratio is greater than a second ratio, and a third ratio is less than a fourth ratio, and / or, the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the first ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, and the second ratio is the ratio of the average energy of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The first audio signal is given by the following ratios: the average energy of its frequency domain signal within the second preset frequency domain range to the average energy of its frequency domain signal across the entire frequency domain; the third ratio is given by the following ratios: the average energy of its frequency domain signal within the third preset frequency domain range to the average energy of its frequency domain signal across the entire frequency domain; the fourth ratio is given by the following ratios: the average energy of its frequency domain signal within the fourth preset frequency domain range to the average energy of its frequency domain signal across the entire frequency domain; the fifth ratio is given by the following ratios: the average energy of its frequency domain signal within the third preset frequency domain range to the average energy of its frequency domain signal across the entire frequency domain; and the sixth ratio is given by the following ratios: the average energy of its frequency domain signal within the fourth preset frequency domain range to the average energy of its frequency domain signal across the entire frequency domain; wherein the frequency points of the second preset frequency domain range are lower than the frequency points of the fourth preset frequency domain range, and the frequency points of the third preset frequency domain range are higher than the frequency points of the fourth preset frequency domain range.
[0019] The aforementioned computing unit can be a computing unit in another earphone, or other smart devices, such as mobile phones. This solution does not specifically limit this.
[0020] The above method determines the user's gesture by analyzing the frequency energy distribution of the frequency domain signal obtained from the audio signals surrounding the two headphones, and then adjusts the headphone mode to the target mode. This eliminates the need for additional parameter settings, avoiding misjudgments caused by improper parameter settings and improving the reliability of headphone control.
[0021] Based on the above implementation, if the ratio of the average energy of the frequency domain signals of the first audio signal and the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is not higher than the third preset threshold, or if the ratio of the average energy of the frequency domain signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the ratio of the average energy of the frequency domain signal in the seventh preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, the Bluetooth signal of the earphone is obtained, and the time intensity distribution of the Bluetooth signal is obtained; if the intensity of the Bluetooth signal in the first time period is lower than the intensity of the Bluetooth signal in the second time period, and the intensity of the Bluetooth signal in the first time period is lower than the fourth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0022] This embodiment further combines Bluetooth signals to determine the user's gesture when the ear-covering gesture cannot be determined based on the audio signal.
[0023] Based on the above implementation, if the ratio of the average energy of the second audio signal in the eighth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the fourth preset threshold, and the ratio of the average energy of the second audio signal in the ninth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the fourth preset threshold, it is determined that the user has performed listening.
[0024] The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
[0025] In this embodiment, in relatively noisy scenarios, the energy peak in the frequency energy distribution of the frequency domain signal can be combined to determine whether the user has performed the gesture of covering their ears, thereby improving the accuracy and reliability of gesture determination.
[0026] Furthermore, the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. The difference between the seventh ratio of the average energy value of the first frequency band to the average energy value of the second frequency band and the eighth ratio of the average energy value of the second frequency band to the average energy value of the third frequency band is greater than a fifth preset threshold.
[0027] In this embodiment, in relatively noisy scenarios, the energy distribution of three consecutive frequency bands can be combined to determine whether the user has performed the gesture of covering their ears, thereby improving the accuracy and reliability of gesture determination.
[0028] As a fifth implementation, when the environmental signal is a Bluetooth signal, determining the gesture performed by the user based on the gesture information includes: obtaining the time intensity distribution of the Bluetooth signal based on the Bluetooth signal; if the intensity of the Bluetooth signal in the first time period in the time intensity distribution is lower than the intensity of the Bluetooth signal in the second time period, and the intensity of the Bluetooth signal in the first time period is lower than a sixth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0029] In this embodiment, the user's gesture is determined based on the Bluetooth signals around the headphones, and the headphone mode is adjusted to the target mode.
[0030] As an optional implementation, when the user's gesture is to cover their ears, the target mode is either noise cancellation on or noise cancellation pass-through off; or, when the current mode of the headphones is noise cancellation pass-through off, the target mode is noise cancellation on; or, when the current mode of the headphones is noise cancellation on, the target mode is pass-through on; or, when the current mode of the headphones is pass-through on, the target mode is noise cancellation pass-through off.
[0031] This approach increases the versatility of headphone mode control and improves the user experience.
[0032] As another optional implementation, when the user's gesture is to cover their ears, the target mode is noise reduction enabled; when the user's gesture is to listen, the target mode is pass-through enabled.
[0033] As an optional implementation, the method further includes: starting to detect gesture information when the strength value of the Bluetooth signal of the headset is lower than a seventh preset threshold.
[0034] This triggering mechanism saves on headphone battery consumption and improves the user experience.
[0035] As another optional implementation, the method further includes: when a preset signal is received, starting to detect gesture information, wherein the preset signal indicates that the wearable device detects that the user has raised their hand.
[0036] This triggering mechanism saves on headphone battery consumption and improves the user experience.
[0037] Secondly, embodiments of this application provide a headphone control method, including: acquiring environmental signals around the headphone and extracting features from the environmental signals; adjusting the headphone mode to a target mode based on the energy intensity of a preset frequency band in the extracted environmental signal features.
[0038] In this embodiment, features are extracted from the environmental signals surrounding the headphones, and the headphone mode is adjusted based on these extracted features. This method allows for implementation without adding additional sensors to existing headphones, simplifies user operation, reduces the need for extensive learning, and improves the user experience.
[0039] The environmental signals include at least one of the following: audio signals, light signals, and ultrasonic signals.
[0040] As one implementation, when the environmental signal is an audio signal, the feature extraction of the environmental signal includes: performing a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculating the ratio of the average energy of the frequency domain signal within a first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and adjusting the mode of the headphones to a target mode based on the energy intensity of a preset frequency band in the extracted environmental signal features, including: if the ratio is greater than a first preset threshold, the target mode is a first mode.
[0041] If the ratio is less than a preset threshold A, the target mode is the second mode, wherein the preset threshold A is less than the first preset threshold.
[0042] As another implementation, when the environmental signal is an audio signal, the feature extraction of the environmental signal includes: performing a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculating the similarity value between the frequency energy distribution of the frequency domain signal and a preset frequency energy distribution; and adjusting the mode of the headphones to a target mode according to the energy intensity of the preset frequency band in the extracted environmental signal features, including: if the similarity value is greater than a second preset threshold, the target mode is a first mode.
[0043] As another implementation, when the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the feature extraction of the environmental signal includes: performing Fourier transform on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; calculating a first ratio of the average energy of the frequency domain signal of the second audio signal in a second preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, a second ratio of the average energy of the frequency domain signal of the first audio signal in the second preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, a third ratio of the average energy of the frequency domain signal of the second audio signal in a third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and a third ratio of the average energy of the frequency domain signal of the second audio signal in a fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain. The fourth ratio of the average energy of the frequency domain signal in the first audio signal to the average energy of the frequency domain signal in the third preset frequency domain, wherein the frequency point of the second preset frequency domain range is lower than the frequency point of the fourth preset frequency domain range, and the frequency point of the third preset frequency domain range is higher than the frequency point of the fourth preset frequency domain range; adjusting the mode of the headphones to the target mode according to the energy intensity of the preset frequency band in the extracted environmental signal features includes: when the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or, the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal in the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal in the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the target mode is the first mode.
[0044] If the ratio of the average energy of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the third preset threshold, and the ratio of the average energy of the second audio signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the third preset threshold, then the target mode is the second mode.
[0045] In one implementation, the energy peak in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
[0046] As another implementation, the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. The difference between the seventh ratio of the average energy value of the first frequency band to the average energy value of the second frequency band and the eighth ratio of the average energy value of the second frequency band to the average energy value of the third frequency band is greater than a fourth preset threshold.
[0047] As one implementation, when the target mode is the first mode, the first mode is either the noise cancellation on mode or the noise cancellation pass-through off mode; or, when the current mode of the headphones is the noise cancellation pass-through off mode, the first mode is the noise cancellation on mode; or, when the current mode of the headphones is the noise cancellation on mode, the first mode is the pass-through on mode; or, when the current mode of the headphones is the pass-through on mode, the first mode is the noise cancellation pass-through off mode.
[0048] As another implementation, when the target mode is the first mode, the first target mode is the noise reduction enabled mode; when the target mode is the second mode, the second mode is the pass-through enabled mode.
[0049] As one implementation, the method further includes: when the strength value of the Bluetooth signal of the earphone is lower than a fifth preset threshold, starting to detect gesture information.
[0050] As another implementation, the method further includes: when a preset signal is received, starting to detect gesture information, wherein the preset signal indicates that the wearable device detects that the user has raised their hand.
[0051] Thirdly, this application provides an earphone control device, comprising: a detection module for detecting gesture information; a signal processing module for determining a gesture performed by a user based on the gesture information, wherein the gesture information is an environmental signal collected by the earphone, and the gesture includes at least one of covering the ear to form a cavity around the earphone and listening to the ear to form an open reflective surface around the earphone; and a noise reduction control module for adjusting the mode of the earphone to a target mode based on the gesture performed by the user.
[0052] The environmental signals include at least one of the following: audio signals, Bluetooth signals, optical signals, and ultrasonic signals.
[0053] In the first implementation, when the environmental signal is an audio signal, the signal processing module is used to: perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculate the ratio of the average energy of the frequency domain signal within a first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; if the ratio is greater than a first preset threshold, determine that the user has covered their ears.
[0054] As a second implementation, when the environmental signal is an audio signal, the signal processing module is used to: perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculate the similarity value between the frequency energy distribution of the frequency domain signal and a preset frequency energy distribution; and if the similarity value is greater than a second preset threshold, determine that the user has covered their ears.
[0055] As a third implementation, when the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the signal processing module is configured to: perform Fourier transform on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; calculate a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, and the frequency domain signal of the second audio signal... The fourth ratio of the average energy value of the first audio signal within the fourth preset frequency domain range to the average energy value of the frequency domain signal across the entire frequency domain is determined, wherein the frequency point of the second preset frequency domain range is lower than the frequency point of the fourth preset frequency domain range, and the frequency point of the third preset frequency domain range is higher than the frequency point of the fourth preset frequency domain range; when the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy value of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy value of the frequency domain signal across the entire frequency domain, and the sixth ratio is the ratio of the average energy value of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy value of the frequency domain signal across the entire frequency domain, thus determining that the user has covered their ears.
[0056] As a fourth implementation, the device further includes a communication module. When the environmental signal is a second audio signal, the communication module sends the second audio signal to the computing unit, so that the computing unit performs Fourier transforms on the second audio signal and the first audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal, wherein the first audio signal and the second audio signal are collected by different headphones. The signal processing module receives information sent by the computing unit and determines, based on the information, that the user has covered their ears. The information indicates that a first ratio is greater than a second ratio, and a third ratio is less than a fourth ratio, and / or, the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the first ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, and the second ratio... The first audio signal is given by the following ratios: a first ratio is the average energy of the frequency domain signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio is the average energy of the frequency domain signal within the third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio is the average energy of the frequency domain signal within the fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a fourth ratio is the average energy of the frequency domain signal within the fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a fifth ratio is the average energy of the frequency domain signal within the third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a sixth ratio is the average energy of the frequency domain signal within the fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range.
[0057] If the ratio of the average energy of the frequency domain signals of the first audio signal and the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is not higher than the third preset threshold, or the ratio of the average energy of the frequency domain signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the ratio of the average energy of the frequency domain signal in the seventh preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, the Bluetooth signal of the earphone is acquired, and the time intensity distribution of the Bluetooth signal is obtained; if the intensity of the Bluetooth signal in the first time period is lower than the intensity of the Bluetooth signal in the second time period, and the intensity of the Bluetooth signal in the first time period is lower than the fourth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0058] As another implementation, if the ratio of the average energy of the second audio signal in the eighth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the fourth preset threshold, and the ratio of the average energy of the second audio signal in the ninth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the fourth preset threshold, it is determined that the user has performed listening.
[0059] As an optional implementation, the energy peak in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
[0060] As another optional implementation, the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band, wherein the last frequency point of the first frequency band is the first frequency point of the second frequency band, the last frequency point of the second frequency band is the first frequency point of the third frequency band, and the difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than a fifth preset threshold.
[0061] As an alternative, when the environmental signal is a Bluetooth signal, the signal processing module is configured to: obtain the time intensity distribution of the Bluetooth signal based on the Bluetooth signal; if the intensity of the Bluetooth signal in the first time period in the time intensity distribution is lower than the intensity of the Bluetooth signal in the second time period, and the intensity of the Bluetooth signal in the first time period is lower than a sixth preset threshold within a preset duration, determine that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0062] Specifically, when the user's gesture is to cover their ears, the target mode is either noise cancellation on or noise cancellation pass-through off; or, when the current mode of the headphones is noise cancellation pass-through off, the target mode is noise cancellation on; or, when the current mode of the headphones is noise cancellation on, the target mode is pass-through on; or, when the current mode of the headphones is pass-through on, the target mode is noise cancellation pass-through off.
[0063] As an alternative, when the user's gesture is to cover their ears, the target mode is noise reduction enabled; when the user's gesture is to listen, the target mode is pass-through enabled.
[0064] Optionally, the device further includes a trigger module for: starting to detect gesture information when the strength value of the Bluetooth signal of the earphone is lower than a seventh preset threshold.
[0065] As an alternative, the device also includes a trigger module for: starting to detect gesture information when a preset signal is received, the preset signal indicating that the wearable device has detected that the user has raised their hand.
[0066] Fourthly, this application provides a headphone control device, comprising: a signal acquisition module for acquiring environmental signals around the headphone; a signal processing module for extracting features from the environmental signals; and a noise reduction control module for adjusting the headphone mode to a target mode based on the energy intensity of a preset frequency band in the extracted environmental signal features.
[0067] The environmental signals include at least one of the following: audio signals, light signals, and ultrasonic signals.
[0068] As one implementation, when the environmental signal is an audio signal, the signal processing module is used to: perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculate the ratio of the average energy of the frequency domain signal within a first preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain; the noise reduction control module is used to: if the ratio is greater than a first preset threshold, the target mode is the first mode.
[0069] As another optional implementation, when the environmental signal is an audio signal, the signal processing module is used to: perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; calculate the similarity value between the frequency energy distribution of the frequency domain signal and a preset frequency energy distribution; the noise reduction control module is used to: if the similarity value is greater than a second preset threshold, the target mode is the first mode.
[0070] As another optional implementation, when the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the signal processing module is configured to: perform Fourier transform on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; calculate a first ratio of the average energy of the frequency domain signal of the second audio signal in a second preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, a second ratio of the average energy of the frequency domain signal of the first audio signal in the second preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, a third ratio of the average energy of the frequency domain signal of the second audio signal in a third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and a fourth preset ratio of the average energy of the frequency domain signal of the second audio signal in a fourth preset frequency domain range. Let there be a fourth ratio between the average energy value within the frequency domain and the average energy value of the frequency domain signal across the entire frequency domain, wherein the frequency point of the second preset frequency domain is lower than the frequency point of the fourth preset frequency domain, and the frequency point of the third preset frequency domain is higher than the frequency point of the fourth preset frequency domain; the noise reduction control module is configured to: when the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or, the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy value of the frequency domain signal of the first audio signal within the third preset frequency domain to the average energy value of the frequency domain signal across the entire frequency domain, and the sixth ratio is the ratio of the average energy value of the frequency domain signal of the first audio signal within the fourth preset frequency domain to the average energy value of the frequency domain signal across the entire frequency domain, and the target mode is the first mode.
[0071] If the ratio of the average energy of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the third preset threshold, and the ratio of the average energy of the second audio signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the third preset threshold, then the target mode is the second mode.
[0072] Optionally, the energy peak in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
[0073] Furthermore, the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. The difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than a fourth preset threshold.
[0074] As one implementation, when the target mode is the first mode, the first mode is either the noise cancellation on mode or the noise cancellation pass-through off mode; or, when the current mode of the headphones is the noise cancellation pass-through off mode, the first mode is the noise cancellation on mode; or, when the current mode of the headphones is the noise cancellation on mode, the first mode is the pass-through on mode; or, when the current mode of the headphones is the pass-through on mode, the first mode is the noise cancellation pass-through off mode.
[0075] As another implementation, when the target mode is the first mode, the first target mode is the noise reduction enabled mode; when the target mode is the second mode, the second mode is the pass-through enabled mode.
[0076] Optionally, the device further includes a trigger module for: starting to detect gesture information when the strength value of the Bluetooth signal of the earphone is lower than a fifth preset threshold.
[0077] Alternatively, the device may further include a trigger module for: starting to detect gesture information when a preset signal is received, the preset signal indicating that the wearable device has detected that the user has raised their hand.
[0078] Fifthly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform a method provided as in any possible implementation of the first aspect and / or any possible implementation of the second aspect.
[0079] Sixthly, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to perform a method provided as in any possible implementation of the first aspect and / or any possible implementation of the second aspect.
[0080] It is understood that the apparatus described in the third aspect, the apparatus described in the fourth aspect, the computer storage medium described in the fifth aspect, or the computer program product described in the sixth aspect are all used to execute the methods provided in any of the first and second aspects. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0081] The accompanying drawings used in the embodiments of this application are described below.
[0082] Figure 1a This is a schematic diagram of an ear-covering gesture provided in an embodiment of this application;
[0083] Figure 1b This is a schematic diagram of a listening gesture provided in an embodiment of this application;
[0084] Figure 1c This is a schematic diagram of the structure of an earphone provided in an embodiment of this application;
[0085] Figure 1d This is a schematic diagram of a device connected to headphones according to an embodiment of this application;
[0086] Figure 2 This is a schematic flowchart of an earphone control method provided in an embodiment of this application;
[0087] Figure 3a This is a flowchart illustrating the first headphone control method provided in the embodiments of this application;
[0088] Figure 3b This is a schematic diagram of the frequency energy distribution of a frequency domain signal provided in an embodiment of this application;
[0089] Figure 4a This is a flowchart illustrating the second headphone control method provided in the embodiments of this application;
[0090] Figure 4b This is a schematic diagram of the frequency energy distribution of a frequency domain signal provided in an embodiment of this application;
[0091] Figure 4c This is a schematic diagram of the frequency energy distribution of another frequency domain signal provided in an embodiment of this application;
[0092] Figure 4d This is a schematic diagram of the frequency energy distribution of another frequency domain signal provided in an embodiment of this application;
[0093] Figure 5a This is a flowchart illustrating the third headphone control method provided in the embodiments of this application;
[0094] Figure 5b This is a schematic diagram of the frequency energy distribution of the frequency domain signal before and after covering the ears, provided in an embodiment of this application.
[0095] Figure 5c This is another schematic diagram of the frequency energy distribution of the frequency domain signal before and after covering the ears, provided in an embodiment of this application;
[0096] Figure 6 This is a flowchart illustrating the fourth headphone control method provided in the embodiments of this application;
[0097] Figure 7a This is a flowchart illustrating the fifth headphone control method provided in this application embodiment;
[0098] Figure 7b This is a schematic diagram of the time intensity curve of a Bluetooth signal provided in an embodiment of this application;
[0099] Figure 8a This is a flowchart illustrating the sixth headphone control method provided in this application embodiment;
[0100] Figure 8b This is a schematic diagram of the energy at different frequency bands for an ear-covering operation provided in an embodiment of this application;
[0101] Figure 8c This is a schematic diagram of energy at different frequency bands when a user's ears are covered and the ambient noise is high, provided in an embodiment of this application.
[0102] Figure 9 This is a flowchart illustrating the seventh headphone control method provided in the embodiments of this application;
[0103] Figure 10 This is a flowchart illustrating another headphone control method provided in an embodiment of this application;
[0104] Figure 11 This is a schematic diagram of the structure of an earphone control device provided in an embodiment of this application;
[0105] Figure 12 This is a schematic diagram of another headphone control device provided in an embodiment of this application;
[0106] Figure 13 This is a schematic diagram of another headphone control device provided in the embodiments of this application. Detailed Implementation
[0107] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0108] Since the existing noise cancellation and other control interactions of headphones are not very user-friendly and require additional hardware, this solution proposes a method, device, and storage medium for controlling headphones. This method does not require additional hardware and can use acoustic, optical, and other detection methods to identify user operations, thereby enabling headphone control.
[0109] In this solution, the embodiments involve ear-covering gestures and listening gestures. Covering one's ears is a natural way to express resistance to external noise in the real world. Therefore, this solution uses technical means to detect this user action and triggers the noise cancellation control of the headphones based on the detected action, thereby achieving a natural and low-cost interaction method on current headphones such as Bluetooth headphones.
[0110] In this solution, covering the ears can be understood as the user wearing headphones and then using their hand to create a completely enclosed cavity around the headphones. This solution defines the action of using one's hand to create a completely enclosed cavity around the headphones as the ear-covering operation. For example... Figure 1a As shown, the user's hand has a certain convex (arched) shape, and there is a cavity between the ear and the hand.
[0111] Listening is an action people take when they can't hear clearly. In this solution, listening can be understood as the user's action of forming an open reflective surface around the headphones with their hand. This solution defines the action of forming an open reflective surface around the headphones as the listening operation. For example... Figure 1b As shown, the user's palm is placed close to the back of the ear, forming an open reflective surface.
[0112] It should be noted that this solution is only used as an example for illustration, and other gestures are also possible, which are not specifically limited in this solution.
[0113] This solution detects user gestures on the headphones, such as changes in the audio signal, to determine the user's action and then control the headphones. Specifically, the microphone in the headphones detects the characteristics of the external sound signal and, based on these characteristics or changes, detects that the user's action is covering their ears. Based on this action, the noise cancellation mode of the headphones is activated or deactivated. Alternatively, based on the headphones' previous state and the ear-covering action, the solution determines whether to activate noise cancellation, deactivate noise cancellation, or enable pass-through mode. The above explanation uses audio signals as an example only. Other modal detection signals can also be used, such as Bluetooth signals, or photosensitive sensors, proximity sensors, etc., to detect the user's specific actions. Of course, cameras can also be used for detection. This solution does not impose any specific limitations on these methods.
[0114] This solution is applicable to wired and wireless headphones, such as over-ear headphones and true wireless stereo Bluetooth headphones.
[0115] Headphones can be categorized by shape into over-ear and in-ear headphones. Over-ear headphones are worn on the head and not inserted into the ear canal, unlike in-ear headphones. Over-ear headphones are generally further divided into over-ear and on-ear headphones. Over-ear headphones have larger earcups that completely cover the ear, with ear pads pressing against the skin outside the ear. On-ear headphones have earcups that press against the ear.
[0116] Earbuds can be divided into in-ear headphones and semi-in-ear headphones. In-ear TWS earbuds have rubber tips that go deep into the ear canal, allowing for a tighter fit. In-ear headphones are generally designed like little beans, and the rubber tips come in different sizes to fit the human ear. Semi-in-ear headphones do not have rubber tips and have a long stem, making them more like hanging in the ear canal when used.
[0117] Among them, true wireless stereo Bluetooth earbuds, with their smaller size, lower latency, and good sound quality, are gradually replacing traditional wired earbuds and becoming a more convenient tool for music, watching dramas, and playing games in daily life.
[0118] Reference Figure 1c The image shown is a structural schematic diagram of an earphone provided in an embodiment of this application. Figure 1c As shown, the headset includes a detection module, a calculation module, a feedback module, a noise reduction control module, a signal processing module, a communication module, a microphone, a speaker, a Bluetooth module, and a CPU. The detection module detects user input, which can include the microphone for sound input (such as the user's voice or ambient sound); it can also include a touch sensor to receive touch input, such as clicks, double-clicks, or swipes on the headset surface. The feedback module provides feedback to the user through sound, vibration, etc. The calculation module performs internal calculations. The noise reduction control module controls the switching of noise reduction modes based on the detected information. The signal processing module processes received signal information, such as audio, Bluetooth, or light signals. Noise reduction processing can be performed within the signal processing module. The communication module exchanges control and audio data when the headset is associated with other devices. The microphone receives external audio information. The speaker transmits the processed sound from the signal processing module to the outside world.
[0119] The headphones can also incorporate various other sensor modules. For example, motion sensors such as accelerometers and gyroscopes can detect the headphones' position and orientation; optical sensors can detect when the headphones are removed from the case; touch sensors can detect finger touches on the headphone surface; and sensors such as capacitive sensors, voltage sensors, impedance sensors, photosensors, proximity sensors, and image sensors can be used for various detection purposes.
[0120] In this solution, the user's operation detected based on changes in the external audio signal can be implemented in either the detection module or the signal processing module. This detection function can be integrated into other modules such as the signal processing module. However, this solution does not impose any specific limitations on this.
[0121] like Figure 1d As shown in the illustration, this application also provides a schematic diagram of a device connected to headphones. This device may be, for example, a mobile phone, tablet computer, smart TV, etc. It is understood that these devices generally have input systems, feedback systems, detection systems, displays, computing units, storage units, and communication units common to electronic devices. In some embodiments, the signal triggering and detection process for gesture information detection in the headphone control method provided in this application is executed by the detection system of the device. For example, sensors integrated into the device can be used for detection during the detection process. It is understood that the detection system and sensors on the device connected to the headphones have the same function as the detection system and sensors installed on the headphones; they may simply be installed on different entities due to commercial or cost considerations.
[0122] The headphone control method provided in the embodiments of this application is described below. Figure 2 The image shows an embodiment of a headphone control method provided in this application, which includes steps 201-203, as detailed below:
[0123] 201. Detect gesture information;
[0124] The earphones can detect gesture information in real time. This gesture information is environmental signals collected by the earphones. Specifically, the earphones' detection module collects environmental signals around the earphones in real time.
[0125] Alternatively, the headphones can periodically detect gesture information. For example, the headphones can periodically collect ambient signals.
[0126] Alternatively, the headphones can also detect gesture information when preset trigger conditions are met. For example, when preset conditions are met, the headphones begin to collect ambient signals.
[0127] The preset condition can be that when the Bluetooth module in the earphone detects that the Bluetooth RSSI value is lower than a preset threshold, it triggers the detection of gesture information.
[0128] Alternatively, the system could detect when a user raises their hand, such as by using the IMU module in the watch. Other triggering methods are also possible, such as using a proximity sensor to detect an object approaching the headphones, or a photosensor to detect that the light intensity is below a certain threshold, thus triggering the headphones to collect ambient signals in real time, such as activating the microphone for audio recording. This solution does not specifically limit these methods.
[0129] By triggering the detection of gesture information in the headphones when preset conditions are met, the power consumption of the headphones can be saved.
[0130] 202. Determine the gesture performed by the user based on the gesture information, wherein the gesture information is an environmental signal collected by the earphone, and the gesture includes at least one of covering the ear to form a cavity around the earphone and listening to the ear to form an open reflective surface around the earphone;
[0131] The aforementioned environmental signals may be, for example, one or more of the following: audio signals, Bluetooth signals, light signals, ultrasonic signals, etc.
[0132] The aforementioned audio signal can be detected by the headphone's detection module, such as an audio signal collected over a period of time via a microphone. It can also be an audio signal collected after being triggered by other methods; this solution does not specifically limit its acquisition.
[0133] The aforementioned Bluetooth signal can be obtained through Bluetooth modules or similar means.
[0134] By extracting features from the collected environmental signals, the gestures performed by the user can be determined based on the extracted features.
[0135] The purpose of feature extraction is to detect whether the collected audio signals / Bluetooth signals / light signals / ultrasonic signals, etc., match the signal characteristics of a user covering their ears, or whether they match the signal characteristics of a user listening, so as to trigger corresponding operations in the future.
[0136] For example, the headphone's computing module determines the user's gesture based on the gesture information, and then instructs the noise cancellation control module to switch the headphone's mode.
[0137] The "ear-covering" feature in this solution can be understood as the user creating a cavity around the earphones with their hand. This cavity can be completely sealed, as described above. Figure 1b As shown.
[0138] The listening function in this solution can be understood as the user's action of forming an open reflective surface around the headphones with their hand. This solution defines the action of forming an open reflective surface around the headphones as the listening operation; please refer to the aforementioned... Figure 1c As shown.
[0139] 203. Adjust the mode of the headphones to the target mode based on the gesture performed by the user.
[0140] For example, when the system detects that the user is covering their ears, the headphones are triggered to enter noise cancellation mode. Alternatively, when the user's gesture is covering their ears, the target mode is noise cancellation pass-through off mode (normal mode).
[0141] Specifically, the dedicated noise-canceling circuitry or audio processing module in the headphones is activated to process the noise of the subsequent audio collected from the microphone.
[0142] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0143] The above is just one example of an implementation method. It can also correspond to other control modes, and this solution does not make specific limitations on them.
[0144] As an alternative implementation, after determining the user's gesture, the internal processing of the ear-covering gesture can be changed based on the current state of the earphone.
[0145] Specifically, when the user performs a gesture of covering the ears, if the current mode of the headphones is noise cancellation pass-through off mode, the target mode is noise cancellation on mode;
[0146] Alternatively, if the current mode of the headphones is noise cancellation enabled, the target mode is pass-through enabled.
[0147] Alternatively, when the current mode of the headphones is pass-through enabled, the target mode is noise cancellation pass-through disabled.
[0148] For example, if the current mode is already noise-canceling mode, the current mode will be switched to non-noise-canceling mode when the ear-covering operation is detected.
[0149] As another optional implementation, noise reduction levels can be differentiated. If the user's environment is a relatively quiet place such as a library, bookstore, or office, and the user is detected covering their ears, a light noise reduction mode is activated. If the user's environment is a moderately noisy place such as a coffee shop or subway, and the user is detected covering their ears, a balanced noise reduction mode is activated. If the user's environment is a very noisy place such as a restaurant or airport, and the user is detected covering their ears, a deep noise reduction mode is activated, and so on.
[0150] Based on the above embodiments, combinations can also be used for mode control. For example, a scheme that differentiates levels based on noise reduction level can be combined with the current state of the headphones, using the ear-covering action in conjunction with the noise reduction level currently used by the headphones to determine the specific headphone operation triggered by the ear-covering action.
[0151] For example, if the headphones are currently in light noise cancellation mode, cover your ears briefly to activate medium noise cancellation; if the headphones are currently in medium noise cancellation mode, cover your ears briefly to activate heavy noise cancellation, and so on.
[0152] In addition to controlling the headphone mode based on the scenario, the noise cancellation level can also be adjusted according to the length of time the ears are covered. For example, if the user keeps their ears covered, the noise cancellation level will gradually increase, which can be confirmed through sound feedback from the headphone's feedback module. Specifically, when the system detects that the user has just covered their ears, the headphone will beep once; if the system detects that the user is not letting go, the noise cancellation level will be increased, and another beep will be heard, or the beep sound will change. The above is just one example, and this solution does not impose any specific limitations on it.
[0153] Alternatively, the ear-covering operation can be determined based on user presets to determine the corresponding headphone mode, etc., but this solution does not make specific limitations on this.
[0154] This embodiment triggers different headphone mode switching through user gesture operation, and can realize multiple modes and state switching based on a single action. It is highly operable, more intelligent, and has a user-friendly interface.
[0155] In this embodiment, the user's gesture is determined based on environmental signals collected by the headphones, and the headphones are then adjusted to the target mode. Using this method, the ear-covering gesture is more natural, user-friendly, and simpler to operate. It requires minimal user learning and no additional hardware is needed to control the headphones, thus improving the user experience.
[0156] The specific implementation of this application will be described in detail below.
[0157] Reference Figure 3a The diagram shown is a flowchart illustrating the first headphone control method provided in this application. The method includes steps 301-305, as detailed below:
[0158] 301. Detect gesture information, wherein the gesture information is the audio signal collected by the headphones;
[0159] One implementation involves using an external microphone on the headphones to collect audio signals from the surrounding area. These audio signals can be a duration of a specific time, such as 30 seconds.
[0160] The specific duration can be determined based on the required accuracy of the calculation and the processing power of the headphones; this solution does not impose specific limitations on this.
[0161] 302. Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0162] Specifically, performing a Fourier transform on the acquired audio signal converts the received time-domain audio signal into a frequency-domain signal. The frequency energy distribution of the frequency-domain signal can be as follows: Figure 3bAs shown, the horizontal axis represents time, the vertical axis represents the frequency range, and the points in the figure represent the energy of the signal.
[0163] 303. Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain;
[0164] The frequency domain signals within the aforementioned first preset frequency domain range can be a relatively concentrated range of frequencies. For example, Figure 3b The ratio of the average energy value between the frequency range f1 and f2 in the frequency domain to the average energy value across the entire frequency band.
[0165] 304. If the ratio is greater than a first preset threshold, it is determined that the user has covered their ears;
[0166] The first preset threshold can be any value, and this solution does not impose any specific restrictions on it.
[0167] Experiments have shown that ambient sounds created by the cavity formed by the hand and ear exhibit a relatively concentrated energy in characteristic frequency bands. Therefore, this audio characteristic can be used to detect when a user covers their ears.
[0168] Furthermore, if the ratio is less than a second preset threshold, it is determined that the user has performed listening, wherein the second preset threshold is less than the first preset threshold.
[0169] Of course, the above ratio can also be greater than a third preset threshold and less than a second preset threshold to determine that the user has performed listening, etc., wherein the second preset threshold is less than the first preset threshold. This solution does not specifically limit this.
[0170] 305. Adjust the mode of the headphones to the target mode.
[0171] For example, when the system detects that the user is covering their ears, it triggers the headphones to enter noise-canceling mode. Specifically, the headphones activate a dedicated noise-canceling circuit or audio processing module to process the subsequent audio collected from the microphone.
[0172] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0173] Furthermore, after determining the user's gesture, the internal processing triggered by the ear-covering gesture can be modified based on the current state of the headphones. For example, if the current mode is already noise-canceling mode, the ear-covering gesture will be detected and the current mode will be switched to non-noise-canceling mode.
[0174] This embodiment determines whether a user has covered their ears based on the phenomenon that ambient sound caused by the cavity formed by the hand and ear will have relatively concentrated energy in a characteristic frequency band.
[0175] In this embodiment, the user's gesture is determined based on the energy information of the frequency domain signal of the audio signal collected by the headphones, and the headphone mode is then adjusted to the target mode. Using this method, the ear-covering gesture is more natural, user-friendly, and simpler to operate, requiring minimal user learning and no additional hardware, thus achieving headphone control and improving the user experience.
[0176] Reference Figure 4a The diagram shown is a flowchart illustrating a second headphone control method provided in this application. The method includes steps 401-405, as detailed below:
[0177] 401. Detect gesture information, wherein the gesture information is the audio signal collected by the headphones;
[0178] One implementation involves using an external microphone on the headphones to collect audio signals from the surrounding area. These audio signals can be a duration of a specific time, such as 30 seconds.
[0179] The specific duration can be determined based on the required accuracy of the calculation and the processing power of the headphones; this solution does not impose specific limitations on this.
[0180] 402. Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0181] Specifically, performing a Fourier transform on the acquired audio signal can convert the received time-domain audio signal into a frequency-domain signal. For details, please refer to the description in the foregoing embodiments; further elaboration will not be repeated here.
[0182] 403. Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution;
[0183] The frequency energy distribution of the aforementioned frequency domain signal can be, for example, a frequency energy curve. Figure 4b , Figure 4c , Figure 4d As shown, the horizontal axis represents frequency, and the vertical axis represents the energy at that frequency. Figure 4b The figure shows the frequency energy curve when the noise source is tilted to the side opposite to the ear being covered; Figure 4c The figure shows the frequency energy curve when the noise source is in the middle of the user's location; Figure 4d The figure shows the frequency energy curve when the noise source is tilted towards the side where the ears are covered.
[0184] The aforementioned preset frequency energy distribution can be the frequency energy distribution corresponding to the ear-covering gesture obtained through learning.
[0185] 404. If the similarity value is greater than the second preset threshold, it is determined that the user has covered their ears;
[0186] After multiple experiments, it was found that when a user covers their ears, the frequency energy curves of noise sources from different directions will show a large peak. Therefore, by matching the frequency energy curve of the detected audio signal with a pre-learned or set curve, when the similarity is greater than a certain threshold, it can be determined that the user has covered their ears.
[0187] 405. Adjust the mode of the headphones to the target mode.
[0188] For example, when the system detects that the user is covering their ears, the headphones are triggered to enter noise cancellation mode.
[0189] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0190] Furthermore, after determining the user's gesture, the internal processing triggered by the ear-covering gesture can be modified based on the current state of the headphones. For example, if the current mode is already noise-canceling, the ear-covering gesture can be detected and the current mode can be switched to non-noise-canceling mode.
[0191] In this embodiment, the similarity value between the frequency energy distribution of the frequency domain signal obtained from the audio signals around the headphones and a preset frequency energy distribution is used to determine the user's gesture, thereby adjusting the headphones to the target mode. Using this method, the ear-covering operation is more natural, user-friendly, and simpler to operate, requiring minimal user learning and no additional hardware, thus achieving headphone control and improving the user experience.
[0192] The above embodiment uses a single audio signal as an example. The following describes how to control headphones based on two audio signals. (Refer to...) Figure 5a The diagram shown is a flowchart illustrating the third headphone control method provided in this application. The method includes steps 501-505, as detailed below:
[0193] 501. Detect gesture information, wherein the gesture information is a first audio signal and a second audio signal collected by the headphones, and the first audio signal is collected earlier than the second audio signal;
[0194] For example, two audio signals around the headphones can be collected at certain time intervals.
[0195] The two audio signals mentioned above can also be acquired from two consecutive time periods; this solution does not impose specific limitations on this.
[0196] 502. Perform Fourier transform on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal;
[0197] The Fourier transform can be found in the description of the foregoing embodiments, and will not be repeated here.
[0198] 503. Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range.
[0199] 504. When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, it is determined that the user has covered their ears;
[0200] like Figure 5b The comparison of low-frequency energy changes in the two audio signals before and after covering the ears is shown, and as shown in the image. Figure 5c The image shows a comparison of high-frequency energy changes in two audio signals before and after the ears were covered. Time period 1 corresponds to the first audio signal before the ears were covered, and time period 2 corresponds to the second audio signal after the ears were covered. By comparing the energy distribution of the frequency domain signals of the two audio signals, if, relative to the first audio signal acquired earlier, the second audio signal acquired later shows a concentration of frequency energy in the low-frequency range (e.g., ...), then... Figure 5b As shown), and there is a sudden drop in energy in the high-frequency range (such as...). Figure 5c If the user covers their ears (as shown in the image), then it is determined that the user has done so.
[0201] 505. Adjust the mode of the headphones to the target mode.
[0202] For example, when the system detects that the user is covering their ears, the headphones are triggered to enter noise cancellation mode.
[0203] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0204] Furthermore, after determining the user's gesture, the internal processing triggered by the ear-covering gesture can be modified based on the current state of the headphones. For example, if the current mode is already noise-canceling mode, the ear-covering gesture will be detected and the current mode will be switched to non-noise-canceling mode.
[0205] The above example illustrates the action of a user covering their ears.
[0206] If, in the frequency domain signals of the first audio signal and the second audio signal, there exists at least one segment of the audio signal whose average energy value in the eighth preset frequency domain range is higher than the average energy value of the frequency domain signal across the entire frequency domain than a third preset threshold, and whose average energy value in the ninth preset frequency domain range is also higher than the third preset threshold, then it is determined that the user has performed listening.
[0207] In other words, if one of the two audio signals is enhanced at both high and low frequencies, it is determined that the user has performed listening.
[0208] Accordingly, the mode of the headphones can be adjusted to the target mode based on the listening gestures performed by the user.
[0209] This embodiment is based on the ratio of the relatively uniform energy distribution before covering the ears to the energy distribution after covering the ears. After covering the ears, the energy is more concentrated in a certain low-frequency region, and the energy drops sharply in the high-frequency region. This feature is used to determine the user's gesture.
[0210] In this embodiment, the user's gesture is determined by the frequency energy distribution of a frequency domain signal obtained from two audio signals at different times surrounding the headphones, thereby adjusting the headphones to the target mode. Using this method, the ear-covering gesture is more natural, user-friendly, and simpler to operate, requiring minimal user learning and no additional hardware, thus achieving headphone control and improving the user experience.
[0211] On the other hand, this solution uses the difference between two audio segments at different times to determine the user's operation, without the need for additional parameter settings. This avoids some misjudgments caused by unreasonable parameter settings and improves the reliability of headphone control.
[0212] Reference Figure 6The diagram shown is a flowchart illustrating the fourth headphone control method provided in this application. This method controls the headphones based on two audio signals from two headphones. The method includes steps 601-604, as follows:
[0213] 601. Detect gesture information, wherein the gesture information is a second audio signal collected by the headphones;
[0214] 602. The second audio signal is sent to the computing unit so that the computing unit performs Fourier transform on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal, wherein the first audio signal and the second audio signal are collected by different headphones;
[0215] The first and second audio signals mentioned above can be collected within the same time period, and this solution does not impose specific limitations on this.
[0216] The aforementioned computing unit may be located in a device such as a mobile phone connected to the earphone, or it may be located in another earphone connected to the earphone.
[0217] Optionally, the left earphone collects the first audio signal, and the right earphone collects the second audio signal.
[0218] 603. Receive information sent by the computing unit, and determine based on the information that the user has covered their ears, wherein the information indicates that a first ratio is greater than a second ratio, and a third ratio is less than a fourth ratio, and / or, the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the first ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal within the full frequency domain, the second ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal within the full frequency domain, and the third ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range. The ratio of the average energy value to the average energy value of the frequency domain signal in the full frequency domain, the fourth ratio being the ratio of the average energy value of the frequency domain signal of the second audio signal in the fourth preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, the fifth ratio being the ratio of the average energy value of the frequency domain signal of the first audio signal in the third preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, and the sixth ratio being the ratio of the average energy value of the frequency domain signal of the first audio signal in the fourth preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, wherein the frequency point of the second preset frequency domain range is lower than the frequency point of the fourth preset frequency domain range, and the frequency point of the third preset frequency domain range is higher than the frequency point of the fourth preset frequency domain range;
[0219] By comparing the frequency energy distribution of the two audio signals, if, relative to the other audio signal, the frequency energy of one audio signal is concentrated in the low-frequency range and suddenly drops in the high-frequency range, it is determined that the user has covered their ears.
[0220] Furthermore, based on the frequency energy distribution, it can be determined which ear the user is covering, and then different noise reduction operations can be performed on the left and right earphones depending on which ear is covered.
[0221] 604. Adjust the mode of the headphones to the target mode.
[0222] For example, when the system detects that the user is covering their ears, the headphones are triggered to enter noise cancellation mode.
[0223] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0224] Furthermore, after determining the user's gesture, the internal processing triggered by the ear-covering gesture can be modified based on the current state of the headphones. For example, if the current mode is already noise-canceling mode, the ear-covering gesture will be detected and the current mode will be switched to non-noise-canceling mode.
[0225] Wherein, when the ratio of the average energy of the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is not higher than the third preset threshold, or when the ratio of the average energy of the frequency domain signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the ratio of the average energy of the frequency domain signal in the seventh preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, the Bluetooth signal of the earphone is obtained, and the time intensity distribution of the Bluetooth signal is obtained;
[0226] If the strength of the Bluetooth signal in the first time period is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than the fourth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0227] In other words, when audio signals alone are insufficient to determine if a user has covered their ears, Bluetooth signals can be used as an auxiliary means of detection. Of course, other sensors such as photosensors and proximity sensors can also be employed; this solution does not impose specific limitations on these methods.
[0228] The above example illustrates the action of a user covering their ears.
[0229] If the ratio of the average energy of the second audio signal in the eighth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the fourth preset threshold, and the ratio of the average energy of the second audio signal in the ninth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the fourth preset threshold, it is determined that the user has performed listening.
[0230] Accordingly, the mode of the headphones can be adjusted to the target mode based on the listening gestures performed by the user.
[0231] In this embodiment, the user's gesture is determined by the frequency energy distribution of the frequency domain signal obtained from the audio signals around the two earphones, and the earphone mode is then adjusted to the target mode. Using this method, the ear-covering gesture is more natural, user-friendly, and simpler to operate. It requires minimal user learning and no additional hardware is needed to recognize the user's gesture, thereby controlling the earphones and improving the user experience.
[0232] On the other hand, this solution does not require additional parameter settings for judgment, avoiding some misjudgments caused by unreasonable parameter settings and improving the reliability of headphone control.
[0233] Reference Figure 7aThe diagram shown is a flowchart illustrating the fifth headphone control method provided in this application. The method includes steps 701-704, as detailed below:
[0234] 701. Detect gesture information, wherein the gesture information is the Bluetooth signal collected by the earphone;
[0235] The Bluetooth signal can be over a period of time.
[0236] 702. Obtain the time intensity distribution of the Bluetooth signal based on the Bluetooth signal;
[0237] like Figure 7b In the time-intensity curve shown, the horizontal axis represents time, and the vertical axis represents the Received Signal Strength Indication (RSSI). The RSSI value is always negative, and a decrease in this value indicates a weakening signal. Specifically, when a user covers their ears (Bluetooth headset), the Bluetooth signal weakens more rapidly and the RSSI value decreases more significantly compared to when the communication device is further away. Although RSSI also decreases when the phone and headset are far apart, this decrease is gradual and fluctuates with distance, without suddenly dropping below a certain threshold.
[0238] 703. If the strength of the Bluetooth signal in the first time period of the time intensity distribution is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than the sixth preset threshold within a preset duration, it is determined that the user has covered his ears, wherein the second time period is earlier than the first time period.
[0239] If the RSSI strength of the Bluetooth signal suddenly drops below a certain threshold and persists for a certain period of time, it is determined that the user has covered their ears.
[0240] 704. Adjust the mode of the headphones to the target mode.
[0241] For example, when the system detects that the user is covering their ears, it triggers the headphones to enter noise-canceling mode. Specifically, the headphones activate a dedicated noise-canceling circuit or audio processing module to process the subsequent audio collected from the microphone.
[0242] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0243] Furthermore, after determining the user's gesture, the internal processing triggered by the ear-covering gesture can be modified based on the current state of the headphones. For example, if the current mode is already noise-canceling mode, the ear-covering gesture will be detected and the current mode will be switched to non-noise-canceling mode.
[0244] In this embodiment, the user's gesture is determined based on the Bluetooth signals surrounding the earphones, and the earphone mode is adjusted to the target mode. Using this method, the ear-covering gesture is more natural, user-friendly, and simpler to operate. It requires minimal user learning and no additional hardware is needed to recognize the user's actions, thereby controlling the earphones and improving the user experience.
[0245] The above embodiments use audio signals and Bluetooth signals as examples for illustration. It can also use photosensitive sensors, cameras, proximity sensors (infrared, ultrasonic, capacitive), etc., to detect light obstruction and determine the user's gestures; this solution does not specifically limit this.
[0246] In situations where the headphones are covered by a hat, clothing, or other coverings, or when the user is in a noisy environment, the aforementioned embodiments may incorrectly identify some noise as the user performing gestures such as covering their ears or listening. For example, if a user is wearing a hat covering their ears and is in a noisy environment, and a vehicle quickly drives past them, the change in noise energy from the vehicle might be interpreted as the user covering their ears. Therefore, referring to... Figure 8a The diagram shown is a flowchart illustrating the sixth headphone control method provided in this application. The method includes steps 801-804, as detailed below:
[0247] 801. Detect gesture information, wherein the gesture information is the audio signal collected by the headphones;
[0248] 802. Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0249] 803. If the frequency domain of the frequency domain signal includes at least a first frequency band, a second frequency band, and a third frequency band, and the last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band, and the difference between the seventh ratio of the average energy value of the first frequency band to the average energy value of the second frequency band and the eighth ratio of the average energy value of the second frequency band to the average energy value of the third frequency band is greater than a fifth preset threshold, it is determined that the user has covered their ears.
[0250] Specifically, the ratio between the energies of three consecutive frequency domain signals was used to determine whether a user covered their ears in a high-noise scenario with external obstruction.
[0251] like Figure 8b The diagram shows the energy distribution at different frequency bands when a user covers their ears. Figure 8c The diagram shows the energy distribution across different frequency bands when a user's ears are covered and the ambient noise level is high. The difference between the energy ratios of three consecutive frequency bands is used to determine if the ears are truly covered.
[0252] The above description uses one implementation as an example. It can also be based on the assumption that the user has covered their ears if the energy peak of the frequency domain signal is within a preset frequency range. This solution does not specifically limit this approach.
[0253] 804. Adjust the mode of the headphones to the target mode.
[0254] For example, when the system detects that the user is covering their ears, the headphones are triggered to enter noise cancellation mode.
[0255] When the system detects that the user is performing a listening gesture, it triggers the headphones to enter pass-through mode.
[0256] In this embodiment, for scenarios with high noise levels, the user's gesture is determined based on the audio signals surrounding the headphones, and the headphones are then adjusted to the target mode. This method further improves the detection accuracy in high-noise scenarios with external obstructions.
[0257] exist Figure 8a Based on the illustrated embodiment, referring to Figure 9 The diagram shown is a flowchart illustrating the seventh headphone control method provided in this application. The method includes steps 901-908, as detailed below:
[0258] 901. Detect gesture information, wherein the gesture information is the audio signal collected by the headphones;
[0259] This embodiment uses the ambient signal collected by the headphones as the audio signal for illustration.
[0260] 902. Determine the user's environment based on the gesture information;
[0261] 903. If the environment is a noisy environment, the user's first gesture is determined based on the gesture information. The first gesture includes at least one of covering the ears to form a cavity around the headphones and listening to the ears to form an open reflective surface around the headphones.
[0262] One method to determine whether an environment is noisy is by detecting the proportion of high-frequency energy obtained from a Fast Fourier Transform (FFT). This proportion of high-frequency energy can be understood as, for example, the ratio of energy in the 6000Hz-12000Hz frequency band to the total energy in the frequency band. If it exceeds a certain threshold, the environment is considered noisy. Alternatively, a specific decibel threshold can be used as the criterion, or the environment can be considered noisy if the total frequency energy of the collected audio signal exceeds a certain range.
[0263] Furthermore, it is possible to determine whether it is a noisy environment by adding additional sensors, etc. This step is not further limited, as long as the user's environment can be identified.
[0264] 904. If the first gesture is covering the ears, confirm whether the peak energy frequency in the frequency domain signal of the audio signal appears within the preset frequency range;
[0265] The preset frequency range can be between 900Hz and 1500Hz. If the peak value is not within this range, it is considered that no user's ear-covering action has been detected, and the detection ends.
[0266] 905. If satisfied, confirm whether the difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band in the frequency domain signal of the audio signal and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than the seventh preset threshold.
[0267] The first, second, and third frequency bands mentioned above are three consecutive frequency ranges. Specifically, the last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band.
[0268] 906. If the difference is greater than the seventh preset threshold, it is confirmed that the user has covered their ears;
[0269] If it is confirmed that the user covered their ears, proceed to step 908.
[0270] 907. If the environment is a quiet environment, the gesture performed by the user is determined based on the gesture information, and the gesture includes at least one of covering the ears to form a cavity around the headphones and listening to the ears to form an open reflective surface around the headphones;
[0271] When the environment is quiet, the gesture performed by the user can be determined directly based on the gesture information according to the method of the aforementioned embodiment, and step 908 can be executed.
[0272] 908. Adjust the mode of the headphones to the target mode.
[0273] In other words, in this embodiment of the application, if it is initially determined that the user has covered their ears in a noisy scene, further confirmation is made based on steps 904 and 905 in order to finally confirm the gesture performed by the user.
[0274] This method improves the accuracy of gesture detection in noisy scenarios by differentiating between noisy and quiet scenarios and then implementing different controls based on those scenarios.
[0275] Based on the above interactive embodiments, such as Figure 10 As shown in the figure, this application embodiment also provides a headphone control method, which includes steps 1001-1002, as follows:
[0276] 1001. Collect environmental signals around the headphones and extract features from the environmental signals;
[0277] The aforementioned environmental signals include at least one of the following: audio signals, light signals, and ultrasonic signals.
[0278] The triggering condition for step 1001 above can be that the Bluetooth signal strength of the earphone is lower than a preset threshold, triggering the gesture detection information; or, the gesture detection information can be triggered when a preset signal is received, wherein the preset signal indicates that the wearable device detects that the user has raised their hand. Other triggering conditions are also possible, and this solution does not specifically limit them.
[0279] This could also involve real-time acquisition of ambient signals around the headphones, or periodic acquisition, etc. This solution does not impose specific limitations on this.
[0280] The triggering conditions mentioned above can be found in the foregoing. Figure 2 The description of the illustrated embodiments will not be repeated here.
[0281] When the environmental signal is an audio signal, the above-mentioned acquisition of the environmental signal around the headphones and feature extraction of the environmental signal can be performed as follows:
[0282] The audio signal around the headphones is collected, and the audio signal is subjected to Fourier transform to obtain the frequency domain signal of the audio signal.
[0283] For details on the Fourier transform described above, please refer to the description of the foregoing embodiments; further details will not be repeated here.
[0284] 1002. Based on the energy intensity of the preset frequency band in the extracted environmental signal features, adjust the mode of the headphones to the target mode.
[0285] As one implementation, step 1002 may include:
[0286] Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain;
[0287] If the ratio is greater than a first preset threshold, the target mode is the first mode.
[0288] In other words, the headphone mode is adjusted based on the comparison result by comparing the energy of the frequency domain signal within the first preset frequency domain range of the aforementioned frequency domain signal with the energy of the entire frequency band.
[0289] The first mode mentioned above can be a noise reduction on mode or a noise reduction pass-through off mode, etc. Of course, other modes are also possible, and this solution does not specifically limit them.
[0290] Furthermore, if the ratio is less than a preset threshold A, the target mode is the second mode, wherein the preset threshold A is less than the first preset threshold.
[0291] The second mode can be a pass-through enabled mode. Of course, other modes are also possible, but this solution does not impose specific limitations on them.
[0292] For an introduction to this implementation method, please refer to the aforementioned document. Figure 3a Examples are not described in detail here.
[0293] Furthermore, adjustments and controls can be made based on the current headphone mode. For example, when the current headphone mode is noise cancellation pass-through off mode, the first mode is noise cancellation on mode; or, when the current headphone mode is noise cancellation on mode, the first mode is pass-through on mode; or, when the current headphone mode is pass-through on mode, the first mode is noise cancellation pass-through off mode, etc. Other control strategies are also possible, and this solution does not specifically limit them.
[0294] As another implementation, step 1002 may include:
[0295] Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution;
[0296] If the similarity value is greater than the second preset threshold, the target pattern is the first pattern.
[0297] The headphone mode is controlled by comparing the frequency energy curve of the frequency domain signal with a preset frequency energy curve.
[0298] For an introduction to this implementation method, please refer to the aforementioned document. Figure 4a The embodiments shown are not described in detail here.
[0299] As another implementation, when the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the feature extraction of the environmental signal includes:
[0300] Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal;
[0301] Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range.
[0302] Step 1002 includes:
[0303] When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the target mode is the first mode.
[0304] If, in the frequency domain signals of the first audio signal and the second audio signal, there exists at least one audio signal whose average energy value in the frequency domain range of the first audio signal is higher than the average energy value in the full frequency domain signal by a ratio greater than a sixth preset threshold, and whose average energy value in the ninth preset frequency domain range is also higher than the sixth preset threshold, then the target mode is the second mode.
[0305] For an introduction to this implementation method, please refer to the aforementioned document. Figure 5a The embodiments shown are not described in detail here.
[0306] In some cases, when the headphones are covered by a hat, clothing, or other coverings, or when the user is in a noisy environment, the aforementioned embodiments may incorrectly identify some noise as the user performing gestures such as covering their ears or listening. For example, if a user is wearing a hat covering their ears and is in a noisy environment, and a vehicle quickly passes by, the change in noise energy from the vehicle may be identified as the user covering their ears. Based on this, this solution, building upon the above embodiments, further determines that if the peak energy of the frequency domain signal is within a preset frequency range, then the target mode is the first mode.
[0307] Furthermore, based on the above embodiments, the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. If the difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than a seventh preset threshold, then the target mode is the first mode.
[0308] For a description of this embodiment, please refer to [link / reference]. Figure 8a The description of the illustrated embodiments will not be repeated here.
[0309] In this embodiment, feature extraction is performed based on environmental signals surrounding the headphones. Then, based on the energy intensity of a preset frequency band within the extracted environmental signal features, the headphone mode is adjusted to a target mode. Using this method, the ear-covering action is more natural, user-friendly, and simpler to operate. It requires minimal user learning and no additional hardware is needed to recognize user actions, thereby controlling the headphones and improving the user experience.
[0310] Reference Figure 11 The diagram shown is a schematic representation of an earphone control device provided in an embodiment of this application. Figure 11 As shown, it includes a detection module 1101, a signal processing module 1102, and a noise reduction control module 1103, wherein:
[0311] The detection module 1101 is used to detect gesture information;
[0312] The signal processing module 1102 is used to determine the gesture performed by the user based on the gesture information, wherein the gesture information is an environmental signal collected by the earphone, and the gesture includes at least one of covering the ear to form a cavity around the earphone and listening to the ear to form an open reflective surface around the earphone.
[0313] The noise reduction control module 1103 is used to adjust the mode of the headphones to the target mode according to the gesture performed by the user.
[0314] The environmental signals include at least one of the following: audio signals, Bluetooth signals, optical signals, and ultrasonic signals.
[0315] When the environmental signal is an audio signal, the signal processing module 1102 is used to:
[0316] Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0317] Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain;
[0318] If the ratio is greater than a first preset threshold, it is determined that the user has covered their ears.
[0319] The signal processing module 1102 is also used for:
[0320] If the ratio is less than a second preset threshold, it is determined that the user has performed listening, wherein the second preset threshold is less than the first preset threshold.
[0321] When the environmental signal is an audio signal, the signal processing module 1102 is used to:
[0322] Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0323] Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution;
[0324] If the similarity value is greater than the second preset threshold, it is determined that the user has covered their ears.
[0325] When the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the signal processing module 1102 is used for:
[0326] Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal;
[0327] Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range.
[0328] When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, it is determined that the user has covered their ears.
[0329] When the environmental signal is a second audio signal, the signal processing module 1102 is used to:
[0330] The second audio signal is sent to the computing unit, so that the computing unit performs Fourier transform on the second audio signal and the first audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal, wherein the first audio signal and the second audio signal are collected by different headphones;
[0331] The system receives information sent by the computing unit and determines, based on the information, that the user has covered their ears. The information indicates that a first ratio is greater than a second ratio, and a third ratio is less than a fourth ratio, and / or, the ratio between the third and fourth ratios is greater than the ratio between the fifth and sixth ratios. The first ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The second ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The third ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range. The ratio of the average value to the average energy value of the frequency domain signal in the full frequency domain, the fourth ratio being the ratio of the average energy value of the frequency domain signal of the second audio signal in the fourth preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, the fifth ratio being the ratio of the average energy value of the frequency domain signal of the first audio signal in the third preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, and the sixth ratio being the ratio of the average energy value of the frequency domain signal of the first audio signal in the fourth preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, wherein the frequency point of the second preset frequency domain range is lower than the frequency point of the fourth preset frequency domain range, and the frequency point of the third preset frequency domain range is higher than the frequency point of the fourth preset frequency domain range.
[0332] The detection module 1101 is further used for:
[0333] If the ratio of the average energy of the frequency domain signal of the first audio signal and the average energy of the frequency domain signal of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is not higher than the fifth preset threshold, or if the ratio of the average energy of the frequency domain signal of the first audio signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the ratio of the average energy of the frequency domain signal of the seventh preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, the Bluetooth signal of the earphone is obtained, and the time intensity distribution of the Bluetooth signal is obtained.
[0334] The signal processing module 1102 is also used for:
[0335] If the strength of the Bluetooth signal in the first time period is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than the ninth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0336] The signal processing module 1102 is also used for:
[0337] If, in the frequency domain signals of the first audio signal and the second audio signal, there exists at least one audio signal whose average energy value in the eighth preset frequency domain range is higher than the average energy value of the frequency domain signal across the entire frequency domain than a sixth preset threshold, and whose average energy value in the ninth preset frequency domain range is also higher than the sixth preset threshold, then it is determined that the user has performed listening.
[0338] The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
[0339] The frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. The difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than a seventh preset threshold.
[0340] When the environmental signal is a Bluetooth signal, the signal processing module 1102 is further configured to:
[0341] The temporal intensity distribution of the Bluetooth signal is obtained based on the Bluetooth signal.
[0342] If the strength of the Bluetooth signal in the first time period of the time intensity distribution is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than a sixth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
[0343] When the user performs a gesture of covering their ears, the target mode is either noise reduction on mode or noise reduction pass-through off mode.
[0344] Alternatively, when the current mode of the headphones is noise cancellation pass-through off mode, the target mode is noise cancellation on mode;
[0345] Alternatively, when the current mode of the headphones is noise cancellation enabled, the target mode is pass-through enabled.
[0346] Alternatively, when the current mode of the headphones is pass-through enabled, the target mode is noise cancellation pass-through disabled.
[0347] Alternatively, when the user performs a gesture of covering their ears, the target mode is noise reduction enabled.
[0348] When the user performs a listening gesture, the target mode is the pass-through enabled mode.
[0349] The device further includes a trigger module, used to trigger the detection gesture information when the strength value of the Bluetooth signal of the earphone is lower than a tenth preset threshold.
[0350] The device also includes a trigger module, used to: trigger the detected gesture information when a preset signal is received, wherein the preset signal indicates that the wearable device has detected that the user has raised their hand.
[0351] For a description of each of the above modules, please refer to the foregoing embodiments; they will not be repeated here.
[0352] Reference Figure 12 The diagram shown is a schematic of another headphone control device provided in an embodiment of this application. Figure 12 As shown, it includes a signal acquisition module 1201, a signal processing module 1202, and a noise reduction control module 1203, wherein:
[0353] Signal acquisition module 1201 is used to acquire environmental signals around the headphones;
[0354] Signal processing module 1202 is used to extract features from the environmental signals;
[0355] The noise reduction control module 1203 is used to adjust the mode of the headphones to the target mode based on the energy intensity of the preset frequency band in the extracted environmental signal features.
[0356] The environmental signals include at least one of the following: audio signals, light signals, and ultrasonic signals.
[0357] When the environmental signal is an audio signal, the signal processing module 1202 is used to:
[0358] Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0359] Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain;
[0360] The noise reduction control module 1203 is used for:
[0361] If the ratio is greater than a first preset threshold, the target mode is the first mode.
[0362] When the environmental signal is an audio signal, the signal processing module 1202 is used to:
[0363] Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal;
[0364] Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution;
[0365] The noise reduction control module 1203 is used for:
[0366] If the similarity value is greater than the second preset threshold, the target pattern is the first pattern.
[0367] When the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the signal processing module 1202 is used to:
[0368] Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal;
[0369] Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range.
[0370] The noise reduction control module 1203 is used for:
[0371] When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain, and the target mode is the first mode.
[0372] If the ratio of the average energy of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the third preset threshold, and the ratio of the average energy of the second audio signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the third preset threshold, then the target mode is the second mode.
[0373] The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
[0374] Furthermore, the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. The difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than a seventh preset threshold.
[0375] When the target mode is the first mode, the first mode is either the noise reduction on mode or the noise reduction pass-through off mode.
[0376] Alternatively, when the current mode of the headphones is noise cancellation pass-through off mode, the first mode is noise cancellation on mode;
[0377] Alternatively, when the current mode of the headphones is noise cancellation enabled, the first mode is pass-through enabled;
[0378] Alternatively, when the current mode of the headphones is pass-through enabled, the first mode is noise cancellation pass-through disabled.
[0379] Alternatively, when the target mode is the first mode, the first target mode is the noise reduction enabled mode;
[0380] When the target mode is the second mode, the second mode is the pass-through enabled mode.
[0381] The device further includes a trigger module for:
[0382] When the Bluetooth signal strength of the headset is lower than the tenth preset threshold, the detection gesture information is triggered.
[0383] The device further includes a trigger module for:
[0384] When a preset signal is received, the detection gesture information is triggered, and the preset signal indicates that the wearable device has detected that the user has raised their hand.
[0385] For a description of each of the above modules, please refer to the foregoing embodiments; they will not be repeated here.
[0386] In this embodiment, the headphone control device is presented in the form of a module. Here, "module" can refer to an application-specific integrated circuit (ASIC), a processor and memory that executes one or more software or firmware programs, integrated logic circuits, and / or other devices that can provide the above-mentioned functions.
[0387] Furthermore, the above-mentioned detection module 1101, signal processing module 1102, and noise reduction control module 1103, as well as the signal acquisition module 1201, signal processing module 1202, and noise reduction control module 1203, can be... Figure 13 The processor 1302 of the headphone control device shown is used for implementation.
[0388] Figure 13 This is a schematic diagram of the hardware structure of the headphone control device provided in the embodiments of this application. Figure 13 The headphone control device 1300 shown (which may specifically be a computer device) includes a memory 1301, a processor 1302, a communication interface 1303, and a bus 1304. The memory 1301, processor 1302, and communication interface 1303 are interconnected via the bus 1304.
[0389] The memory 1301 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM).
[0390] The memory 1301 can store programs. When the program stored in the memory 1301 is executed by the processor 1302, the processor 1302 and the communication interface 1303 are used to execute the various steps of the headphone control method of the present application embodiment.
[0391] The processor 1302 may be a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute relevant programs to implement the functions required by the units in the headphone control device of this application embodiment, or to execute the headphone control method of this application method embodiment.
[0392] The processor 1302 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the headphone control method of this application can be completed by the integrated logic circuitry in the hardware of the processor 1302 or by instructions in software form. The processor 1302 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory 1301. The processor 1302 reads the information in the memory 1301 and, in conjunction with its hardware, performs the functions required by the units included in the headphone control device of this application embodiment, or executes the headphone control method of this application method embodiment.
[0393] The communication interface 1303 uses a transceiver device, such as, but not limited to, a transceiver, to enable communication between the device 1300 and other devices or communication networks. For example, data can be acquired through the communication interface 1303.
[0394] Bus 1304 may include a pathway for transmitting information between various components of device 1300 (e.g., memory 1301, processor 1302, communication interface 1303).
[0395] It should be noted that, although Figure 13 The illustrated device 1300 only shows the memory, processor, and communication interface. However, those skilled in the art should understand that in specific implementations, device 1300 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that device 1300 may also include hardware devices for implementing other additional functions. Moreover, those skilled in the art should understand that device 1300 may only include the devices necessary for implementing the embodiments of this application, and may not necessarily include... Figure 13 All the devices shown.
[0396] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods.
[0397] This application also provides a computer program product containing instructions. When the computer program product is run on a computer or processor, it causes the computer or processor to perform one or more steps of any of the methods described above.
[0398] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the specific descriptions of the corresponding steps in the foregoing method embodiments, and will not be repeated here.
[0399] It should be understood that in the description of this application, unless otherwise stated, " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B can represent A or B; where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Additionally, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" do not necessarily imply difference. In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0400] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling, direct coupling, or communication connection shown or discussed between each other may be indirect coupling or communication connection through some interfaces, apparatuses, or units, and may be electrical, mechanical, or other forms.
[0401] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0402] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be read-only memory (ROM), random access memory (RAM), or magnetic media, such as floppy disks, hard disks, magnetic tapes, magnetic disks, or optical media, such as digital versatile discs (DVDs), or semiconductor media, such as solid-state disks (SSDs).
[0403] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A headphone control method, characterized in that, include: Detect gesture information; The gesture performed by the user is determined based on the gesture information, which is an environmental signal collected by the headphones. The gesture includes at least one of covering the ears to form a cavity around the headphones and listening to the ears to form an open reflective surface around the headphones. The environmental signal includes an audio signal, and the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. If the difference between the seventh ratio of the average energy value of the first frequency band to the average energy value of the second frequency band and the eighth ratio of the average energy value of the second frequency band to the average energy value of the third frequency band is greater than a fifth preset threshold, it is determined that the user has performed an ear-covering operation. Based on the gesture performed by the user, the mode of the headphones is adjusted to the target mode.
2. The method according to claim 1, characterized in that, The environmental signals include at least one of the following: audio signals, Bluetooth signals, optical signals, and ultrasonic signals.
3. The method according to claim 2, characterized in that, When the environmental signal is an audio signal, determining the gesture performed by the user based on the gesture information includes: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; If the ratio is greater than a first preset threshold, it is determined that the user has covered their ears.
4. The method according to claim 2, characterized in that, When the environmental signal is an audio signal, determining the gesture performed by the user based on the gesture information includes: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution; If the similarity value is greater than the second preset threshold, it is determined that the user has covered their ears.
5. The method according to claim 2, characterized in that, When the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, determining the gesture performed by the user based on the gesture information includes: Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range. When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, it is determined that the user has covered their ears.
6. The method according to claim 2, characterized in that, When the environmental signal is a second audio signal, determining the gesture performed by the user based on the gesture information includes: The second audio signal is sent to the computing unit, so that the computing unit performs Fourier transform on the second audio signal and the first audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal, wherein the first audio signal and the second audio signal are collected by different headphones; The system receives information sent by the computing unit and determines, based on the information, that the user has covered their ears. The information indicates that a first ratio is greater than a second ratio, and a third ratio is less than a fourth ratio, and / or, the ratio between the third and fourth ratios is greater than the ratio between the fifth and sixth ratios. The first ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The second ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The third ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range. The ratio of the average value to the average energy value of the frequency domain signal in the full frequency domain, the fourth ratio being the ratio of the average energy value of the frequency domain signal of the second audio signal in the fourth preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, the fifth ratio being the ratio of the average energy value of the frequency domain signal of the first audio signal in the third preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, and the sixth ratio being the ratio of the average energy value of the frequency domain signal of the first audio signal in the fourth preset frequency domain range to the average energy value of the frequency domain signal in the full frequency domain, wherein the frequency point of the second preset frequency domain range is lower than the frequency point of the fourth preset frequency domain range, and the frequency point of the third preset frequency domain range is higher than the frequency point of the fourth preset frequency domain range.
7. The method according to claim 6, characterized in that, If the ratio of the average energy of the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is not higher than the third preset threshold, or the ratio of the average energy of the frequency domain signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the ratio of the average energy of the frequency domain signal in the seventh preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, the Bluetooth signal of the earphone is obtained, and the time intensity distribution of the Bluetooth signal is obtained. If the strength of the Bluetooth signal in the first time period is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than the fourth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
8. The method according to any one of claims 5 to 7, characterized in that, If the ratio of the average energy of the second audio signal in the eighth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the fourth preset threshold, and the ratio of the average energy of the second audio signal in the ninth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the fourth preset threshold, it is determined that the user has performed listening.
9. The method according to any one of claims 3 to 7, characterized in that, The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
10. The method according to claim 2, characterized in that, When the environmental signal is a Bluetooth signal, determining the gesture performed by the user based on the gesture information includes: The temporal intensity distribution of the Bluetooth signal is obtained based on the Bluetooth signal. If the strength of the Bluetooth signal in the first time period of the time intensity distribution is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than a sixth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
11. The method according to claim 1, characterized in that, When the user performs a gesture of covering their ears, the target mode is either noise reduction on mode or noise reduction pass-through off mode. Alternatively, when the current mode of the headphones is noise cancellation pass-through off mode, the target mode is noise cancellation on mode; Alternatively, when the current mode of the headphones is noise cancellation enabled, the target mode is pass-through enabled. Alternatively, when the current mode of the headphones is pass-through enabled, the target mode is noise cancellation pass-through disabled.
12. The method according to claim 1, characterized in that, When the user performs a gesture of covering their ears, the target mode is noise reduction enabled. When the user performs a listening gesture, the target mode is the pass-through enabled mode.
13. The method according to claim 1, characterized in that, The method further includes: When the Bluetooth signal strength of the headset is lower than the seventh preset threshold, gesture information detection begins.
14. The method according to claim 1, characterized in that, The method further includes: When a preset signal is received, the device begins to detect gesture information. The preset signal indicates that the wearable device has detected that the user has raised their hand.
15. A headphone control method, characterized in that, include: Collect environmental signals around the headphones and extract features from the environmental signals; Based on the energy intensity of a preset frequency band in the extracted environmental signal features, the mode of the headphones is adjusted to the target mode. The environmental signal includes an audio signal, and the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. If the difference between the seventh ratio of the average energy value of the first frequency band to the average energy value of the second frequency band and the eighth ratio of the average energy value of the second frequency band to the average energy value of the third frequency band is greater than a fifth preset threshold, the mode of the headphones is adjusted to the target mode.
16. The method according to claim 15, characterized in that, The environmental signal includes at least one of the following: audio signal, light signal, and ultrasonic signal.
17. The method according to claim 16, characterized in that, When the environmental signal is an audio signal, the feature extraction of the environmental signal includes: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; The step of adjusting the headphone mode to the target mode based on the energy intensity of the preset frequency band in the extracted environmental signal features includes: If the ratio is greater than a first preset threshold, the target mode is the first mode.
18. The method according to claim 16, characterized in that, When the environmental signal is an audio signal, the feature extraction of the environmental signal includes: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution; The step of adjusting the headphone mode to the target mode based on the energy intensity of the preset frequency band in the extracted environmental signal features includes: If the similarity value is greater than the second preset threshold, the target pattern is the first pattern.
19. The method according to claim 16, characterized in that, When the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the feature extraction of the environmental signal includes: Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range. The step of adjusting the headphone mode to the target mode based on the energy intensity of the preset frequency band in the extracted environmental signal features includes: When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the target mode is the first mode.
20. The method according to claim 19, characterized in that, If the ratio of the average energy of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the third preset threshold, and the ratio of the average energy of the second audio signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the third preset threshold, then the target mode is the second mode.
21. The method according to any one of claims 17 to 20, characterized in that, The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
22. The method according to any one of claims 15 to 20, characterized in that, When the target mode is the first mode, the first mode is either the noise reduction on mode or the noise reduction pass-through off mode. Alternatively, when the current mode of the headphones is noise cancellation pass-through off mode, the first mode is noise cancellation on mode; Alternatively, when the current mode of the headphones is noise cancellation enabled, the first mode is pass-through enabled; Alternatively, when the current mode of the headphones is pass-through enabled, the first mode is noise cancellation pass-through disabled.
23. The method according to any one of claims 15 to 20, characterized in that, When the target mode is the first mode, the first mode is the noise reduction enabled mode; When the target mode is the second mode, the second mode is the pass-through enabled mode.
24. The method according to any one of claims 15 to 20, characterized in that, The method further includes: When the Bluetooth signal strength of the headset is lower than the fifth preset threshold, gesture information detection begins.
25. The method according to any one of claims 15 to 20, characterized in that, The method further includes: When a preset signal is received, the device begins to detect gesture information. The preset signal indicates that the wearable device has detected that the user has raised their hand.
26. A headphone control device, characterized in that, include: The detection module is used to detect gesture information; A signal processing module is used to determine the gesture performed by the user based on the gesture information, wherein the gesture information is an environmental signal collected by the earphone, and the gesture includes at least one of covering the ears to form a cavity around the earphone and listening to the ears to form an open reflective surface around the earphone; the environmental signal includes an audio signal, and the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band, wherein the last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band; if the difference between the seventh ratio of the average energy of the first frequency band to the average energy of the second frequency band and the eighth ratio of the average energy of the second frequency band to the average energy of the third frequency band is greater than a fifth preset threshold, it is determined that the user has performed an ear-covering operation; The noise reduction control module is used to adjust the mode of the headphones to the target mode based on the gestures performed by the user.
27. The apparatus according to claim 26, characterized in that, The environmental signals include at least one of the following: audio signals, Bluetooth signals, optical signals, and ultrasonic signals.
28. The apparatus according to claim 27, characterized in that, When the environmental signal is an audio signal, the signal processing module is used to: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; If the ratio is greater than a first preset threshold, it is determined that the user has covered their ears.
29. The apparatus according to claim 27, characterized in that, When the environmental signal is an audio signal, the signal processing module is used to: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution; If the similarity value is greater than the second preset threshold, it is determined that the user has covered their ears.
30. The apparatus according to claim 27, characterized in that, When the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the signal processing module is used to: Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range. When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, it is determined that the user has covered their ears.
31. The apparatus according to claim 27, characterized in that, The device also includes a communication module, which is used when the environmental signal is a second audio signal. The communication module is used to send the second audio signal to the computing unit, so that the computing unit performs Fourier transform on the second audio signal and the first audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal, wherein the first audio signal and the second audio signal are collected by different headphones; The signal processing module receives information sent by the computing unit and determines, based on the information, that the user has covered their ears. The information indicates that a first ratio is greater than a second ratio, and a third ratio is less than a fourth ratio, and / or, the ratio between the third and fourth ratios is greater than the ratio between the fifth and sixth ratios. The first ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The second ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The third ratio is the ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range. The fourth ratio is the ratio of the average energy value of the second audio signal within a fourth preset frequency range to the average energy value of the frequency domain signal across the entire frequency domain. The fifth ratio is the ratio of the average energy value of the first audio signal within a third preset frequency range to the average energy value of the frequency domain signal across the entire frequency domain. The sixth ratio is the ratio of the average energy value of the first audio signal within a fourth preset frequency range to the average energy value of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency range are lower than those within the fourth preset frequency range, and the frequency points within the third preset frequency range are higher than those within the fourth preset frequency range.
32. The apparatus according to claim 31, characterized in that, If the ratio of the average energy of the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is not higher than the third preset threshold, or the ratio of the average energy of the frequency domain signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the ratio of the average energy of the frequency domain signal in the seventh preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, the Bluetooth signal of the earphone is obtained, and the time intensity distribution of the Bluetooth signal is obtained. If the strength of the Bluetooth signal in the first time period is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than the fourth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
33. The apparatus according to any one of claims 30 to 32, characterized in that, If the ratio of the average energy of the second audio signal in the eighth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the fourth preset threshold, and the ratio of the average energy of the second audio signal in the ninth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the fourth preset threshold, it is determined that the user has performed listening.
34. The apparatus according to any one of claims 28 to 32, characterized in that, The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
35. The apparatus according to claim 27, characterized in that, When the environmental signal is a Bluetooth signal, the signal processing module is used to: The temporal intensity distribution of the Bluetooth signal is obtained based on the Bluetooth signal. If the strength of the Bluetooth signal in the first time period of the time intensity distribution is lower than the strength of the Bluetooth signal in the second time period, and the strength of the Bluetooth signal in the first time period is lower than a sixth preset threshold within a preset duration, it is determined that the user has covered their ears, wherein the second time period is earlier than the first time period.
36. The apparatus according to claim 26, characterized in that, When the user performs a gesture of covering their ears, the target mode is either noise reduction on mode or noise reduction pass-through off mode. Alternatively, when the current mode of the headphones is noise cancellation pass-through off mode, the target mode is noise cancellation on mode; Alternatively, when the current mode of the headphones is noise cancellation enabled, the target mode is pass-through enabled. Alternatively, when the current mode of the headphones is pass-through enabled, the target mode is noise cancellation pass-through disabled.
37. The apparatus according to claim 26, characterized in that, When the user performs a gesture of covering their ears, the target mode is noise reduction enabled. When the user performs a listening gesture, the target mode is the pass-through enabled mode.
38. The apparatus according to claim 26, characterized in that, The device further includes a trigger module for: When the Bluetooth signal strength of the headset is lower than the seventh preset threshold, gesture information detection begins.
39. The apparatus according to claim 26, characterized in that, The device further includes a trigger module for: When a preset signal is received, the device begins to detect gesture information. The preset signal indicates that the wearable device has detected that the user has raised their hand.
40. A headphone control device, characterized in that, include: The signal acquisition module is used to collect environmental signals around the headphones; The signal processing module is used to extract features from the environmental signals; A noise reduction control module is used to adjust the mode of the headphones to a target mode based on the energy intensity of a preset frequency band in the extracted environmental signal features. The environmental signal includes an audio signal, and the frequency domain of the audio signal includes at least a first frequency band, a second frequency band, and a third frequency band. The last frequency point of the first frequency band is the first frequency point of the second frequency band, and the last frequency point of the second frequency band is the first frequency point of the third frequency band. When the difference between the seventh ratio of the average energy value of the first frequency band to the average energy value of the second frequency band and the eighth ratio of the average energy value of the second frequency band to the average energy value of the third frequency band is greater than a fifth preset threshold, the mode of the headphones is adjusted to the target mode.
41. The apparatus according to claim 40, characterized in that, The environmental signal includes at least one of the following: audio signal, light signal, and ultrasonic signal.
42. The apparatus according to claim 41, characterized in that, When the environmental signal is an audio signal, the signal processing module is used to: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the ratio of the average energy of the frequency domain signal within the first preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; The noise reduction control module is used for: If the ratio is greater than a first preset threshold, the target mode is the first mode.
43. The apparatus according to claim 41, characterized in that, When the environmental signal is an audio signal, the signal processing module is used to: Perform a Fourier transform on the audio signal to obtain the frequency domain signal of the audio signal; Calculate the similarity value between the frequency energy distribution of the frequency domain signal and the preset frequency energy distribution; The noise reduction control module is used for: If the similarity value is greater than the second preset threshold, the target pattern is the first pattern.
44. The apparatus according to claim 41, characterized in that, When the environmental signal includes a first audio signal and a second audio signal, and the acquisition time of the first audio signal is earlier than the acquisition time of the second audio signal, the signal processing module is used to: Fourier transforms are performed on the first audio signal and the second audio signal respectively to obtain the frequency domain signal of the first audio signal and the frequency domain signal of the second audio signal; Calculate the following ratios: a first ratio of the average energy of the frequency domain signal of the second audio signal within a second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a second ratio of the average energy of the frequency domain signal of the first audio signal within the second preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; a third ratio of the average energy of the frequency domain signal of the second audio signal within a third preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain; and a fourth ratio of the average energy of the frequency domain signal of the second audio signal within a fourth preset frequency domain range to the average energy of the frequency domain signal across the entire frequency domain. The frequency points within the second preset frequency domain range are lower than the frequency points within the fourth preset frequency domain range, and the frequency points within the third preset frequency domain range are higher than the frequency points within the fourth preset frequency domain range. The noise reduction control module is used for: When the first ratio is greater than the second ratio, if the third ratio is less than the fourth ratio, and / or the ratio between the third ratio and the fourth ratio is greater than the ratio between the fifth ratio and the sixth ratio, wherein the fifth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the third preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the sixth ratio is the ratio of the average energy of the frequency domain signal of the first audio signal within the fourth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain, and the target mode is the first mode.
45. The apparatus according to claim 44, characterized in that, If the ratio of the average energy of the second audio signal in the fifth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is higher than the third preset threshold, and the ratio of the average energy of the second audio signal in the sixth preset frequency domain range to the average energy of the frequency domain signal in the full frequency domain is also higher than the third preset threshold, then the target mode is the second mode.
46. The apparatus according to any one of claims 42 to 45, characterized in that, The energy peak value in the frequency energy distribution of the frequency domain signal is located within a preset frequency range.
47. The apparatus according to any one of claims 40 to 45, characterized in that, When the target mode is the first mode, the first mode is either the noise reduction on mode or the noise reduction pass-through off mode. Alternatively, when the current mode of the headphones is noise cancellation pass-through off mode, the first mode is noise cancellation on mode; Alternatively, when the current mode of the headphones is noise cancellation enabled, the first mode is pass-through enabled; Alternatively, when the current mode of the headphones is pass-through enabled, the first mode is noise cancellation pass-through disabled.
48. The apparatus according to any one of claims 40 to 45, characterized in that, When the target mode is the first mode, the first mode is the noise reduction enabled mode; When the target mode is the second mode, the second mode is the pass-through enabled mode.
49. The apparatus according to any one of claims 40 to 45, characterized in that, The device further includes a trigger module for: When the Bluetooth signal strength of the headset is lower than the fifth preset threshold, gesture information detection begins.
50. The apparatus according to any one of claims 40 to 45, characterized in that, The device further includes a trigger module for: When a preset signal is received, the device begins to detect gesture information. The preset signal indicates that the wearable device has detected that the user has raised their hand.
51. A headphone control device, characterized in that, It includes a processor and a memory; wherein the memory is used to store program code, and the processor is used to call the program code to perform the method as claimed in any one of claims 1 to 14, and / or the method as claimed in any one of claims 15 to 25.
52. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method of any one of claims 1 to 14, and / or the method of any one of claims 15 to 25.
53. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 14, and / or the method as described in any one of claims 15 to 25.
Citation Information
Patent Citations
Acoustic gesture detection for control of a hearable device
CN113196797A