Voice detection method, device, earphone and storage medium
By obtaining the voice signals of the wearer's ear canal and environment in the headset, and calculating energy parameters using the occlusion effect frequency interval, the accuracy of the headset detects the wearer's speech is solved, and efficient voice detection without additional hardware is achieved.
Patent Information
- Application Number
- CN202211042440.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-08-29
AI Technical Summary
In the prior art, the method of detecting whether the wearer is speaking requires adding sensor hardware or susceptible to ambient sound interference, resulting in high costs or inaccurate detection.
By obtaining voice signals from the wearer's ear canal and environment, and using the frequency interval (200Hz to 500Hz) generated by the occlusion effect to calculate the energy parameters of the voice signals to determine whether the wearer has issued a sound signal.
Without adding sensor hardware and microphone array, accurately determine whether the headphone wearer emits a sound signal to avoid interference from external environment.
Smart Images

Figure CN115278441B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of voice detection technology, and in particular to a voice detection method, device, earphone, and storage medium. Background Art
[0002] With the popularity of TWS (True Wireless Stereo) headphones, when users wear headphones, the headphones need to enable the sound transmission function when the wearer speaks, so that the user can transmit clear voice information when communicating with others through the headphones; or turn on the voice recognition function of the headphones when the wearer speaks to recognize the wearer's voice information and perform related operations. Therefore, the headphones need to accurately detect whether the wearer is speaking, and then determine whether to maintain the current state or enable other functions based on the wearer's voice signal based on the detection result.
[0003] Currently, in order to accurately detect whether the wearer of headphones is speaking, one solution in the related art is to set auxiliary sensors on the headphones, such as vibration sensors and acceleration sensors, to determine whether the user is speaking. This detection method increases the cost of the headphones because it requires the addition of auxiliary sensors. Another solution in the related art is to set a microphone array on the headphones to estimate the source direction of the collected sound to determine whether the wearer is speaking. With this solution, the estimation of the incoming wave direction is easily interfered by the ambient sound. When the external sound is louder than the wearer's voice, the direction of the sound source cannot be accurately estimated, resulting in inaccurate detection results. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a voice detection method, device, earphone and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a voice detection method is provided, which is applied to an in-ear headset. The voice detection method includes:
[0006] Acquire a first voice signal and a second voice signal, wherein the first voice signal is a sound signal acquired from the wearer's ear canal, and the second voice signal is a sound signal acquired from the wearer's environment;
[0007] Obtaining a first energy parameter and a second energy parameter, wherein the first energy parameter represents an energy value of the first speech signal in a preset frequency band, and the second energy parameter represents an energy value of the second speech signal in the preset frequency band, wherein the preset frequency band is a frequency interval where an occlusion effect occurs;
[0008] A numerical relationship between the first energy parameter and the second energy parameter is obtained, and if the numerical relationship indicates that the first energy parameter is greater than the second energy parameter, it is determined that the wearer has emitted a sound signal.
[0009] In an exemplary embodiment, the preset frequency band includes a starting frequency point and an ending frequency point;
[0010] The obtaining of the first energy parameter includes: obtaining a first starting energy value, a first ending energy value, and a first average energy value of the first speech signal, wherein the first starting energy value represents the energy value of the first speech signal at the starting frequency point, the first ending energy value represents the energy value of the first speech signal at the ending frequency point, and the first average energy value represents the average energy value of the first speech signal in the preset frequency band;
[0011] The obtaining of the second energy parameter includes: obtaining a second starting energy value, a second ending energy value and a second average energy value of the second speech signal, wherein the second starting energy value represents the energy value of the second speech signal at the starting frequency point, the second ending energy value represents the ending energy value of the second speech signal at the ending frequency point, and the second average energy value represents the average energy value of the second speech signal in the preset frequency band.
[0012] In an exemplary embodiment, obtaining the first average energy value includes:
[0013] In the preset frequency band, a first reference frequency point is set at intervals of a first preset frequency difference, and the preset frequency band includes a plurality of the first reference frequency points;
[0014] Obtaining a first reference energy value for each first reference frequency point in the first speech signal respectively;
[0015] summing each of the first reference energy values and taking an average value as the first average energy value;
[0016] The obtaining of the second average energy value includes:
[0017] In the preset frequency band, a second reference frequency point is set at intervals of a second preset frequency difference, and the preset frequency band includes a plurality of second reference frequency points;
[0018] Obtaining a second reference energy value for each second reference frequency point in the second speech signal respectively;
[0019] The second reference energy values are summed and the average value is taken as the second average energy value.
[0020] In an exemplary embodiment, if the numerical relationship indicates that the first energy parameter is greater than the second energy parameter, determining that the wearer has emitted a sound signal includes:
[0021] Obtaining a first energy difference value, a second energy difference value, and an average energy difference value, wherein the first energy difference value represents an energy difference between the first starting energy value and the second starting energy value, the second energy difference value represents an energy difference between the first ending energy value and the second ending energy value, and the average energy difference value represents an energy difference between the first average energy value and the second average energy value;
[0022] According to the first energy difference value, the second energy difference value and the average energy difference value, it is determined that the wearer has emitted a sound signal.
[0023] In an exemplary embodiment, determining that the wearer has emitted a sound signal based on the first energy difference value, the second energy difference value, and the average energy difference value includes:
[0024] If the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value, it is determined that the wearer has emitted a sound signal.
[0025] In an exemplary embodiment, the voice detection method further includes:
[0026] If the first energy difference is less than or equal to the second energy difference, and / or the average energy difference is less than or equal to the preset reference value, it is determined that the wearer does not emit a sound signal.
[0027] In an exemplary embodiment, the preset frequency range is 200 Hz to 500 Hz.
[0028] According to a second aspect of an embodiment of the present disclosure, a voice detection device is provided, which is applied to an in-ear headset. The voice detection device includes:
[0029] an acquisition module configured to acquire a first voice signal and a second voice signal, wherein the first voice signal is a sound signal acquired from the wearer's ear canal, and the second voice signal is a sound signal acquired from the wearer's environment;
[0030] a calculation module configured to obtain a first energy parameter and a second energy parameter, wherein the first energy parameter represents an energy value of the first speech signal in a preset frequency band, and the second energy parameter represents an energy value of the second speech signal in the preset frequency band, wherein the preset frequency band is a frequency interval where an occlusion effect occurs;
[0031] The determination module is configured to obtain a numerical relationship between the first energy parameter and the second energy parameter, and determine that the wearer has emitted a sound signal if the numerical relationship indicates that the first energy parameter is greater than the second energy parameter.
[0032] In an exemplary embodiment, the preset frequency band includes a starting frequency point and an ending frequency point;
[0033] The computing module is further configured to:
[0034] Obtaining a first starting energy value, a first ending energy value, and a first average energy value of the first speech signal, wherein the first starting energy value represents the energy value of the first speech signal at the starting frequency point, the first ending energy value represents the energy value of the first speech signal at the ending frequency point, and the first average energy value represents the average energy value of the first speech signal in the preset frequency band;
[0035] The computing module is further configured to:
[0036] Obtain a second starting energy value, a second ending energy value and a second average energy value of the second speech signal, wherein the second starting energy value represents the energy value of the second speech signal at the starting frequency point, the second ending energy value represents the energy value of the second speech signal at the ending frequency point, and the second average energy value represents the average energy value of the second speech signal in the preset frequency band.
[0037] In an exemplary embodiment, the calculation module is further configured to:
[0038] In the preset frequency band, a first reference frequency point is set at intervals of a first preset frequency difference, and the preset frequency band includes a plurality of the first reference frequency points;
[0039] Obtaining a first reference energy value for each first reference frequency point in the first speech signal respectively;
[0040] summing each of the first reference energy values and taking an average value as the first average energy value;
[0041] The computing module is further configured to:
[0042] In the preset frequency band, a second reference frequency point is set at intervals of a second preset frequency difference, and the preset frequency band includes a plurality of second reference frequency points;
[0043] Obtaining a second reference energy value for each second reference frequency point in the second speech signal respectively;
[0044] The second reference energy values are summed and the average value is taken as the second average energy value.
[0045] In an exemplary embodiment, the determining module is further configured to:
[0046] Obtaining a first energy difference value, a second energy difference value, and an average energy difference value, wherein the first energy difference value represents an energy difference between the first starting energy value and the second starting energy value, the second energy difference value represents an energy difference between the first ending energy value and the second ending energy value, and the average energy difference value represents an energy difference between the first average energy value and the second average energy value;
[0047] According to the first energy difference value, the second energy difference value and the average energy difference value, it is determined that the wearer has emitted a sound signal.
[0048] In an exemplary embodiment, the determining module is further configured to:
[0049] If the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value, it is determined that the wearer has emitted a sound signal.
[0050] In an exemplary embodiment, the determining module is further configured to:
[0051] If the first energy difference is less than or equal to the second energy difference, and / or the average energy difference is less than or equal to the preset reference value, it is determined that the wearer does not emit a sound signal.
[0052] In an exemplary embodiment, the preset frequency range is 200 Hz to 500 Hz.
[0053] According to a third aspect of an embodiment of the present disclosure, there is provided an earphone, comprising:
[0054] processor;
[0055] a memory for storing processor-executable instructions;
[0056] The processor is configured to execute the speech detection method as described in the first aspect of the embodiment of the present disclosure.
[0057] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided. When the instructions in the storage medium are executed by the processor of the headset, the headset can perform the voice detection method as described in the first aspect of the embodiment of the present disclosure.
[0058] The above method disclosed in the present invention has the following beneficial effects: without the need for additional sensor hardware and a microphone array composed of headphones, it is possible to accurately determine whether the headphone wearer has emitted a sound signal.
[0059] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0061] Figure 1 is a flow chart of a voice detection method according to an exemplary embodiment;
[0062] Figure 2 1 is a schematic structural diagram of an in-ear TWS headset according to an exemplary embodiment;
[0063] Figure 3 is a block diagram of a speech detection device according to an exemplary embodiment;
[0064] Figure 4 is a block diagram of a speech detection apparatus according to an exemplary embodiment. DETAILED DESCRIPTION
[0065] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0066] In the related art, when a user wears headphones, there are two methods for the headphones to detect whether the wearer is talking: first, an auxiliary sensor is built into the headphones, such as a vibration sensor or an acceleration sensor, and the sensor signal is used to determine whether the wearer is talking. However, this method requires the user to wear the headphones in a preset posture, which affects the user experience, and the cost of the sensor is relatively high; second, a microphone array is formed by multiple headphones on the headphones to collect sound signals, and an algorithm related to sound direction finding is used to estimate the incoming wave direction of the collected sound signal to determine whether the direction of the sound signal is the direction of the wearer's mouth, thereby determining whether the wearer is talking. However, this method is easily interfered by ambient sound. When the ambient sound is louder than the wearer's voice, detection errors may occur.
[0067] In an exemplary embodiment of the present disclosure, in order to overcome the problems in the related art, a voice detection method is provided, which is applied to in-ear headphones. A first voice signal is obtained from the wearer's ear canal, and a second voice signal is obtained from the wearer's environment. The frequency interval generated by the occlusion effect is selected as a preset frequency band, and a first energy parameter of the first voice signal in the preset frequency band and a second energy parameter of the second voice signal in the preset frequency band are obtained. The numerical relationship between the first energy parameter and the second energy parameter is obtained. If the numerical relationship indicates that the first energy parameter is greater than the second energy parameter, it is determined that the wearer has emitted a sound signal. Using the voice detection method of the present disclosure, it is possible to accurately determine whether the wearer of the headphones has emitted a sound signal without the need for additional sensor hardware and a microphone array composed of headphones.
[0068] In an exemplary embodiment of the present disclosure, a voice detection method is provided, which is applied to an in-ear headset. Figure 1 FIG. 1 is a flow chart showing a method for voice detection according to an exemplary embodiment. Figure 1 As shown, the voice detection method includes the following steps:
[0069] Step S101, obtaining a first voice signal and a second voice signal, wherein the first voice signal is a sound signal obtained from the ear canal of the wearer, and the second voice signal is a sound signal obtained from the environment in which the wearer is located;
[0070] Step S102: obtaining a first energy parameter and a second energy parameter, wherein the first energy parameter represents the energy value of the first speech signal in a preset frequency band, and the second energy parameter represents the energy value of the second speech signal in a preset frequency band, wherein the preset frequency band is a frequency interval where the occlusion effect occurs;
[0071] Step S103: obtaining a numerical relationship between the first energy parameter and the second energy parameter. If the numerical relationship indicates that the first energy parameter is greater than the second energy parameter, it is determined that the wearer has emitted a sound signal.
[0072] In step S101 , due to the structural characteristics of the in-ear earphone, voice signals can be acquired from the ear canal of the wearer and the environment in which the wearer is located. Figure 2 : is a schematic structural diagram of an in-ear TWS headset according to an exemplary embodiment. Figure 2 As shown in the figure, the acoustic components of the in-ear headphones mainly include a feedforward microphone 1, a feedback microphone 2 and a call microphone 3. Among them, the feedforward microphone 1 is placed outside the wearer's auricle and can obtain the sound signal of the wearer's environment. The feedback microphone 2 is placed inside the wearer's ear canal and can obtain the sound signal in the wearer's ear canal. The call microphone 3 is placed on the earphone handle and can obtain the wearer's sound signal to perform related operations on the wearer's sound signal. Figure 2 The feedback microphone of the headset shown in the figure obtains a first speech signal to represent the sound signal in the wearer's ear canal, and obtains a second speech signal from the feedforward microphone to represent the sound signal in the wearer's environment. Figure 2 In addition to wireless in-ear headphones, any wired in-ear headphones can use the voice detection method of the present disclosure as long as they are equipped with a feedforward microphone and a feedback microphone.
[0073] It should be noted that since the present disclosure is intended to determine whether the headphone wearer has emitted a sound signal, when the wearer plays music or videos while wearing the headphones, an echo cancellation algorithm is required to remove interference from the music or video sound signal and extract the human voice signal from the acquired sound signal as the acquired voice signal. In addition, to ensure the clarity of the extracted voice signal, noise reduction processing can be performed on the extracted voice signal.
[0074] In step S102, when the user is not wearing an in-ear TWS headset, the sound signal emitted will be transmitted to the ear through the bones, and then diffused to the outside of the ear through the ear canal. However, when wearing an in-ear headset, the sound signal will be blocked in the ear by the headset after being transmitted to the ear through the bones, and the sound signal in the ear will be strengthened. Therefore, when the user wears an in-ear headset, the energy of the wearer's sound signal obtained from the ear canal will be greater than the energy of the wearer's sound signal obtained from the outside of the ear. This phenomenon is called the occlusion effect. Since the occlusion effect only occurs in the frequency band of 200Hz to 500Hz, the frequency range in which the occlusion effect occurs is determined as a preset frequency band, which is 200Hz to 500Hz. The preset frequency band includes a starting frequency point of 200Hz and an ending frequency point of 500Hz. Obtain a first energy parameter of the first voice signal in a preset frequency band, and a second energy parameter of the second voice signal in the preset frequency band. The first energy parameter and the second energy parameter are parameters that can reflect the energy characteristics of the first voice signal and the second voice signal in the preset frequency band, and may include the energy value of the voice signal corresponding to a special frequency point in the preset frequency band, and the average energy value in the preset frequency band, such as the energy values of the voice signal at the starting frequency point and the ending frequency point in the preset frequency band.
[0075] In step S103, due to the occlusion effect in the preset frequency band, when the user speaks while wearing in-ear headphones, a portion of the sound signal will be conducted to the ear canal through the bones and strengthened. Therefore, the energy of the sound signal in the ear canal is greater than the energy of the sound signal in the external ear environment, that is, the first energy parameter of the first voice signal is greater than the second energy parameter of the second voice signal. After obtaining the first energy parameter and the second energy parameter, the numerical relationship between the first energy parameter and the second energy parameter is calculated. The numerical relationship can be any numerical relationship that can characterize the size between the first energy parameter and the second energy parameter, such as by numerical size or by the positive or negative difference or other compound operations. If the numerical relationship between the first energy parameter and the second energy parameter characterizes that the first energy parameter is greater than the second energy parameter, it means that the wearer is speaking while wearing headphones, that is, the wearer has emitted a sound signal.
[0076] In an exemplary embodiment of the present disclosure, when a user wears in-ear headphones, a first voice signal is obtained from the wearer's ear canal, and a second voice signal is obtained from the wearer's environment. A frequency range resulting from the occlusion effect is selected as a preset frequency band, and a first energy parameter of the first voice signal in the preset frequency band and a second energy parameter of the second voice signal in the preset frequency band are obtained. A numerical relationship between the first energy parameter and the second energy parameter is obtained. If the numerical relationship indicates that the first energy parameter is greater than the second energy parameter, it is determined that the wearer has emitted a sound signal. Without requiring additional sensor hardware and without requiring the headphones to form a microphone array, interference from external environmental sounds can be avoided, accurately determining whether the headphone wearer has emitted a sound signal.
[0077] In an exemplary embodiment of the present disclosure, obtaining the first energy parameter of the first speech signal in a preset frequency band in step S102 includes:
[0078] Obtain a first starting energy value, a first ending energy value, and a first average energy value of the first speech signal, wherein the first starting energy value represents the energy value of the first speech signal at the starting frequency point, the first ending energy value represents the energy value of the first speech signal at the ending frequency point, and the first average energy value represents the average energy value of the first speech signal in a preset frequency band.
[0079] The preset frequency band includes a starting frequency point and an ending frequency point. When the preset frequency band is 200Hz to 500Hz, the starting frequency point is 200Hz and the ending frequency point is 500Hz. The first energy parameter includes a first starting energy value, a first ending energy value, and a first average energy value. The first starting energy value is the energy of the first speech signal at the starting frequency point of the preset frequency band. When the starting frequency point is 200Hz, the first starting energy value is recorded as A 200HzThe first end energy value is the energy of the first speech signal at the end frequency point of the preset frequency band. When the end frequency point is 500Hz, the first end energy value is recorded as A 500Hz The first average energy value is the average energy of the first speech signal in the preset frequency band. When the preset frequency band is 200Hz to 500Hz, the calculation method of the first average energy can be set according to actual needs. It can be the average value of the energy sum of each frequency point in the preset frequency band, or the average value of the energy sum of multiple frequency points in the preset frequency band. The first average energy value is recorded as A 200Hz~500Hz .
[0080] In one embodiment, the first average energy value of the first speech signal in the preset frequency band is obtained by:
[0081] In the preset frequency band, a first reference frequency point is set at every interval of a first preset frequency difference, and the preset frequency band includes a plurality of first reference frequency points;
[0082] Obtaining a first reference energy value for each first reference frequency point in the first speech signal respectively;
[0083] Each first reference energy value is summed and the average value is taken as the first average energy value.
[0084] In a preset frequency band, a first reference frequency point is set at intervals of a first preset frequency difference. The first preset frequency difference can be set according to actual needs. The preset frequency band includes multiple first reference frequency points. A larger value of the first preset frequency difference and a greater number of first reference frequency points result in a greater computational effort and higher accuracy. After obtaining the multiple first reference frequency points, a first reference energy value is obtained for each first reference frequency point in the first speech signal. The sum of each first reference energy value is then averaged to obtain the first average energy value.
[0085] When the preset frequency range is 200Hz to 500Hz, for example, the first preset frequency difference is 60Hz, the preset frequency range includes 5 first reference frequency points, and the corresponding first reference energy values are recorded as E11, E12, E13, E14, and E15 respectively. The first average energy value A 200Hz~500Hz =(E11+E12+E13+E14+E15) / 5.
[0086] In an exemplary embodiment of the present disclosure, obtaining the second energy parameter of the second speech signal in a preset frequency band in step S102 includes:
[0087] Obtain a second starting energy value, a second ending energy value, and a second average energy value of the second speech signal, wherein the second starting energy value represents the energy value of the second speech signal at the starting frequency point, the second ending energy value represents the energy value of the second speech signal at the ending frequency point, and the second average energy value represents the average energy value of the second speech signal in a preset frequency band.
[0088] When the preset frequency band is 200Hz to 500Hz, the starting frequency point is 200Hz and the ending frequency point is 500Hz. The second energy parameter includes a second starting energy value, a second ending energy value, and a second average energy value. The second starting energy value is the energy of the second speech signal at the starting frequency point of the preset frequency band. When the starting frequency point is 200Hz, the second starting energy value is recorded as B 200Hz The second end energy value is the energy of the second speech signal at the end frequency point of the preset frequency band. When the end frequency point is 500Hz, the second end energy value is recorded as B 500Hz The second average energy value is the average energy of the second speech signal in the preset frequency band. When the preset frequency band is 200Hz to 500Hz, the calculation method of the second average energy can be set according to actual needs. It can be the average value of the energy sum of each frequency point in the preset frequency band, or the average value of the energy sum of multiple frequency points in the preset frequency band. The second average energy value is recorded as B 200Hz~500Hz .
[0089] In one embodiment, the second average energy value of the second speech signal in the preset frequency band is obtained by:
[0090] In the preset frequency band, a second reference frequency point is set at intervals of a second preset frequency difference, and the preset frequency band includes a plurality of second reference frequency points;
[0091] Obtaining a second reference energy value for each second reference frequency point in the second speech signal respectively;
[0092] Each second reference energy value is summed and the average value is taken as the second average energy value.
[0093] In the preset frequency band, a second reference frequency point is set at intervals of a second preset frequency difference. The second preset frequency difference can be set according to actual needs. The preset frequency band contains multiple second reference frequency points. A larger second preset frequency difference value and a greater number of second reference frequency points result in a greater computational effort and higher accuracy. After obtaining the multiple second reference frequency points, a second reference energy value is obtained for each second reference frequency point in the second speech signal. The sum of each second reference energy value is then averaged to obtain the second average energy value.
[0094] When the preset frequency range is 200Hz to 500Hz, for example, the second preset frequency difference is 60Hz, the preset frequency range includes 5 second reference frequency points, and the corresponding second reference energy values are recorded as E21, E22, E23, E24, and E25 respectively. The second average energy value B 200Hz~500Hz =(E21+E22+E23+E24+E25) / 5.
[0095] In an exemplary embodiment of the present disclosure, the method for obtaining the numerical relationship between the first energy parameter and the second energy parameter in step S103 includes:
[0096] Obtain a first energy difference value, a second energy difference value and an average energy difference value, wherein the first energy difference value represents the energy difference between the first starting energy value and the second starting energy value, the second energy difference value represents the energy difference between the first ending energy value and the second ending energy value, and the average energy difference value represents the energy difference between the first average energy value and the second average energy value.
[0097] Obtain the first energy difference between the first starting energy value and the second starting energy value, and record the first starting energy value as A 200Hz , the second starting energy value is recorded as B 200Hz , the first energy difference is recorded as Δ1, then Δ1=A 200Hz -B 200Hz ; Obtain the second energy difference between the first termination energy value and the second termination energy value, the first termination energy value is recorded as A 500Hz , the second termination energy value is recorded as B 500Hz , the second energy difference is recorded as Δ2, then Δ2=A 500Hz -B 500Hz ; Obtain the average energy difference between the first average energy value and the second average energy value, and the first average energy value is recorded as A 200Hz~500Hz , the second average energy value is recorded as B 200Hz~500Hz , the average energy difference is recorded as Δ, then Δ=A 200Hz~500Hz -B 200Hz~500Hz .
[0098] In an exemplary embodiment of the present disclosure, determining whether the wearer emits a sound signal according to the numerical relationship between the first energy parameter and the second energy parameter in step S103 includes the following two situations:
[0099] The first type: if the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value, it is determined that the wearer has emitted a sound signal;
[0100] The second type: if the first energy difference is less than or equal to the second energy difference, and / or the average energy difference is less than or equal to a preset reference value, it is determined that the wearer does not emit a sound signal.
[0101] The preset reference value is recorded as F, and the specific value can be determined according to actual needs, for example, 6dB. If the first energy difference Δ1 is greater than the second energy difference Δ2, and the average energy difference Δ is greater than the preset reference value F, it is determined that the wearer has emitted a sound signal, that is, when (A 200Hz -B 200Hz )>(A 500Hz -B 500Hz ), and (A 200Hz~500Hz -B 200Hz~500Hz )>F, it is determined that the wearer has issued a sound signal. After determining that the wearer has issued a sound signal, the sound transmission function or the voice recognition function can be turned on to collect the wearer's voice information and execute corresponding instructions according to the collected voice information. If the first energy difference Δ1 is less than or equal to the second energy difference Δ2, and / or the average energy difference Δ is less than or equal to the preset reference value F, it is determined that the wearer has not issued a sound signal, that is, when (A 200Hz -B 200Hz )>(A 500Hz -B 500Hz ) and (A 200Hz~500Hz -B 200Hz~500Hz )>F, it is determined that the wearer has not issued a sound signal. When it is determined that the wearer has not issued a sound signal, the headset maintains the current voice signal receiving state and detects in real time whether the wearer has issued a sound signal.
[0102] In an exemplary embodiment of the present disclosure, a voice detection device is provided for use in an in-ear headset. Figure 3 FIG. 1 is a block diagram of a speech detection device according to an exemplary embodiment. Figure 3 As shown, the voice detection device includes:
[0103] An acquisition module 301 is configured to acquire a first voice signal and a second voice signal, wherein the first voice signal is a sound signal acquired from the ear canal of the wearer, and the second voice signal is a sound signal acquired from the environment in which the wearer is located;
[0104] The calculation module 302 is configured to obtain a first energy parameter and a second energy parameter, wherein the first energy parameter represents an energy value of the first speech signal in a preset frequency band, and the second energy parameter represents an energy value of the second speech signal in a preset frequency band, wherein the preset frequency band is a frequency interval where the occlusion effect occurs;
[0105] The determination module 303 is configured to obtain a numerical relationship between the first energy parameter and the second energy parameter, and determine that the wearer has emitted a sound signal if the numerical relationship indicates that the first energy parameter is greater than the second energy parameter.
[0106] In an exemplary embodiment, the preset frequency band includes a starting frequency point and an ending frequency point;
[0107] The calculation module 302 is further configured to:
[0108] Obtaining a first starting energy value, a first ending energy value, and a first average energy value of the first speech signal, wherein the first starting energy value represents the energy value of the first speech signal at a starting frequency point, the first ending energy value represents the energy value of the first speech signal at an ending frequency point, and the first average energy value represents the average energy value of the first speech signal in a preset frequency band;
[0109] The calculation module 302 is further configured to:
[0110] Obtain a second starting energy value, a second ending energy value, and a second average energy value of the second speech signal, wherein the second starting energy value represents the energy value of the second speech signal at the starting frequency point, the second ending energy value represents the energy value of the second speech signal at the ending frequency point, and the second average energy value represents the average energy value of the second speech signal in a preset frequency band.
[0111] In an exemplary embodiment, the calculation module 302 is further configured to:
[0112] In the preset frequency band, a first reference frequency point is set at every interval of a first preset frequency difference, and the preset frequency band includes a plurality of first reference frequency points;
[0113] Obtaining a first reference energy value for each first reference frequency point in the first speech signal respectively;
[0114] summing each first reference energy value and taking the average value as the first average energy value;
[0115] The calculation module 302 is further configured to:
[0116] In the preset frequency band, a second reference frequency point is set at intervals of a second preset frequency difference, and the preset frequency band includes a plurality of second reference frequency points;
[0117] Obtaining a second reference energy value for each second reference frequency point in the second speech signal respectively;
[0118] Each second reference energy value is summed and the average value is taken as the second average energy value.
[0119] In an exemplary embodiment, the determination module 303 is further configured to:
[0120] Obtaining a first energy difference value, a second energy difference value, and an average energy difference value, wherein the first energy difference value represents an energy difference between a first starting energy value and a second starting energy value, the second energy difference value represents an energy difference between a first ending energy value and a second ending energy value, and the average energy difference value represents an energy difference between a first average energy value and a second average energy value;
[0121] It is determined that the wearer has emitted a sound signal according to the first energy difference value, the second energy difference value and the average energy difference value.
[0122] In an exemplary embodiment, the determination module 303 is further configured to:
[0123] If the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value, it is determined that the wearer has emitted a sound signal.
[0124] In an exemplary embodiment, the determination module 303 is further configured to:
[0125] If the first energy difference is less than or equal to the second energy difference, and / or the average energy difference is less than or equal to a preset reference value, it is determined that the wearer does not emit a sound signal.
[0126] In an exemplary embodiment, the preset frequency range is 200 Hz to 300 Hz.
[0127] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0128] The present disclosure also provides an earphone, which includes a processor and a memory for storing executable instructions of the processor, and the processor is configured to execute the voice detection method shown in the embodiment of the present disclosure.
[0129] Figure 4 is a block diagram showing a speech detection device 400 according to an exemplary embodiment.
[0130] Reference Figure 4 , apparatus 400 may include one or more of the following components: a processing component 402 , a memory 404 , a power component 406 , a multimedia component 408 , an audio component 410 , an input / output (I / O) interface 412 , a sensor component 414 , and a communication component 416 .
[0131] Processing component 402 generally controls the overall operation of device 400, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. Processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the above-described method. In addition, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.
[0132] The memory 404 is configured to store various types of data to support operations on the device 400. Examples of such data include instructions for any application or method operating on the device 400, contact data, phone book data, messages, pictures, videos, etc. The memory 404 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0133] The power supply component 406 provides power to the various components of the device 400. The power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 400.
[0134] The multimedia component 408 includes a screen that provides an output interface between the device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. When the device 400 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0135] The audio component 410 is configured to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC) that is configured to receive external audio signals when the device 400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 404 or transmitted via the communication component 416. In some embodiments, the audio component 410 also includes a speaker for outputting audio signals.
[0136] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0137] The sensor assembly 414 includes one or more sensors for providing various aspects of the status assessment of the device 400. For example, the sensor assembly 414 can detect the open / closed state of the device 400, the relative positioning of components, such as the display and keypad of the device 400. The sensor assembly 414 can also detect changes in the position of the device 400 or a component of the device 400, the presence or absence of user contact with the device 400, the orientation or acceleration / deceleration of the device 400, and temperature changes of the device 400. The sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 414 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 414 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0138] The communication component 416 is configured to facilitate wired or wireless communication between the device 400 and other devices. The device 400 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0139] In an exemplary embodiment, the apparatus 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.
[0140] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions. The instructions can be executed by the processor 420 of the apparatus 400 to perform the above-described speech detection method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0141] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a device, enables the device to perform the voice detection method in the above embodiment.
[0142] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.
[0143] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A speech detection method, characterized in that: Applied to in-ear headphones, the voice detection method includes: Acquire a first voice signal and a second voice signal, wherein the first voice signal is a sound signal acquired from the wearer's ear canal, and the second voice signal is a sound signal acquired from the wearer's environment; Obtaining a first energy parameter and a second energy parameter, wherein the first energy parameter represents an energy value of the first speech signal in a preset frequency band, and the second energy parameter represents an energy value of the second speech signal in the preset frequency band, wherein the preset frequency band is a frequency interval where an occlusion effect occurs; obtaining a numerical relationship between the first energy parameter and the second energy parameter, and determining that the wearer has emitted a sound signal if the numerical relationship indicates that the first energy parameter is greater than the second energy parameter; Among them, the first energy parameter is greater than the second energy parameter, which includes: the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value; the first energy difference represents the difference between the energy value of the first voice signal at the starting frequency point and the energy value of the second voice signal at the starting frequency point, the second energy difference represents the difference between the energy value of the first voice signal at the ending frequency point and the energy value of the second voice signal at the ending frequency point, and the average energy difference represents the difference between the average energy value of the first voice signal in the preset frequency band and the average energy value of the second voice signal in the preset frequency band.
2. The speech detection method according to claim 1, wherein: The preset frequency band includes a starting frequency point and an ending frequency point; The obtaining of the first energy parameter includes: obtaining a first starting energy value, a first ending energy value, and a first average energy value of the first speech signal, wherein the first starting energy value represents the energy value of the first speech signal at the starting frequency point, the first ending energy value represents the energy value of the first speech signal at the ending frequency point, and the first average energy value represents the average energy value of the first speech signal in the preset frequency band; The obtaining of the second energy parameter includes: obtaining a second starting energy value, a second ending energy value and a second average energy value of the second speech signal, wherein the second starting energy value represents the energy value of the second speech signal at the starting frequency point, the second ending energy value represents the energy value of the second speech signal at the ending frequency point, and the second average energy value represents the average energy value of the second speech signal in the preset frequency band.
3. The speech detection method according to claim 2, wherein: The obtaining of the first average energy value includes: In the preset frequency band, a first reference frequency point is set at intervals of a first preset frequency difference, and the preset frequency band includes a plurality of the first reference frequency points; Obtaining a first reference energy value for each first reference frequency point in the first speech signal respectively; summing each of the first reference energy values and taking an average value as the first average energy value; The obtaining of the second average energy value includes: In the preset frequency band, a second reference frequency point is set at intervals of a second preset frequency difference, and the preset frequency band includes a plurality of second reference frequency points; Obtaining a second reference energy value for each second reference frequency point in the second speech signal respectively; The second reference energy values are summed and the average value is taken as the second average energy value.
4. The speech detection method according to claim 3, wherein: If the numerical relationship indicates that the first energy parameter is greater than the second energy parameter, determining that the wearer has emitted a sound signal includes: Obtaining a first energy difference value, a second energy difference value, and an average energy difference value, wherein the first energy difference value represents an energy difference between the first starting energy value and the second starting energy value, the second energy difference value represents an energy difference between the first ending energy value and the second ending energy value, and the average energy difference value represents an energy difference between the first average energy value and the second average energy value; According to the first energy difference value, the second energy difference value and the average energy difference value, it is determined that the wearer has emitted a sound signal.
5. The speech detection method according to claim 4, characterized in that The determining, based on the first energy difference, the second energy difference, and the average energy difference, that the wearer has emitted a sound signal includes: If the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value, it is determined that the wearer has emitted a sound signal.
6. The speech detection method according to claim 5, characterized in that The voice detection method further comprises: If the first energy difference is less than or equal to the second energy difference, and / or the average energy difference is less than or equal to the preset reference value, it is determined that the wearer does not emit a sound signal.
7. The speech detection method according to any one of claims 1 to 6, characterized in that: The preset frequency band is 200 Hz to 500 Hz.
8. A speech detection device, characterized in that: Applied to in-ear headphones, the voice detection device includes: an acquisition module configured to acquire a first voice signal and a second voice signal, wherein the first voice signal is a sound signal acquired from the wearer's ear canal, and the second voice signal is a sound signal acquired from the wearer's environment; a calculation module configured to obtain a first energy parameter and a second energy parameter, wherein the first energy parameter represents an energy value of the first speech signal in a preset frequency band, and the second energy parameter represents an energy value of the second speech signal in the preset frequency band, wherein the preset frequency band is a frequency interval where an occlusion effect occurs; a determination module configured to obtain a numerical relationship between the first energy parameter and the second energy parameter, and determine that the wearer has emitted a sound signal if the numerical relationship indicates that the first energy parameter is greater than the second energy parameter; Among them, the first energy parameter is greater than the second energy parameter, which includes: the first energy difference is greater than the second energy difference, and the average energy difference is greater than a preset reference value; the first energy difference represents the difference between the energy value of the first voice signal at the starting frequency point and the energy value of the second voice signal at the starting frequency point, the second energy difference represents the difference between the energy value of the first voice signal at the ending frequency point and the energy value of the second voice signal at the ending frequency point, and the average energy difference represents the difference between the average energy value of the first voice signal in the preset frequency band and the average energy value of the second voice signal in the preset frequency band.
9. A headset, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the speech detection method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by the processor of the headset, the headset is enabled to execute the voice detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for detecting voice of earphone wearer and storage medium
CN111933140A
Method for reducing earphone blocking effect and related device
CN113132841A