Sound energy processing device and processing method, earphone, and readable storage medium
By setting microphones on the earphones that are biased towards and away from the ear canal, and combining the sound energy detection module and the processing control module, the ambient sound energy is automatically detected and the microphone detection mode is adjusted, which solves the problem of unclear call effects in different environments and enables clear calls in different environments.
Patent Information
- Application Number
- CN202111365119.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-11-17
AI Technical Summary
Existing headphones cannot recognize different environments and cannot collect and process audio according to different environmental requirements, resulting in unclear call effects.
By setting microphones towards and away from the ear canal on the earphones, combined with the sound energy detection module and the processing control module, the ambient sound energy is automatically detected, the environmental information is calculated, and the microphone detection mode is adjusted according to the environmental information and the initial state of the microphone, realizing automatic switching in different environments.
It achieves clear call effects in different environments and improves call efficiency and quality.
Smart Images

Figure CN114038478B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of audio signal processing technology and acoustics, and in particular relates to a sound energy processing device and method, headphones, and a readable storage medium. Background Art
[0002] With changes in lifestyles and work styles, more and more people are wearing headphones for mobile remote communications, such as audio and video calls, conference calls or video conferences, and voice messaging. Furthermore, the people communicating often move around, changing environments, and the noise levels of the living environment. Many people also need to communicate outdoors in windy and rainy weather. Different environments have different impacts on communication sound quality, and corresponding audio acquisition and processing solutions are needed to meet the needs of different environments.
[0003] Current headsets generally use universal echo suppression and noise cancellation algorithms to ensure relatively clear calls in noisy environments, but they cannot identify different environments and cannot provide corresponding audio collection and processing methods for different environmental requirements. Summary of the Invention
[0004] To address the problem in existing earphones that are unable to identify different environments and thus perform targeted audio collection and processing, the present invention proposes a sound energy processing device and method, earphones, and a readable storage medium. These devices detect the energy of the wearer's environment and automatically switch to different call modes based on the energy level, ensuring clear calls at all times and improving interpersonal communication efficiency. The technical solution is as follows:
[0005] In one aspect, the present invention provides a sound energy processing device, comprising:
[0006] A sound energy detection module is used to pick up ambient sound through a microphone, detect the sound energy of the ambient sound, generate energy information, and calculate the environmental information of the environment in which the microphone is located based on the energy information;
[0007] The processing control module adjusts the microphone detection mode of the sound energy detection module according to the energy information, the environment information and the initial state information of the microphone;
[0008] The sound energy detection module is further used to switch the microphone detection mode based on the adjustment instruction sent by the processing control module, and record the status information of the microphone's pickup or closed state after switching.
[0009] By detecting the ambient sound through the sound energy detection module, some environmental information that identifies the environmental characteristics is obtained. Then, the ambient sound and the identifying environmental information are analyzed through the processing and control module. The microphone detection mode and earphone status of the sound energy detection module can be adjusted in a targeted manner according to the switching of the environment, thereby automatically switching different call modes according to the different energy levels, so that the call can remain clear at all times.
[0010] Preferably, the sound energy detection module includes:
[0011] The earphones are equipped with at least one internal microphone facing toward the ear canal and at least one external microphone facing away from the ear canal. These microphones receive adjustment instructions from the processing control module to switch detection modes. Multiple microphones ensure diverse sound extraction, enrich the input data for the sound energy detection module, and achieve better detection results. Therefore, multiple microphones can be selected based on cost. Furthermore, the position of the microphones is also strongly correlated with the detection results. Generally, the detection data from the two orientations, facing toward the ear canal and facing away from the ear canal, differ significantly. Therefore, installing microphones in both orientations can achieve better detection results.
[0012] Preferably, the communication between the sound energy detection module and the processing control module adopts the Socket communication protocol or the GPIO communication protocol. Both protocols can realize the real-time transmission of signaling and data between modules, and can also maintain good performance and consistency in complex environments.
[0013] In another aspect, the present invention provides a method for processing sound energy, wherein the method is applied to an earphone, wherein the earphone is provided with at least one internal microphone biased toward the ear canal and at least one external microphone facing away from the ear canal;
[0014] The steps of the treatment method include:
[0015] Detect sound energy, pick up ambient sound through a microphone, detect the sound energy of the ambient sound, and generate energy information;
[0016] Processing energy information to calculate environmental information for identifying the type of environment;
[0017] Get the initial status information of the microphone;
[0018] Determining a target detection mode for switching the microphone according to the environment information and the initial state information;
[0019] Send a control signal to switch the microphone's initial detection mode to target detection mode.
[0020] Furthermore, the environmental information includes first environmental information, second environmental information, and third environmental information. Processing the energy information and calculating the environmental information for identifying the environmental type includes:
[0021] If the energy information includes the magnitude of sound energy at different frequencies, then calculating the signal-to-noise ratio information based on the magnitude of the energy, determining the magnitude of the ambient noise, and determining the first environmental information based on the magnitude of the ambient noise;
[0022] If the energy information includes sounds of different frequencies, identifying the main audio components of the environment in which the microphone is located according to the sound frequencies, and obtaining the second environment information according to the main audio components;
[0023] If the energy information includes the direction of the sound source, the position of the microphone closest to the mouth is determined based on the direction of the sound source, and the third environment information is obtained based on the microphone position.
[0024] Furthermore, the initial state information includes a first preset threshold and a second preset threshold, and the first preset threshold is smaller than the second preset threshold;
[0025] Determining a target detection mode for switching a microphone based on the environment information and the initial state information includes the following steps:
[0026] When the first environmental information is less than a first preset threshold, confirming that the target detection mode is to use only the external microphone for environmental sound detection;
[0027] When the first environmental information belongs to the preset interval, confirming that the target detection mode is to use the external microphone and the internal microphone for environmental sound detection simultaneously;
[0028] When the first environmental information exceeds a second preset threshold, it is confirmed that the target detection mode is to use only the internal microphone for environmental sound detection.
[0029] Further, the initial state information includes an internal position of the internal microphone and an external position of the external microphone;
[0030] Determining a target detection mode for switching a microphone based on the environment information and the initial state information includes the following steps:
[0031] determining a target position corresponding to a microphone position of the third environmental information among the internal position and the external position;
[0032] Determine the target microphone among the internal and external microphones that corresponds to the target location, and confirm that the target detection mode is to use only microphones other than the target microphone for ambient sound detection.
[0033] Furthermore, the step of obtaining the second environmental information according to the audio principal component includes:
[0034] When the main component of the audio is the vocal component, the vocal ratio is calculated based on the vocal component, and the vocal ratio is used as the second environmental information;
[0035] When the main component of the audio is a noise component, the noise ratio is calculated according to the noise component, and the noise ratio is used as the second environmental information.
[0036] Furthermore, the initial state information includes a third preset threshold and a fourth preset threshold, and the fourth preset threshold is greater than the third preset threshold;
[0037] Determining a target detection mode for switching a microphone based on the environment information and the initial state information includes the following steps:
[0038] When the second environmental information is greater than a third preset threshold, confirming that the target detection mode is to use only the external microphone for environmental sound detection;
[0039] When the second environmental information is greater than a fourth preset threshold, it is confirmed that the target detection mode is to use the external microphone and the internal microphone for environmental sound detection at the same time, and to use the external microphone for noise reduction.
[0040] Furthermore, after the step of confirming the target detection mode for switching the microphone according to the environmental information and the initial state information, the following steps are further included:
[0041] According to the initial state information, effective sound is extracted from the environmental sound.
[0042] In another aspect, the present invention further provides an earphone, comprising:
[0043] The sound energy processing device described above comprises at least one processor and a memory in communication with the at least one processor; the memory stores instructions executed by the processor to implement the steps of the above sound energy processing method. The headset can automatically detect sound energy and, based on the detected sound energy, can identify different environments and adapt accordingly, ensuring consistent call quality.
[0044] On the other hand, the present invention provides a readable storage medium, on which a detection program for a sound energy processing device is stored. When the detection program for the sound energy processing device is executed by a processor, the steps of the processing method of the sound energy processing device are implemented.
[0045] The beneficial effects of the present invention are as follows: by using the scheme of the present invention, the sound energy of the environment is detected by picking up the sound of the environment, and the environmental information of the environment is calculated; then the energy information, environmental information and state information are integrated to extract effective sound, and the microphone detection mode of the sound energy detection module and the headphone state are adjusted. Therefore, the scheme of the present invention can detect the environmental sound through the sound energy detection module to obtain some indicative environmental information that identifies the characteristics of the environment, and then analyze the environmental sound and the indicative environmental information through the processing control module, and then the microphone detection mode of the sound energy detection module can be adjusted in a targeted manner according to the switching of the environment, thereby realizing automatic switching of different call modes according to different environments, so that the call can maintain a clear call effect at any time. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a structural diagram of a sound energy processing device according to an embodiment of the present invention;
[0047] Figure 2 is a flow chart of a method for processing sound energy in an embodiment of the present invention;
[0048] Figure 3 is a diagram of the internal modules of the sound energy processing device according to an embodiment of the present invention;
[0049] Figure 4 is a flow chart of another method for processing sound energy according to an embodiment of the present invention;
[0050] Figure 5 It is a structural diagram of a storage medium in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] To facilitate understanding of the present invention, the present invention will be described in more detail below with reference to the accompanying drawings and specific embodiments. Preferred embodiments of the present invention are shown in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of the present invention.
[0052] It should be noted that, unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0053] Example 1
[0054] On the one hand, please refer to Figure 1 , Figure 1This is a structural diagram of a sound energy processing device of the present invention, which can automatically detect the sound energy in the environment and automatically switch to different call modes based on the sound energy information. The sound energy processing device includes:
[0055] The sound energy detection module 1 is used to pick up ambient sound through a microphone, detect the sound energy of the ambient sound, generate energy information, and calculate the environmental information of the environment in which the microphone is located based on the energy information;
[0056] The processing control module 2 adjusts the microphone detection mode of the sound energy detection module 1 according to the energy information, the environment information and the initial state information of the microphone;
[0057] The sound energy detection module 1 is further configured to switch the microphone detection mode based on the adjustment instruction sent by the processing control module 2, and record the state information of the microphone's pickup or closed state after switching.
[0058] Specifically, ambient sound refers to background sounds other than human voices picked up by the microphone during a call. These sounds vary in different environments, such as offices, cafes, train stations, subways, and airports, all with distinct characteristics, such as the frequency band, signal-to-noise ratio, and direction of the sound source. These characteristics are known as sound energy. After calculating this energy information, we can then use this energy information to generate environmental information that characterizes the environment. For example, we can determine whether the environment is dominated by human voices or noise. For environments dominated by human voices, we can further categorize them into quiet, normal, and noisy environments based on the calculated proportion of human voices.
[0059] Specifically, the initial state information of the microphone refers to whether the current microphone is in the pickup or closed state, and whether the environment detection function is used. When there are multiple microphones, the initial state information of multiple microphones needs to be recorded.
[0060] Specifically, the microphone detection mode refers to which microphones are enabled for environment detection, such as using both internal and external microphones for environment detection, or only using external microphones for detection, or not using the detection function of the microphone.
[0061] The microphone detection mode needs to be switched according to the environment, so it also includes an initial detection mode and a target detection mode. The initial detection mode is the currently active microphone detection mode; the generation and use of the target detection mode is performed by the processing control module 2 based on energy information, environmental information, and the initial state information of the microphone. It ultimately determines whether the microphone detection mode needs to be switched. If so, it sends a command to the sound energy detection module 1 to switch the microphone to the target detection mode.
[0062] It can be seen that by detecting the ambient sound through the sound energy detection module 1 to obtain some identifying environmental information that identifies the environmental characteristics, and then analyzing the ambient sound and the identifying environmental information through the processing control module 2, the microphone detection mode and earphone status of the sound energy detection module 1 can be adjusted in a targeted manner according to the switching of the environment, thereby realizing automatic switching of different call modes according to different energies, so that the call can maintain a clear call effect at all times.
[0063] Preferably, the sound energy detection module 1 further comprises: at least one internal microphone positioned toward the ear canal and at least one external microphone positioned away from the ear canal, disposed on the earphone. The internal and external microphones receive adjustment commands from the processing control module 2 to switch between target detection modes. Multiple microphones ensure diverse sound extraction, enrich the input data for the sound energy detection module 1, and achieve better detection results. Therefore, multiple microphones can be selected based on cost. Furthermore, the location of the microphones is also strongly correlated with detection results. For example, since the specific data of ambient sound detected from the same earphone facing toward and away from the ear canal differ significantly, placing microphones in both directions can achieve better detection results. Furthermore, the present invention is not limited to microphones positioned toward and away from the ear canal. Microphones can also be placed in any other location based on specific product requirements to increase the diversity of ambient sound detection data. For example, the fixed radio receiver on the headphone cable of wired headphones, the microphone head of a headset, and the body of a Bluetooth headset all have different sizes, shapes, and circuit designs. Therefore, the location and number of microphones need to be selected based on different product requirements.
[0064] Further, refer to Figure 3 , a typical internal module diagram of a sound energy processing device is given. Specifically, the sound energy detection module 1 includes a collection unit 101, a storage unit 102, and a switching unit 103, and the processing control module 2 includes a processing unit 201, an extraction unit 202, and an adjustment unit 203.
[0065] Among them, the acquisition unit 101 includes at least one internal microphone biased towards the ear canal and at least one external microphone away from the ear canal arranged on the earphone as mentioned above; the acquisition unit 101 is also used to realize the function of controlling the microphone to pick up environmental sounds, detect the sound energy of environmental sounds, and generate energy information. This function can be implemented by hardware, or by software, or by a combination of hardware and software. For example, audio encoding and decoding, audio parameter extraction and calculation, etc., can be implemented by software, or by using hardware support commonly used in the industry. As a transducer for converting sound waves into electrical signals, the technology of microphones is quite mature, and the structural forms are mostly electret condenser type, dynamic type and carbon crystal type. Preferably, the microphone of the present invention can be selected from bone conduction microphones, electret microphones or MEMS digital microphones.
[0066] Storage unit 102 is used to store the raw ambient sound and environmental information collected by acquisition unit 101. This environmental information is derived by performing calculations on data such as the frequency of the raw ambient sound, the direction of the sound source, and the energy level of different frequencies. Specifically, the data is stored in registers and can be accessed in specific data formats, such as data tables, formatted files, and key-value pairs. This data can be read locally and remotely.
[0067] The switching unit 103 is configured to receive the adjustment instruction sent by the processing control module 2 to switch the microphone detection mode, and record the state information of the microphone's pickup or closed state after switching.
[0068] The processing unit 201 is used to read relevant data from the sound energy detection module 1, including energy information, environmental information, and the initial state information of the microphone, and determine the target detection mode for switching the microphone based on this information. In addition to the above active data reading method, the sound energy detection module 1 can also actively report data.
[0069] Extraction unit 202 is used to extract effective ambient sound based on the initial state information. During a call, human voice recognition can be achieved in a variety of ways, and the position and state of the pickup microphone can also affect voice recognition. Therefore, different extraction methods can be selected based on the microphone's enabled or disabled state and whether energy detection is enabled, as reported in the initial state information, to more accurately extract the call audio for transmission to the other party.
[0070] The adjustment unit 203 is configured to send an instruction to the sound energy detection module 1 to adjust the microphone detection mode and earphone status of the sound energy detection module 1 .
[0071] Furthermore, the communication between the sound energy monitoring module 1 and the processing and control module 2 includes multiple modes. They can be set on the same physical medium for local communication; or they can be set on different physical media for remote communication via a communication protocol, including a Socket communication protocol or a GPIO communication protocol. Preferably, the communication between the sound energy detection module 1 and the processing and control module 2 adopts the Socket communication protocol or the GPIO communication protocol. Both protocols can achieve real-time transmission of signaling and data between modules, while also maintaining good performance and consistency in complex environments.
[0072] In summary, the sound energy processing device of the present invention has the following advantages:
[0073] It detects the environmental energy of the wearer in different environments and automatically switches to different call modes according to the different energy levels, so that calls can be clear at all times and improve the efficiency of communication between people.
[0074] Example 2
[0075] On the other hand, please refer to Figure 2 The present invention provides a method for processing sound energy.
[0076] The processing method is applied to headphones, which are provided with at least one internal microphone facing toward the ear canal and at least one external microphone facing away from the ear canal. The internal microphone and the external microphone receive adjustment instructions and switch detection modes.
[0077] The steps of the treatment method include:
[0078] S101 detects sound energy by picking up ambient sound through a microphone, detecting the sound energy of the ambient sound, and generating energy information;
[0079] S102 processes the energy information and calculates environmental information for identifying the environment type;
[0080] S103 obtains initial state information of the microphone;
[0081] S104: confirming a target detection mode for switching the microphone according to the environment information and the initial state information;
[0082] S105 sends a control signal to switch the initial detection mode of the microphone to the target detection mode.
[0083] Furthermore, in step S101, the energy information may include data such as the different frequencies of sounds in the environment, the direction of the sound source, and the magnitude of the sound energy. This energy information may vary depending on the environment, such as an office, a coffee shop, a train station, a subway, or an airport. Some environments may be noisy, while others may be quiet. In this case, the signal-to-noise ratio (SNR) can be calculated and then divided into specific intervals to identify different environments. Furthermore, some environments are dominated by human voices, while others are dominated by noise. In this case, the sound frequency can be calculated and then divided into specific intervals to identify different environments.
[0084] Furthermore, the environmental information includes first environmental information, second environmental information, and third environmental information. For step S102, the energy information is processed and the environmental information used to identify the environmental type is calculated, including:
[0085] If the energy information includes the magnitude of sound energy at different frequencies, the signal-to-noise ratio information is calculated based on the energy magnitude to determine the magnitude of the ambient noise, and the first environmental information is determined based on the magnitude of the ambient noise. Specifically, the first environmental information may be the ambient noise obtained by calculating the signal-to-noise ratio information.
[0086] If the energy information includes sounds of different frequencies, the primary audio component of the microphone's environment is identified based on the sound frequencies, and the second environmental information is obtained based on the primary audio component. Specifically, environmental sounds with frequencies between 300 Hz and 2.5 kHz can be defined as primarily human voices, while environmental sounds with frequencies below 300 Hz and above 2.5 kHz can be defined as primarily noise. The second environmental information can be the percentage of different sound frequencies obtained by calculating the environmental sounds.
[0087] If the energy information includes the direction of the sound source, the position of the microphone closest to the mouth is determined based on the direction of the sound source, and the third environmental information is obtained based on the microphone position. Specifically, the third environmental information can be the position of the closest microphone obtained by calculating the ambient sound, as well as the sound angle at which the microphone picks up the sound.
[0088] Furthermore, in step S103, obtaining the initial state information of the microphone includes:
[0089] On the one hand, it is necessary to detect whether the current microphone is in the pickup or off state, and whether the environmental detection function is used. When there are multiple microphones, the positions of multiple microphones must be recorded, etc. The initial state information is mainly used to confirm the target detection mode and to extract effective information for ambient sound; on the other hand, the initial state information also includes some preset thresholds, which are used when confirming the target detection mode. The threshold size can be adjusted according to different detection methods and different hardware requirements to achieve the best judgment effect.
[0090] Furthermore, the initial state information includes a first preset threshold and a second preset threshold, and the first preset threshold is less than the second preset threshold. In this case, in step S104, determining the target detection mode for switching the microphone based on the environment information and the initial state information includes the following steps:
[0091] When the first environmental information is less than a first preset threshold, confirming that the target detection mode is to use only the external microphone for environmental sound detection;
[0092] When the first environmental information belongs to the preset interval, confirming that the target detection mode is to use the external microphone and the internal microphone for environmental sound detection simultaneously;
[0093] When the first environmental information exceeds a second preset threshold, it is confirmed that the target detection mode is to use only the internal microphone for environmental sound detection, wherein the first preset threshold is less than the second preset threshold.
[0094] Specifically, the first preset threshold can be set to 60db, the preset interval can be set to 65db-85db, and the second preset threshold can be set to 85db. After setting in this way, the environment can be divided into three situations: low ambient noise, normal ambient noise, and high ambient noise.
[0095] When the first environmental information, i.e., the environmental noise, is less than the first preset threshold of 60db, the environmental noise is relatively small. At this time, only the external microphone can be enabled to achieve the expected call quality. Therefore, the target detection mode is confirmed to use only the external microphone for environmental sound detection.
[0096] When the first environmental information, i.e., the ambient noise level, falls within the preset range of 65dB-85dB, the ambient noise level is normal. In this case, both the external and internal microphones can be activated simultaneously. In particular, if multiple external microphones are present, the internal microphone can be disabled and both external microphones can be activated simultaneously. Therefore, the target detection mode is confirmed to use both the external and internal microphones for ambient sound detection.
[0097] When the first environmental information, i.e., the environmental noise, is greater than the second preset threshold of 85dB, the environmental noise is relatively large, and only the internal microphone can be enabled. Therefore, the target detection mode is confirmed to be the use of only the internal microphone for environmental sound detection.
[0098] Furthermore, the initial state information includes the internal position of the internal microphone and the external position of the external microphone; in this case, for step S104, determining the detection mode that needs to be adjusted based on the environmental information and the state information further includes the following steps:
[0099] determining a target position corresponding to a microphone position of the third environmental information among the internal position and the external position;
[0100] Specifically, as previously described, the third environmental information includes the location of the closest microphone, as determined by calculating the ambient sound, and the angle at which the microphone picks up the sound. For example, the angle could be 45° or 90° from the mouth. Based on this location information and the angle, the closest microphone position in the initial state information, i.e., the target position, can be calculated.
[0101] Determine the target microphone among the internal and external microphones that corresponds to the target location, and confirm that the target detection mode is to use only microphones other than the target microphone for ambient sound detection.
[0102] Specifically, the unique identifier of the corresponding target microphone can be calculated based on the target position, so that an instruction can be issued to use the closest microphone to pick up the call, and not use the microphone for ambient sound detection, which can achieve a better sound reception effect; in addition, the sound angle can also be used as a calculation parameter for later sound extraction, so as to focus on extracting the sound at the best sound angle, so the target detection mode is confirmed to use the closest microphone and not use the microphone for ambient sound detection.
[0103] Furthermore, the step of obtaining the second environmental information according to the audio principal component includes:
[0104] When the main component of the audio is the vocal component, the vocal ratio is calculated based on the vocal component, and the vocal ratio is used as the second environmental information;
[0105] When the main component of the audio is a noise component, the noise ratio is calculated according to the noise component, and the noise ratio is used as the second environmental information.
[0106] Furthermore, the initial state information includes a third preset threshold and a fourth preset threshold, and the fourth preset threshold is greater than the third preset threshold;
[0107] Regarding step S104, determining the detection mode that needs to be adjusted based on the environmental information and the state information further includes the following steps:
[0108] When the second environmental information is greater than a third preset threshold, confirming that the target detection mode is to use only the external microphone for environmental sound detection;
[0109] When the second environmental information is greater than a fourth preset threshold, it is confirmed that the target detection mode is to use the external microphone and the internal microphone for environmental sound detection at the same time, and to use the external microphone for noise reduction.
[0110] Specifically, the ambient sound can be divided into those dominated by human voices and those dominated by noise. Sounds with frequencies between 300 Hz and 2.5 kHz can be identified as human voices, while sounds with frequencies below 300 Hz or above 2.5 kHz can be identified as noise. The endpoints of these frequency ranges can be adjusted based on the algorithm and actual test results.
[0111] When the environment is dominated by human voices, the second environmental information is the proportion of human voices, that is, the proportion of the sound energy between 300Hz and 2.5kHz to the total environmental sound energy. The third preset threshold can be set to a ratio greater than 50%, for example, 80%. When the second environmental information at this time is greater than the third preset threshold, it means that the human voice accounts for a large proportion of the environmental sound, which should be a relatively quiet environment. At this time, the target detection mode is to use only the external microphone for environmental sound detection.
[0112] When the environment is dominated by noise, the second environmental information is the noise ratio, that is, the ratio of the sound energy of sounds with a frequency lower than 300hz or higher than 2.5khz to the total environmental sound energy. The fourth preset threshold can be set to a ratio value greater than 50%, for example, 70%. When the second environmental information at this time is greater than the fourth preset threshold, it means that the noise ratio in the environmental sound is large, and it should be in a relatively noisy environment. At this time, it is confirmed that the target detection mode is to use both the external microphone and the internal microphone for environmental sound detection, and use the external microphone for noise reduction.
[0113] Further, refer to Figure 4 After determining the target detection mode for switching the microphone according to the environment information and the initial state information, the method further includes the following steps:
[0114] S106 extracts effective sounds from the ambient sounds according to the initial state information.
[0115] During a call, voice recognition involves many methods, and the position and state of the pickup microphone can also affect voice recognition. Therefore, based on the initial status information (e.g., whether the microphone is enabled or disabled, and whether energy detection is enabled), different extraction methods can be selected to extract more accurate call audio for transmission to the other party.
[0116] Specifically, referring to the microphone settings mentioned above, if both the internal and external microphones are enabled, the ambient sound picked up by the internal microphone closest to the mouth will be used for call audio extraction, while the ambient sound picked up by the other microphones will be ignored. This allows the most accurate call audio to be extracted and sent to the other party.
[0117] In summary, the sound energy processing method of the present invention has the following advantages:
[0118] By picking up the sounds of the environment to detect the sound energy of the environment, the environmental information is calculated; then the energy information, environmental information and status information are combined to extract effective sounds, and the microphone detection mode and earphone status of the sound energy detection module 1 are adjusted, so as to achieve the effect of maintaining clear calls when the caller enters different environments.
[0119] Example 3
[0120] In another aspect, the present invention further provides a headset for automatically detecting sound energy, comprising:
[0121] The sound energy processing device as described above, at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executed by the processor to implement the steps of the sound energy processing method in the embodiment as described above.
[0122] This headset can automatically detect sound energy. It can identify different environments based on the detected sound energy and adapt to them to keep the call quality consistent.
[0123] Specifically, headphones include wired headphones, wireless headphones, headsets, etc.
[0124] As mentioned above, since the specific data of the ambient sound detected in the two directions of the same earphone, toward the ear canal and away from the ear canal, are quite different, better detection results can be achieved by setting microphones in both directions. Furthermore, the present invention is not limited to setting microphones in the two directions of toward the ear canal and away from the ear canal. Microphones can also be set in any other position according to specific product requirements to increase the diversity of the detection data of ambient sound. For example, the fixed radio device on the headphone cable of wired headphones, the radio microphone head of earphones, the body of Bluetooth headphones, etc. all have different sizes and shapes as well as circuit designs. Therefore, it is necessary to select the position and number of microphones according to different product requirements.
[0125] Example 4
[0126] On the other hand, the present invention provides a readable storage medium, on which a detection program for a sound energy processing device is stored. When the detection program for the sound energy processing device is executed by a processor, the steps of the processing method of the sound energy processing device are implemented.
[0127] In practical applications, Figure 5 It is a structural diagram of the hardware operating environment involved in the readable storage medium of the present invention.
[0128] like Figure 5 As shown, the hardware operating environment may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0129] Those skilled in the art will understand that Figure 5 The hardware structure of the control method for the energy storage charging system shown in the figure does not constitute a limitation on the equipment for operating the control method for the energy storage charging system, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0130] like Figure 5 As shown, memory 1005, a readable storage medium, may include an operating system, a network communication module, a user interface module, and a control program for the energy storage and charging system. The operating system is a management and control program that supports the operation of the network communication module, the user interface module, the control program for the energy storage and charging system, and other programs or software. The network communication module is used to manage and control the network interface 1004; the user interface module is used to manage and control the user interface 1003.
[0131] exist Figure 5 In the hardware structure shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with the backend server; the user interface 1003 is mainly used to connect to the client (user end) and communicate data with the client; the processor 1001 can call the control program of the energy storage charging system stored in the memory 1005 and execute the various method processes involved in the aforementioned energy storage charging system control method.
[0132] The above are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied to other related technical fields, are included in the patent protection scope of the present invention.
Claims
1. A sound energy processing device, characterized in that: include: A sound energy detection module, the sound energy detection module comprising a plurality of microphones, the plurality of microphones comprising at least one internal microphone disposed on the earphone and biased toward the ear canal, and at least one external microphone disposed away from the ear canal; the sound energy detection module being configured to pick up ambient sound by controlling a detection mode of the plurality of microphones, detect the sound energy of the ambient sound, generate energy information, and calculate environmental information of the environment in which the plurality of microphones are located based on the energy information; a processing control module, adjusting the microphone detection mode of the sound energy detection module according to the environmental information and the initial state information of the multiple microphones; the microphone detection mode includes an initial detection mode and a target detection mode; the internal microphone and the external microphone receive the adjustment instruction of the processing control module and switch the detection mode; The sound energy detection module is further configured to switch the microphone detection mode based on the adjustment instruction sent by the processing control module, switch the initial detection mode of the microphone to the target detection mode, and record the state information of the pickup or closed state of the microphone after the switching; The microphone detection mode includes any one of the following: enabling both the internal microphone and the external microphone for ambient sound detection, enabling only the internal microphone or only the external microphone for ambient sound detection, and not enabling the microphone for ambient sound detection.
2. The sound energy processing device according to claim 1, characterized in that: The communication between the sound energy detection module and the processing control module adopts the Socket communication protocol or the GPIO communication protocol.
3. A method for processing sound energy, characterized in that: The processing method is applied to headphones, wherein a plurality of microphones are provided on the headphones, including at least one internal microphone biased toward the ear canal and at least one external microphone facing away from the ear canal; The steps of the processing method include: detecting ambient sound energy, picking up ambient sound through the internal microphone and / or the external microphone, detecting sound energy of the ambient sound, and generating energy information; Processing the energy information to calculate environmental information for identifying an environment type; Acquiring initial state information of the multiple microphones; determining, based on the environment information and the initial state information, a target detection mode for switching the plurality of microphones; Sending a control signal to switch the initial detection mode of the plurality of microphones to the target detection mode; The target detection mode includes any one of the following: enabling both the internal microphone and the external microphone for environmental detection, enabling only the internal microphone or only the external microphone for environmental sound detection, and not enabling the microphone for environmental sound detection.
4. The processing method according to claim 3, characterized in that The environmental information includes first environmental information, second environmental information, and third environmental information, and processing the energy information to calculate environmental information for identifying the environmental type includes: If the energy information includes the magnitude of sound energy at different frequencies, calculating signal-to-noise ratio information based on the magnitude of the energy, determining the magnitude of the ambient noise, and determining the first environmental information based on the magnitude of the ambient noise; If the energy information includes sounds of different frequencies, identifying a main audio component of the environment in which the enabled microphone is located according to the sound frequency, and obtaining second environment information according to the main audio component; If the energy information includes the direction of the sound source, the position of the activated microphone closest to the mouth is determined according to the direction of the sound source, and the third environmental information is obtained according to the microphone position.
5. The processing method according to claim 4, characterized in that: The initial state information includes a first preset threshold and a second preset threshold, and the first preset threshold is smaller than the second preset threshold; Determining, based on the environment information and the initial state information, a target detection mode for switching the plurality of microphones comprises the following steps: When the first environmental information is less than a first preset threshold, confirming that the target detection mode is to use only the external microphone for environmental sound detection; When the first environmental information belongs to a preset interval, confirming that the target detection mode is to use the external microphone and the internal microphone to perform environmental sound detection simultaneously; When the first environmental information exceeds a second preset threshold, it is confirmed that the target detection mode is to use only the internal microphone for environmental sound detection.
6. The processing method according to claim 4, characterized in that: The initial state information includes an internal position of the internal microphone and an external position of the external microphone; Determining, based on the environment information and the initial state information, a target detection mode for switching the plurality of microphones comprises the following steps: determining a target position, among the internal position and the external position, corresponding to a microphone position of the third environmental information; A target microphone corresponding to the target position among the internal microphone and the external microphone is determined, and it is confirmed that the target detection mode is to perform ambient sound detection using only microphones other than the target microphone.
7. The processing method according to claim 4, characterized in that: The step of obtaining the second environmental information according to the audio main component includes: When the main component of the audio is a vocal component, calculating a vocal ratio according to the vocal component, and using the vocal ratio as the second environmental information; When the audio main component is a noise component, a noise ratio is calculated according to the noise component, and the noise ratio is used as the second environmental information.
8. The processing method according to claim 4, characterized in that: The initial state information includes a third preset threshold and a fourth preset threshold, and the fourth preset threshold is greater than the third preset threshold; Determining, based on the environment information and the initial state information, a target detection mode for switching the plurality of microphones comprises the following steps: When the second environmental information is greater than a third preset threshold, confirming that the target detection mode is to use only the external microphone for environmental sound detection; When the second environmental information is greater than a fourth preset threshold, it is confirmed that the target detection mode is to use the external microphone and the internal microphone simultaneously for environmental sound detection, and use the external microphone for noise reduction.
9. The processing method according to any one of claims 3 to 8, characterized in that: After the step of determining the target detection mode for switching the plurality of microphones according to the environmental information and the initial state information, the following steps are further included: Effective sound is extracted from the environmental sound according to the initial state information.
10. A headset, characterized in that: include: The sound energy processing device according to any one of claims 1 to 2, at least one processor, and a memory communicatively connected to the at least one processor; The memory stores instructions executed by the processor to implement the steps of the processing method according to any one of claims 3 to 9.
11. A readable storage medium, characterized in that: The readable storage medium stores a detection program for a sound energy processing device, and when the detection program for the sound energy processing device is executed by a processor, the steps of the processing method according to any one of claims 3 to 9 are implemented.
Citation Information
Patent Citations
Earphone noise reduction method and device
CN111327985A