Audio processing method of open earphone, open earphone, medium and product
By echo cancellation, audio analysis, sound effect processing and reverb processing of the audio data collected by open headphones, the problem of the lack of live karaoke sound effects in the headphones by the karaoke software, achieving better user experience and sound effect presentation.
Patent Information
- Application Number
- CN202510027986.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-16
AI Technical Summary
The existing karaoke software lacks the live feeling of karaoke sound when using headphones, and users cannot feel the audio effect of the audio system.
Through the audio processing method of open headphones, the original audio data collected by the microphone is obtained, echo cancellation, audio analysis, sound effect processing and reverb processing are performed, target audio data is generated, and output through the speaker.
The use of open headphones to present karaoke sound effects is achieved, which improves the user experience, eliminates the ear plugging effect, and simulates the audio system during the traditional karaoke process.
Smart Images

Figure CN120018008A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of audio processing technology, and in particular relates to an audio processing method for open-type headphones, open-type headphones, a medium and a product. Background Art
[0002] At present, more and more users like to use Karaoke software to sing Karaoke, and Karaoke software can easily meet users' Karaoke needs.
[0003] However, the above karaoke operation lacks the presence of karaoke sound effects, especially the user cannot feel the audio effects presented by the sound system, so how to use headphones to present better karaoke sound effects is a problem that urgently needs to be solved. Summary of the invention
[0004] The embodiments of the present application provide an audio processing method for open-type headphones, open-type headphones, a medium and a product, which can achieve the effect of presenting karaoke sound effects using open-type headphones and improve user experience.
[0005] In a first aspect, an embodiment of the present application provides an audio processing method for an open-ear headset, comprising:
[0006] Get the raw audio data collected by the microphone of the open headphone;
[0007] Performing echo cancellation processing on the original audio data to obtain first audio processing data;
[0008] Performing audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data, and performing sound effect processing on the first audio processing data to obtain second audio processing data;
[0009] Performing reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data;
[0010] Controls the speaker of the open-back headphone to output target audio data.
[0011] In some embodiments, performing audio analysis on the first audio processing data to obtain a target reverberation parameter of the first audio processing data includes:
[0012] Performing reverberation estimation on the first audio processing data to obtain an original reverberation time of the first audio processing data, where the original reverberation time is used to represent the time it takes for the first audio processing data to decay to a preset energy value;
[0013] Based on the original reverberation time and the preset reverberation time, target reverberation parameters of the first audio processing data are calculated, where the target reverberation parameters at least include a target reverberation intensity.
[0014] In some embodiments, performing reverberation estimation on the first audio processing data to obtain the original reverberation time of the first audio processing data includes:
[0015] determining audio energy of the first audio processing data;
[0016] The audio energy is calculated using a preset reverberation algorithm to obtain the original reverberation time of the first audio processing data.
[0017] In some embodiments, the number of the microphone is one, and accordingly, the number of the original audio data and the number of the first audio processed data are both one;
[0018] Performing sound effect processing on the first audio processing data to obtain second audio processing data includes:
[0019] The first audio processing data is subjected to audio effect processing to obtain second audio processing data, wherein the audio effect processing at least includes howling suppression, equalization processing, excitation processing and compression processing.
[0020] In some embodiments, the number of microphones is at least two, and the number of raw audio data corresponds to the number of microphones;
[0021] Performing echo cancellation processing on the original audio data to obtain first audio processing data includes:
[0022] Performing echo cancellation processing on each of the at least two original audio data respectively to obtain at least two first audio processing data;
[0023] Accordingly, performing audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data includes:
[0024] Audio analysis is performed on any one of the at least two first audio processing data to obtain a target reverberation parameter of the first audio processing data.
[0025] In some embodiments, performing sound effect processing on the first audio processing data to obtain the second audio processing data includes:
[0026] Performing beamforming processing on at least two first audio processing data to obtain beam-processed data, wherein the beamforming processing is used to enhance the audio data in the direction from the user's sound source to the microphone in the at least two first audio processing data, and to weaken the audio data in other directions in the at least two first audio processing data;
[0027] The beam processed data is subjected to howling suppression, equalization processing, excitation processing and compression processing in sequence to obtain second audio processing data.
[0028] In a second aspect, an embodiment of the present application provides an open-type earphone, comprising:
[0029] A microphone, configured in an open-type headset, for collecting raw audio data;
[0030] An echo cancellation module, used for performing echo cancellation processing on the original audio data to obtain first audio processing data;
[0031] A processing module, configured to perform audio analysis on the first audio processing data to obtain a target reverberation parameter of the first audio processing data;
[0032] An audio effect processing module, used to perform audio effect processing on the first audio processing data to obtain second audio processing data;
[0033] A reverberation module, used to perform reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data;
[0034] The output module is used to control the speaker of the open earphone to output the target audio data.
[0035] In some embodiments, the audio processing module includes:
[0036] A beamforming unit, configured to perform beamforming processing on at least two first audio processing data to obtain beam-processed data;
[0037] A howling suppression unit, used for performing howling suppression on the beam processed data;
[0038] An equalizer, used to adjust the frequency response of the audio data after howling suppression;
[0039] an exciter for adding high frequency harmonics to the frequency adjusted audio data;
[0040] A compressor is used to control the volume range of the audio data after being processed by the exciter.
[0041] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements any method of the first aspect.
[0042] In a fourth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an open-ear headset, the open-ear headset executes any one of the methods in the first aspect.
[0043] The embodiment of the present application provides an audio processing method for open-type headphones, open-type headphones, a medium and a product, the method comprising: obtaining original audio data collected by a microphone of the open-type headphones; performing echo cancellation processing on the original audio data to obtain first audio processing data; performing audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data, and performing sound effect processing on the first audio processing data to obtain second audio processing data; performing reverberation processing on the second audio processing data according to the target reverberation parameters to obtain target audio data corresponding to the original audio data; and controlling the speaker of the open-type headphones to output the target audio data. Utilizing the above technical solution, according to the structural characteristics of the open-type headphones, the original audio data can be initially transmitted to the user's ears without any obstacles. At the same time, the open-type headphones perform sound effect processing on the original audio data to obtain the target audio data, so that the output target audio data forms a time difference with the original audio data initially transmitted to the user, thereby presenting the sound effect of karaoke using the open-type headphones and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 It is a flowchart of an audio processing method provided by the prior art;
[0046] Figure 2 It is a flowchart of an audio processing method for an open-type headset provided in one embodiment of the present application;
[0047] Figure 3 is a flowchart of an audio processing method for an open-type headset provided in another embodiment of the present application;
[0048] Figure 4 It is a structural schematic diagram of an audio processing method for an open-type earphone provided in an embodiment of the present application;
[0049] Figure 5 This is a schematic diagram of the time when a user obtains audio data provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic structural diagram of an open-type earphone provided in one embodiment of the present application. DETAILED DESCRIPTION
[0051] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0052] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0053] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0054] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0055] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0056] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0057] It should be noted that the information collection process (such as facial image collection process, audio data collection process, etc.) / feature extraction process involved in this application is performed with the user's knowledge and permission, that is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.
[0058] Figure 1 It is a flowchart of an audio processing method provided by the prior art, such as Figure 1 As shown, in order to achieve karaoke sound effects, the traditional sound system processes the audio data collected by the microphone through a series of sound effects, including equalization, excitation, compression, reverberation and other functional processing, and then outputs the processed audio data through the speaker of the sound system. The combined effect of the above sound effects processing can beautify the audio data collected by the microphone, cover up the defects of the sound to a certain extent, make the user's voice fuller and more pleasant, thereby enhancing the confidence and fun of karaoke, providing high-quality sound effects, and enhancing the user's singing experience.
[0059] It can be understood that the sound received by the user's ears during the above-mentioned traditional karaoke process can be divided into two parts. The first part is the sound emitted from the user's mouth and transmitted to the ear through the bones and air. The second part is the sound output by the sound system after the above-mentioned audio processing.
[0060] In order to meet the demand for karaoke in any venue, if users use karaoke software to sing karaoke while wearing closed headphones, there will be an ear-blocking effect, which will affect the user experience to a certain extent. Open headphones have a natural advantage in simulating sound systems. On the one hand, due to the structural characteristics of open headphones that usually do not need to be worn in the ears, users can use karaoke software to sing karaoke while wearing open headphones to eliminate the problem of ear blockage, because the sound emitted by the user's mouth can be directly transmitted to the user's ears through bones and air without obstacles, completely eliminating the ear-blocking effect. On the other hand, open headphones can simulate the sound system in the traditional karaoke process by adopting specific audio processing methods, and output audio data with similar sound effects, so that users can get a karaoke experience consistent with karaoke. More specifically, the karaoke audio data based on open headphones is based on the superposition of the real-world conducted sound and the sound processed by the headphone sound effects, so that it can more realistically feedback the sound effects similar to karaoke.
[0061] Figure 2 1 is a flow chart of an audio processing method for open-type headphones provided in one embodiment of the present application. As an example but not a limitation, the method can be applied to open-type headphones.
[0062] S101: Acquire original audio data collected by a microphone of an open-ear headset.
[0063] The original audio data may refer to the audio data originally collected by the microphone of the open-ear headset. The number of the original audio data may correspond to the number of microphones, and different microphones respectively collect the original audio data.
[0064] Exemplarily, one or more first microphones may be configured on the outside of the open-type earphones, and a second microphone may be configured on the inside of the open-type earphones. The first microphone may be considered as a microphone that uses feed-forward (FF) active noise reduction technology, which can collect all noise outside the open-type earphones. The position and number of the first microphones are not limited and can be configured according to actual needs. For example, the first microphone can be configured at the edge of the outside of the open-type earphones. The arrangement of multiple first microphones can be configured according to preset positions, for example, they can be respectively configured at positions close to the top corners of the outer surface of the open-type earphones, which can better collect audio data from different directions.
[0065] The second microphone can be considered as a microphone that uses feedback (Feed Back, FB) active noise reduction technology, which is configured on the inner side of the open-type earphones and can obtain the noise inside the earphone shell. For example, the second microphone can be configured at the center position of the inner side of the open-type earphones. Furthermore, the second microphone can be set on the inner side of the speaker close to the user's ear canal hole, or configured at other positions according to other needs.
[0066] This embodiment can acquire original audio data collected by the microphone of the open-earphone, and perform subsequent sound effect processing on the acquired original audio data to obtain high-quality sound effects.
[0067] S102: Perform echo cancellation processing on original audio data to obtain first audio processing data.
[0068] In open headphones, since the microphone and the speaker are fixed in the same cavity and are relatively close, the vibration of the speaker will be transmitted into the microphone, and the sound of the speaker will also be transmitted to the microphone through the air, resulting in the echo formed by the conduction, which will have a negative impact on the subsequent sound effect processing. Therefore, this step can perform echo cancellation processing on the acquired original audio data to obtain the first audio processing data. The means of echo cancellation processing are not limited, such as using time domain or frequency domain methods to perform echo cancellation, so that after echo cancellation, a cleaner sound signal can be obtained, which includes the user's audio data and environmental sounds.
[0069] Furthermore, this step can also perform echo cancellation processing according to the specific amount of original audio data. For example, when there are multiple original audio data, the same or different echo cancellation processing means can be adopted for specific processing respectively, or multiple original audio data can be mixed together for echo cancellation. This embodiment does not limit this.
[0070] S103: Perform audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data, and perform sound effect processing on the first audio processing data to obtain second audio processing data.
[0071] The target reverberation parameters may be used to perform reverberation processing on the first audio processing data, such as including control parameters related to the reverberation processing.
[0072] Specifically, this embodiment can perform audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data. The specific means of audio analysis are not limited. For example, the first audio processing data can be input into the audio analysis model by adopting an audio analysis model to directly output the target reverberation parameters of the first audio processing data. The audio analysis model can be a pre-trained neural network model. The target reverberation parameters of the first audio processing data can also be obtained by performing certain calculations on the first audio processing data. This embodiment does not further expand the specific calculation process, as long as the target reverberation parameters of the first audio processing data can be obtained.
[0073] At the same time, the present embodiment can also perform sound effect processing on the first audio processing data to obtain second audio processing data. The purpose of the sound effect processing can be to cover up audio defects and improve the sound quality of the first audio processing data. The specific means of sound effect processing can be different according to the number of microphones. For example, when the number of microphones is one, the number of original audio data and the first audio processing data is also one. At this time, the sound effect processing can at least include howling suppression, equalization processing, excitation processing and compression processing. When the number of microphones is multiple, the number of original audio data is multiple. At this time, the sound effect processing can be further processed on the basis of the above processing means in combination with beamforming technology to complete the sound effect processing more accurately.
[0074] In some embodiments, the number of microphones is at least two, and the number of raw audio data corresponds to the number of microphones;
[0075] Performing echo cancellation processing on the original audio data to obtain first audio processing data includes:
[0076] Performing echo cancellation processing on each of the at least two original audio data respectively to obtain at least two first audio processing data;
[0077] Accordingly, performing audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data includes:
[0078] Audio analysis is performed on any one of the at least two first audio processing data to obtain a target reverberation parameter of the first audio processing data.
[0079] In a specific implementation, a multi-microphone channel processing strategy can be adopted according to the number of microphones. For example, when the number of microphones is at least two, the number of corresponding original audio data is also multiple. The echo cancellation processing process can be carried out on each original audio data in units of microphone channels to obtain at least two first audio processing data. Accordingly, the specific process of performing audio analysis to obtain the target reverberation parameters can be carried out by performing a comprehensive analysis based on at least two first audio processing data, or by performing audio analysis based on any one of the first audio processing data to obtain the corresponding target reverberation parameters.
[0080] In some embodiments, performing sound effect processing on the first audio processing data to obtain the second audio processing data includes:
[0081] Performing beamforming processing on at least two first audio processing data to obtain beam-processed data, wherein the beamforming processing is used to enhance the audio data in the direction from the user's sound source to the microphone in the at least two first audio processing data, and to weaken the audio data in other directions in the at least two first audio processing data;
[0082] The beam processed data is subjected to howling suppression, equalization processing, excitation processing and compression processing in sequence to obtain second audio processing data.
[0083] In a specific implementation, in order to better recognize the user's voice, beamforming technology can be further used to pick up more directional sounds. For example, in the design of two microphones, the connection between the microphones can point to the direction of the user's sound source. The first audio processing data may include audio data from the user's sound source to the microphone, surrounding environmental noise, or audio data of surrounding people (such as sound that is not emitted from the user's mouth but transmitted to the microphone). Using beamforming technology to process at least two first audio processing data can enhance the user's audio data in the first audio processing data and weaken the surrounding environmental noise and the audio data of surrounding people, so that the voice in the direction of the human mouth can be picked up more clearly, and the environmental noise in other directions can be suppressed to obtain beam processed data; then the beam processed data can be subjected to howling suppression, equalization processing, excitation processing and compression processing in turn to obtain second audio processing data.
[0084] Among them, since the microphone will recognize the ambient sound and play it through the speaker, it is easy to cause howling. This howling is caused by the mid-to-high frequency noise formed by the sound signal looping back to the microphone from the speaker. The mid-to-high frequency noise is sharp and harsh, so the howling suppression processing of the data after beam processing can be performed through feedback elimination and spectrum adjustment technologies to avoid the howling effect and effectively help users get a more stable and clear listening experience in different environments.
[0085] Equalization can be used to adjust the frequency response of audio data. It can be fine-tuned in the low, medium and high frequency bands. For example, by enhancing or cutting the frequency of a specific frequency band, the sound can be made clearer, fuller or softer to meet the needs of different singers and styles. Excitation processing can make the audio sound more dynamic and bright by adding high-frequency harmonics, enhancing the clarity and appeal of the sound.
[0086] Compression processing can include controlling the dynamic range of the audio signal, reducing the volume of parts that are too high or too low, making the sound more stable and consistent, helping to avoid sudden increases or decreases in the sound, making the audio data smoother, and protecting the device from being damaged by peak volume.
[0087] S104: Perform reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data.
[0088] The target audio data can be understood as the audio data after the final sound effect processing.
[0089] After obtaining the target reverberation parameters and the processed second audio processing data through the above steps, specific reverberation processing can be performed on the second audio processing data according to the target reverberation parameters to obtain the target audio data corresponding to the original audio data. The reverberation processing can specifically generate different reverberation effects according to different reverberation parameters. For example, it can simulate the reflection effect of sound in different environments to make the sound have a sense of space and depth. For example, in karaoke, the reverberation processing can cover up some flaws in the singing and make the sound more "live", making the audio data fuller and more natural, and enhancing the singing atmosphere.
[0090] S105: Control the speaker of the open-type earphone to output target audio data.
[0091] The present embodiment provides an audio processing method for open-type headphones, which obtains original audio data collected by a microphone of the open-type headphones; performs echo cancellation processing on the original audio data to obtain first audio processing data; performs audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data, and performs sound effect processing on the first audio processing data to obtain second audio processing data; performs reverberation processing on the second audio processing data according to the target reverberation parameters to obtain target audio data corresponding to the original audio data; and controls the speaker of the open-type headphones to output the target audio data. By using this method, based on the structural characteristics of the open-type headphones, the original audio data can be initially transmitted to the user's ears without any obstacles. At the same time, the open-type headphones perform sound effect processing on the original audio data to obtain the target audio data, so that the output target audio data forms a time difference with the original audio data initially transmitted to the user, thereby presenting a karaoke sound effect using the open-type headphones and improving the user experience.
[0092] Figure 3 This is a flow chart of an audio processing method for an open-earphone provided by another embodiment of the present application. In this embodiment, audio analysis is performed on the first audio processing data to obtain the target reverberation parameters of the first audio processing data, which are further optimized as follows: reverberation estimation is performed on the first audio processing data to obtain the original reverberation time of the first audio processing data, which is used to characterize the time it takes for the first audio processing data to decay to a preset energy value; based on the original reverberation time and the preset reverberation time, the target reverberation parameters of the first audio processing data are calculated, and the target reverberation parameters at least include the target reverberation intensity. Figure 3 As shown, the method includes:
[0093] S201: Acquire original audio data collected by a microphone of an open-ear headset.
[0094] S202: Perform echo cancellation processing on original audio data to obtain first audio processing data.
[0095] S203: Estimating reverberation of the first audio processing data to obtain original reverberation time of the first audio processing data, where the original reverberation time is used to represent the time it takes for the first audio processing data to decay to a preset energy value.
[0096] The original reverberation time is used to represent the time taken for the first audio processing data to decay to a preset energy value, and the preset energy value may be, for example, -60 dB.
[0097] In a specific implementation, since the environment in which users use open-back headphones is different, the reverberation effect of each environment is also very different. For example, there is basically no reverberation when a user sings outdoors, but there will be strong reverberation when singing in the bathroom. Therefore, this embodiment can first estimate the reverberation of the first audio processing data to obtain the original reverberation time of the first audio processing data. There is no limit to the way to determine the original reverberation time. For example, it can be estimated directly through a model, or by determining the audio energy of the first audio processing data and using a preset reverberation algorithm to calculate the audio energy, thereby obtaining the original reverberation time of the first audio processing data. The preset reverberation algorithm can be used to calculate the original reverberation time of a section of audio data.
[0098] S204: Calculate target reverberation parameters of the first audio processing data based on the original reverberation time and the preset reverberation time, where the target reverberation parameters at least include target reverberation intensity, and perform sound effect processing on the first audio processing data to obtain second audio processing data.
[0099] The preset reverberation time may be a pre-set reverberation time, such as a reverberation time under an ideal state, and the specific value may be determined by an empirical value. The target reverberation parameter may include at least a target reverberation intensity, and may also include parameters such as an actual reverberation time.
[0100] S205 . Perform reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data.
[0101] S206: Control the speaker of the open-type earphone to output the target audio data.
[0102] The present embodiment provides an audio processing method for open-type headphones, which obtains the original reverberation time and the preset reverberation time based on reverberation estimation, and calculates the target reverberation parameters of the first audio processing data, thereby providing a feasible means for determining the target reverberation parameters and further ensuring that the open-type headphones can present karaoke sound effects.
[0103] The following is an example of an open-type headset configured with two microphones to describe the audio processing method of the open-type headset:
[0104] Figure 4 is a structural diagram of an audio processing method for an open-type earphone provided in an embodiment of the present application. Figure 4As shown, the original audio speech collected by each microphone is echo-cancelled to obtain two echo-cancelled speech. On the one hand, a speech after echo cancellation can be processed by the reverberation estimation module, and the reverberation time of the environmental system is estimated by the algorithm. The reverberation time can be represented by RT60, that is, the time required for the sound energy to decay to 1 / 1000 (i.e. -60dB) of the initial value; the reverberation strategy module can analyze the reverberation time obtained by the reverberation estimation, and compare it with the reverberation target preset by the headphone system, calculate the parameters for further increasing the reverberation, which may include the reverberation time, the reverberation intensity, etc., and then send the calculated parameters to the reverberation module for processing; on the other hand, the beamforming technology can be used to process the two echo-cancelled speech, which can more clearly pick up the speech in the direction of the human mouth and suppress the noise in other directions, and then can be processed by the howling suppression module, the equalizer, the exciter, the compressor and the reverberation module in sequence to obtain the target audio data, and finally control the speaker to output the target audio data.
[0105] Figure 5 is a time diagram of a user acquiring audio data provided by an embodiment of the present application, such as Figure 5 As shown, the user can first hear the original audio speech emitted by the user's mouth and the conducted sound transmitted to the ear through the air and bones (i.e., the original audio speech), then hear the early reflected sound processed by the equalizer, compressor, etc., and finally hear the late reverberation processed by sound effects (i.e., the target audio data). The whole process simulates the effect of the user singing using a karaoke speaker, allowing the user to enjoy the karaoke experience at home.
[0106] Furthermore, in practical applications, on the basis of echo cancellation, the active noise cancellation (ANC) function of multiple microphone channels can be used to further perform noise reduction processing on at least two original audio data. The specific noise reduction processing process can be as follows:
[0107] First, the original audio data may include first audio data collected by at least two first microphones respectively. This embodiment can obtain second audio data collected by the second microphone of the open-type earphone while obtaining the first audio data, wherein the first microphone and the second microphone are respectively arranged on the outside and inside of the open-type earphone.
[0108] Then, gain processing may be performed on each of the first audio data and the second audio data respectively to obtain first gain-processed data corresponding to each of the first audio data and second gain-processed data corresponding to the second audio data.
[0109] It can be considered that the gain of the audio data originally collected by the first microphone and the second microphone is relatively low, and the sensitivity is low, and the audio data needs to be gain-amplified, so as to adjust the audio data to a suitable amplitude, so as to ensure that no data is lost. Therefore, in this embodiment, each first audio data and the second audio data can be gain-processed respectively to obtain the first gain-processed data corresponding to each first audio data and the second gain-processed data corresponding to the second audio data.
[0110] Secondly, for each first gain processed data, audio analysis can be performed on each first gain processed data based on the second gain processed data to obtain audio information of each first gain processed data.
[0111] The audio information may be various information related to the first gain processing data, and may specifically include delay parameters, or other information that can accurately determine noise reduction parameters. Specifically, for each first gain processing data, audio analysis may be performed on each first gain processing data based on the second gain processing data to obtain the audio information of each first gain processing data. Different audio information may correspond to different audio analysis processes, or the audio analysis processes corresponding to different first gain processing data may also be different.
[0112] Exemplarily, the audio information may include at least a delay parameter and a wind noise parameter. From the physical properties of acoustics, it can be seen that if the noise signal is directional, the first microphone close to the noise signal can receive the noise signal earliest, and the time delay between the audio signal collected by the first microphone close to the noise signal and the audio signal collected by the second microphone is the largest; if the noise signal comes from the front side of the open earphone, the time delay between the audio signal collected by each first microphone and the audio signal collected by the second microphone is the same, so this embodiment can accurately calculate the time delay parameter formed between each first audio data and the second audio data. If the time delay between the first audio signal collected by a first microphone and the audio signal collected by the second microphone is large, it means that the direction of the noise source is closer to the arrangement direction of the first microphone, and the first audio signal collected by the first microphone can obtain a better noise reduction effect, and the weight of the first audio data in the noise reduction process can be increased subsequently.
[0113] Among them, the method of calculating the delay parameter is not limited, such as the cross-correlation method can be used to calculate the cross-correlation function of the two signals, find the peak position of the cross-correlation function, and the time difference corresponding to the peak position is the delay parameter of the two signals; the generalized cross-correlation method can also be used, and a pre-processing weighting function can be added on the basis of the cross-correlation method to enhance the display of the peak position; the short-time Fourier transform method can also be used to analyze the signal delay in the frequency domain, divide the signal into short time windows, and calculate the frequency components and phase differences segment by segment, so as to estimate the obtained time difference as the delay parameter. Furthermore, it is also necessary to analyze the influence of wind noise on the microphone. In general wind noise scenes, due to the great difference in the configuration position and opening direction of the first microphone, each microphone is affected by different wind noise, and the microphone facing the wind is the most affected. The purpose of this embodiment is to reduce or close the microphone channel that is more seriously affected by wind noise, so that the open-type headphones can achieve a good noise reduction effect even under the influence of wind noise. Therefore, in this embodiment, wind noise analysis can be performed on each first gain processed data to obtain wind noise parameters of each first gain processed data. The process of wind noise analysis may include, for example, analyzing whether there is a wind noise spectrum and determining the wind noise intensity.
[0114] Subsequently, the noise reduction parameter corresponding to each first gain processed data may be determined according to the audio information of each first gain processed data.
[0115] After determining the audio information of each first gain processed data, this step can determine the noise reduction parameters corresponding to each first gain processed data separately according to the audio information of each first gain processed data. For example, the noise reduction parameters corresponding to each first gain processed data can be obtained by calculating each parameter in the audio information of each first gain processed data. The importance of each first gain processed data can also be measured by comparing the audio information between multiple first gain processed data to determine the noise reduction parameters corresponding to each first gain processed data. The measurement standard can be determined according to the degree of influence of different audio information on the noise reduction processing. At the same time, multiple noise reduction parameters need to comply with specified regular requirements or meet certain logic.
[0116] In some embodiments, determining the noise reduction parameter corresponding to each first gain processed data according to the audio information of each first gain processed data includes:
[0117] Setting the priority of the wind noise parameter to the first priority, and setting the priority of the delay parameter to the second priority;
[0118] According to the delay parameter and the wind noise parameter of each first gain processing data, the noise reduction parameter corresponding to each first gain processing data is determined respectively by using the first priority and the second priority.
[0119] In a specific implementation, when the audio information includes at least a delay parameter and a wind noise parameter, the sound source direction and the impact of wind noise are analyzed according to the delay parameters, wind noise data and other parameters, the weights of each microphone channel in active noise reduction are dynamically changed, and the optimal weight distribution coefficient is selected. For example, when the delay parameters corresponding to each first microphone are basically the same, the same weight can be assigned to each microphone, that is, the noise reduction parameters corresponding to each first gain processing data can be the same; if it is analyzed that the current noise has obvious directionality, the weight of the microphone that receives the noise earliest can be set to the highest. For example, the gain parameters corresponding to each microphone channel can be changed. The larger the gain parameter, the higher the energy output by the microphone channel, and the greater the weight in active noise reduction, thereby improving the noise reduction effect.
[0120] Furthermore, priorities between different parameters can also be set, such as setting the priority of the wind noise parameter as the first priority, setting the priority of the delay parameter as the second priority, weighing different audio information according to different priorities, and determining the noise reduction parameters corresponding to each first gain processing data respectively. For example, when specifically determining the noise reduction parameters, if the delay parameter of a certain first gain processing data is the smallest, but when it is detected that the first gain processing data has a wind noise spectrum, the wind noise parameter is mainly used, and the noise reduction parameter corresponding to the first gain processing data is reduced according to the specific size of the wind noise parameter; more specifically, when it is detected that the first gain processing data of a certain first microphone has a wind noise spectrum, if the wind noise intensity of the first gain processing data is detected to be greater than 65dB, the weight of the first audio data in the noise reduction processing can be reduced, such as the gain of the microphone can be attenuated by 12dB, and if the wind noise intensity of the first gain processing data is detected to be greater than 75dB, the weight of the first audio data in the noise reduction processing can be directly set to zero, that is, the noise reduction processing of this microphone channel is turned off, and accordingly, the noise reduction parameters of other microphone channels can be adaptively adjusted, such as increasing their corresponding weight coefficients.
[0121] Finally, according to the noise reduction parameters corresponding to each first audio data, each first audio data is subjected to noise reduction processing to obtain the first noise reduction data corresponding to each first audio data, so that the obtained first noise reduction data can be subjected to subsequent sound effect processing to achieve better sound quality effects.
[0122] In some embodiments, the noise reduction parameter is a gain parameter for gain processing, and according to the noise reduction parameter corresponding to each first audio data, noise reduction processing is performed on each first audio data to obtain first noise reduction data corresponding to each first audio data, including:
[0123] Performing filtering processing on each first gain processed data respectively to obtain first filtered data corresponding to each first gain processed data;
[0124] According to the gain parameters corresponding to each first gain processed data, gain processing is performed on the first filtered data corresponding to each first gain processed data to obtain the first noise reduction data corresponding to each first audio data.
[0125] The noise reduction parameter may be a gain parameter used for gain processing, such as an increased or decreased gain value.
[0126] In a specific implementation, each first microphone corresponds to its own filter and gain module, and can process the collected audio data respectively, and filter each first gain-processed data in turn, and perform gain processing on the first filtered data corresponding to each first gain-processed data according to the gain parameters corresponding to each first gain-processed data, so as to obtain the first noise reduction data corresponding to each first audio data, wherein, for the first filtered data corresponding to each first gain-processed data, the process of performing gain processing specifically according to the gain parameters can include determining a target gain value based on an initial gain value, so as to adjust the amplitude of the corresponding audio data according to the target gain value, such as first calculating the target gain value corresponding to each first filtered data according to the gain parameters and the initial gain value corresponding to each first gain-processed data, and then using the target gain value to perform gain processing on each first filtered data, so as to obtain the first noise reduction data corresponding to each first audio data.
[0127] The audio processing method of the open-earphone corresponding to the above embodiment, Figure 6 This is a schematic diagram of the structure of an open-type earphone provided in one embodiment of the present application. For the sake of ease of explanation, only the parts related to the embodiment of the present application are shown.
[0128] Reference Figure 6 , the open-back headphones include:
[0129] A microphone 301 is provided in an open-type headset and is used to collect raw audio data;
[0130] The echo cancellation module 302 is used to perform echo cancellation processing on the original audio data to obtain first audio processing data;
[0131] The processing module 303 is used to perform audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data;
[0132] The audio effect processing module 304 is used to perform audio effect processing on the first audio processing data to obtain second audio processing data;
[0133] A reverberation module 305 is used to perform reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data;
[0134] The output module 306 is used to control the speaker of the open-ear headset to output the target audio data.
[0135] An open-type headset provided in an embodiment of the present application collects original audio data through a microphone; performs echo cancellation processing on the original audio data through an echo cancellation module to obtain first audio processing data; performs audio analysis on the first audio processing data through a processing module to obtain target reverberation parameters of the first audio processing data; performs sound effect processing on the first audio processing data through a sound effect processing module to obtain second audio processing data; performs reverberation processing on the second audio processing data according to the target reverberation parameters through a reverberation module to obtain target audio data corresponding to the original audio data; and controls the speaker of the open-type headset to output the target audio data through an output module. By using the open-type headset, according to the structural characteristics of the open-type headset, the original audio data can be initially transmitted to the user's ears without obstacles. At the same time, the open-type headset performs sound effect processing on the original audio data to obtain the target audio data, so that the output target audio data forms a time difference with the original audio data initially transmitted to the user, thereby using the open-type headset to present the sound effect of karaoke, improving the user experience.
[0136] In some embodiments, the audio processing module includes:
[0137] A beamforming unit, configured to perform beamforming processing on at least two first audio processing data to obtain beam-processed data;
[0138] A howling suppression unit, used for performing howling suppression on the beam processed data;
[0139] An equalizer, used to adjust the frequency response of the audio data after howling suppression;
[0140] an exciter for adding high frequency harmonics to the frequency adjusted audio data;
[0141] A compressor is used to control the volume range of the audio data after being processed by the exciter.
[0142] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0143] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0144] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0145] An embodiment of the present application provides a computer program product. When the computer program product is run on an open-ear headset, the open-ear headset can implement the steps in the above-mentioned various method embodiments.
[0146] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / open earphone, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0147] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0148] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0149] In the embodiments provided in the present application, it should be understood that the disclosed device / open-earphone and method can be implemented in other ways. For example, the device / open-earphone embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0150] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An audio processing method for open-type headphones, characterized in that: include: Get the raw audio data collected by the microphone of the open headphone; Performing echo cancellation processing on the original audio data to obtain first audio processing data; Performing audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data, and performing sound effect processing on the first audio processing data to obtain second audio processing data; Performing reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data; A speaker of the open headphone is controlled to output the target audio data.
2. The audio processing method of open earphones according to claim 1, characterized in that: The performing audio analysis on the first audio processing data to obtain a target reverberation parameter of the first audio processing data includes: Performing reverberation estimation on the first audio processing data to obtain an original reverberation time of the first audio processing data, where the original reverberation time is used to represent the time it takes for the first audio processing data to decay to a preset energy value; Based on the original reverberation time and the preset reverberation time, target reverberation parameters of the first audio processing data are calculated, and the target reverberation parameters at least include a target reverberation intensity.
3. The audio processing method of open-earphones according to claim 2, characterized in that: The performing reverberation estimation on the first audio processing data to obtain the original reverberation time of the first audio processing data includes: determining audio energy of the first audio processing data; The audio energy is calculated using a preset reverberation algorithm to obtain an original reverberation time of the first audio processing data.
4. The audio processing method for open earphones according to claim 1, characterized in that: The number of the microphone is one, and accordingly, the number of the original audio data and the number of the first audio processed data are both one; The performing sound effect processing on the first audio processing data to obtain second audio processing data includes: The first audio processing data is subjected to audio effect processing to obtain second audio processing data, wherein the audio effect processing at least includes howling suppression, equalization processing, excitation processing and compression processing.
5. The audio processing method of open earphones according to claim 1, characterized in that: The number of the microphones is at least two, and the number of the original audio data corresponds to the number of the microphones; performing echo cancellation processing on the original audio data to obtain first audio processing data includes: Performing echo cancellation processing on each of the at least two original audio data respectively to obtain at least two first audio processing data; Accordingly, performing audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data includes: Audio analysis is performed on any one of the at least two first audio processing data to obtain a target reverberation parameter of the first audio processing data.
6. The audio processing method of open-earphones according to claim 5, characterized in that: The performing sound effect processing on the first audio processing data to obtain second audio processing data includes: Performing beamforming processing on the at least two first audio processing data to obtain beam-processed data, wherein the beamforming processing is used to enhance the audio data in the direction from the user's sound source to the microphone in the at least two first audio processing data, and to weaken the audio data in other directions in the at least two first audio processing data; The beam processed data is sequentially subjected to howling suppression, equalization processing, excitation processing and compression processing to obtain second audio processing data.
7. An open-type headphone, characterized in that: include: A microphone, configured in an open-type headset, for collecting raw audio data; An echo cancellation module, configured to perform echo cancellation processing on the original audio data to obtain first audio processing data; a processing module, configured to perform audio analysis on the first audio processing data to obtain target reverberation parameters of the first audio processing data; An audio effect processing module, used for performing audio effect processing on the first audio processing data to obtain second audio processing data; a reverberation module, configured to perform reverberation processing on the second audio processing data according to the target reverberation parameter to obtain target audio data corresponding to the original audio data; The output module is used to control the speaker of the open-type earphone to output the target audio data.
8. The earphone according to claim 7, characterized in that The sound effect processing module comprises: A beamforming unit, configured to perform beamforming processing on at least two first audio processing data to obtain beam-processed data; A howling suppression unit, used for performing howling suppression on the beam processed data; An equalizer, used to adjust the frequency response of the audio data after howling suppression; an exciter for adding high frequency harmonics to the frequency adjusted audio data; A compressor is used to control the volume range of the audio data after being processed by the exciter.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, enables the method according to any one of claims 1 to 6 to be performed.