Earphone voice enhancement method, earphone and storage medium
By obtaining the qualified attitude angle and quantity of the headphones, determining the microphone array model and weight parameters, and collecting and fusing voice characteristics, the audio signal quality problem caused by changes in the headphone wear situation is solved, and a higher quality voice call is achieved.
Patent Information
- Application Number
- CN202411988187.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The existing headphones are different from the wearable habits of users and their wearable conditions, which leads to changes in the position and quantity of microphone arrays, affecting the quality of audio signals and containing noise signals, reducing call quality.
By obtaining the qualified attitude angle and the qualified number of the headphones, the array model of the microphone array is determined, and the target weight parameters are obtained. Based on these parameters, the audio signal is collected, the STFT spectrum and the Mel spectrum are extracted, and the speech characteristics are fused to obtain the enhanced speech signal.
It realizes adaptive matching of the weight parameters of the microphone array according to the headset wear and use scenarios, improves the accuracy of directional acquisition, improves the quality of voice signals, and improves call quality.
Smart Images

Figure CN119996889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of headset speech enhancement, and in particular to a headset speech enhancement method, a headset and a storage medium. Background Art
[0002] With the rapid development of electronic technology, people are using earphones more and more widely. Earphones are usually provided with a microphone array consisting of multiple microphones, which collects the user's voice signal through beamforming to ensure the call quality when the user wears the earphone for a call.
[0003] In existing headsets, fixed array pointing is usually used to collect the voice signal of the user during a call. However, in the actual use of the headset by the user, due to the different wearing habits and wearing conditions of the user, the position and number of the microphones for collecting the user's voice will change, which will affect the quality of the audio signal collected by the headset. At the same time, the audio signal still contains noise signals, which will affect the call quality of the user using the headset.
[0004] In view of this, it is necessary to provide a method for earphone speech enhancement, earphone and storage medium to solve the above problems. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a method for headset voice enhancement, a headset and a storage medium, which are intended to solve the problem of relatively poor call quality when a user uses a headset to make a voice call.
[0006] To achieve the above-mentioned purpose, a first aspect of the present invention provides a method for headphone speech enhancement, the steps of which include:
[0007] In response to triggering the voice call mode, obtaining a qualified posture angle and a qualified number of the earphone;
[0008] Determine an array model of the microphone array according to the qualified attitude angle and the qualified number;
[0009] Obtain the weight parameter corresponding to the array model as the target weight parameter;
[0010] Based on the target weight parameters, the microphone array on the headset collects the audio signal;
[0011] According to the audio signal, the STFT spectrum and the Mel spectrum are obtained respectively;
[0012] From the STFT spectrum and the Mel spectrum, the first speech feature and the second speech feature are obtained respectively;
[0013] The first speech feature and the second speech feature are fused to obtain an enhanced speech signal.
[0014] In one embodiment, in response to triggering the voice call mode, the step of obtaining the qualified attitude angle and qualified number of the headset includes:
[0015] Obtaining the wearing posture angle of the headset in the wearing state;
[0016] Determine whether the wearing posture angle falls within the preset qualified posture angle range;
[0017] If yes, the wearing posture angle is marked as a qualified posture angle;
[0018] The number of qualified attitude angles is determined to obtain a qualified number.
[0019] In one embodiment, the step of determining the array model of the microphone array according to the qualified attitude angle and the qualified number includes:
[0020] According to the qualified number, determine the acquisition mode corresponding to the trigger;
[0021] Acquire a model database corresponding to an acquisition mode and a number of corresponding determination intervals;
[0022] Based on the qualified attitude angle, a number of determination intervals and a model database, an array model of a corresponding microphone array is determined.
[0023] In one embodiment, the acquisition mode includes a first acquisition mode and a second acquisition mode; and according to the qualified number, the step of determining the corresponding triggered acquisition mode includes:
[0024] If the qualified number is 2, the first acquisition mode is triggered;
[0025] If the qualified number is 1, the second acquisition mode is triggered.
[0026] In one embodiment, the step of determining the array model of the corresponding microphone array based on the qualified attitude angle, the plurality of determination intervals and the model database includes:
[0027] Get the judgment interval into which the qualified posture angle falls;
[0028] Get the eigenvalue corresponding to the judgment interval;
[0029] According to the characteristic value, a corresponding preset array model is determined from the first database, and the preset array model is used as the array model of the microphone array.
[0030] In one embodiment, the step of obtaining the first speech feature and the second speech feature from the STFT spectrum and the Mel spectrum includes:
[0031] Obtaining a preset first extraction mode and a second extraction mode;
[0032] Obtaining a first speech feature according to the STFT spectrum and the first extraction mode;
[0033] A second speech feature is obtained according to the Mel spectrum and the second extraction mode.
[0034] In one embodiment, the step of fusing the first speech feature with the second speech feature to obtain an enhanced speech signal includes:
[0035] Based on the second extraction mode, the second speech feature is integrated into the first speech feature to obtain a specific speech feature;
[0036] The specific speech features are inversely transformed to obtain an enhanced speech signal.
[0037] In one embodiment, after the step of obtaining the weight parameter corresponding to the array model as the target weight parameter, the following step is further included:
[0038] Obtain the wearing status and wearing posture angle of the headset in real time;
[0039] Based on the wearing state and the wearing posture angle, detecting whether to obtain a new target weight parameter of the microphone array;
[0040] If yes, obtain the qualified attitude angle and qualified number of the headset.
[0041] A second aspect of the present invention provides a headset, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of any one of the above-mentioned methods for headset speech enhancement are implemented.
[0042] A third aspect of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned methods for headphone speech enhancement are implemented.
[0043] The beneficial effects of the present invention are as follows: by utilizing the qualified posture angle and qualified number of the headset, the microphone array model of the headset under voice call is determined, and then the target weight parameters that meet the current usage scenario are quickly determined, and then the first voice feature and the second voice feature are extracted and fused from the audio signal collected based on the target weight parameters to obtain an enhanced voice signal. This can realize adaptive matching of the corresponding microphone array weight parameters according to the specific circumstances of the headset wearing and usage scenario, and can effectively solve the problem of relatively low matching between the microphone array enhancement direction and the actual usage situation, and improve the accuracy of directional acquisition. At the same time, based on the characteristics of the human voice and the auditory characteristics of the human ear, the first voice feature and the second voice feature are fused, which can realize relatively more accurate extraction of the user's voice signal, improve the quality of the obtained voice signal, and thus improve the call quality of the user using the headset for calls. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 The first flowchart of the method for headphone speech enhancement disclosed in an embodiment of the present invention is shown in FIG.
[0045] Figure 2 This is a second flow chart of the method for headphone speech enhancement disclosed in an embodiment of the present invention.
[0046] Figure 3 This is a schematic diagram of the module structure of the earphone disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In the present invention, the terms "disposed", "provided with" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection, or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be an internal connection between two devices, elements or components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0048] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0049] In addition, some of the above terms may be used to express other meanings in addition to indicating orientation or positional relationship. For example, the term "on" may also be used to express a certain dependency or connection relationship in some cases. For those skilled in the art, the specific meanings of these terms in the present invention can be understood according to specific circumstances.
[0050] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0051] The following is the content of the first aspect of the present invention:
[0052] Please refer to Figure 1 In this embodiment, the method for headphone speech enhancement includes the following steps:
[0053] S1. In response to triggering a voice call mode, obtaining a qualified posture angle and a qualified number of the earphone.
[0054] The qualified posture angle is the wearing posture angle when the headset is in a relatively effective position to collect voice signals when a voice call is triggered. The qualified number is the number of qualified posture angles when a voice call is triggered, that is, when only one headset's wearing posture angle meets the requirements, the qualified number is 1, and when both headsets' wearing posture angles meet the requirements, the qualified number is 2.
[0055] Specifically, after triggering the voice call mode, by obtaining the qualified posture angle and qualified number of the headset, the wearing condition and usage condition of the headset can be preliminarily determined, thereby preparing for the subsequent division of the user's use of the headset for calls based on the qualified posture angle and qualified number of the headset.
[0056] S2. Determine an array model of the microphone array according to the qualified attitude angle and the qualified number;
[0057] S3, obtaining a weight parameter corresponding to the array model as a target weight parameter;
[0058] S4, based on the target weight parameter, the microphone array on the headset collects the audio signal;
[0059] Among them, the array model is a model obtained by a large number of tests in advance, which can be used to reflect the distribution characteristics of the microphones in the earphone. Several groups of array models are pre-set inside the earphone, and a mapping relationship between qualified posture angles and qualified numbers is set for several groups of array models. That is, by determining the qualified posture angles and the qualified number, the array model corresponding to the current microphone array can be quickly determined.
[0060] The weight parameter is the importance ratio of each microphone in the microphone array on the headset when processing sound signals. The weight parameter is the data obtained by a large number of pre-tests based on the array model. Each array model has a corresponding weight parameter.
[0061] When using headphones, different users wear headphones in different postures, and users use headphones in different scenarios. There are many scenarios for wearing and using headphones. During a voice call, different wearing postures and usage of headphones will affect the quality of the collected audio signal. For example, the user may wear two headphones and both headphones meet the wearing requirements, or the user may wear two headphones but only one headphone meets the wearing requirements, or the user may only wear a single headphone and meet the wearing requirements.
[0062] Specifically, after the voice call mode is triggered and the qualified posture angle and qualified number of the headset are obtained, a preliminary determination can be made on the usage of the headset according to the qualified number to obtain the corresponding array models, and then combined with the specific situation of the qualified posture angle, according to the preset mapping relationship, the array model of the corresponding microphone array can be quickly determined from the array models. In the actual use scenario of the headset, the qualified number may be 1 or 2. For different qualified numbers, the configuration of the array model library is different. The determination method and mapping relationship of the qualified posture angle can be the same or different, which can be selected according to the actual design requirements. In the use scenario of the headset, the qualified number may also be 0, that is, before determining the microphone array model, if the qualified number is 0, that is, the wearing posture of the headset is relatively unreasonable. In order to ensure the quality of the collected audio signal, the headset can output the corresponding prompt information to remind the user to adjust the posture of the headset and re-obtain the qualified posture angle and qualified number of the headset.
[0063] After determining the array model of the microphone array according to the qualified posture angle and qualified number of the headset, the weight parameters corresponding to the array model of the microphone array are obtained, and the weight parameters are used as target weight parameters for the current usage scenario of the headset.
[0064] It is understandable that since the distance between the microphones in the headphones is fixed, and after the user wears the headphones, the distance between the headphones and the user's mouth is also relatively fixed, several microphone array models can be set in advance, and then the usage scenarios of the headphones can be divided according to the qualified posture angles and qualified numbers of the headphones, thereby realizing the adaptive acquisition of the corresponding microphone array model and the corresponding target weight parameters according to the wearing condition and usage condition of the headphones, which can effectively solve the problem of low matching between the microphone array enhancement direction and the actual usage condition, and improve the quality of directional audio signal collection by the headphone microphone array.
[0065] S5, obtaining the STFT spectrum and the Mel spectrum respectively according to the audio signal;
[0066] S6. Obtaining a first speech feature and a second speech feature from the STFT spectrum and the Mel spectrum respectively;
[0067] S7. Fuse the first speech feature and the second speech feature to obtain an enhanced speech signal.
[0068] The STFT spectrum is the audio spectrum obtained by short-time Fourier transform of the audio signal, and the Mel spectrum is the spectrum of the STFT spectrum converted to Mel scale. The first speech feature is the feature information about the user's speech audio obtained from the STFT spectrum, and the second speech feature is the feature information about the user's speech audio obtained from the Mel spectrum.
[0069] Specifically, after the microphone array collects the audio signal, the audio signal is preprocessed, and then the preprocessed audio signal is short-time Fourier transformed to obtain an STFT spectrum, and then the STFT spectrum is converted based on the Mel scale to obtain a Mel spectrum.
[0070] After obtaining the STFT spectrum and the Mel spectrum, based on the preset extraction mode, the first speech feature and the second speech feature are obtained from the STFT spectrum and the Mel spectrum respectively. Specifically, a number of first speech features are extracted from the full frequency band of the STFT, and a number of second speech features are extracted from the specific frequency band of the Mel spectrum. Specifically, the first speech feature and the second speech feature include but are not limited to one or more or a combination of features such as amplitude features and energy intensity features, which can be selected according to design requirements and are not limited here.
[0071] After obtaining the first voice feature and the second voice feature, the first voice feature and the second voice feature need to be fused to obtain a voice feature with relatively more comprehensive information. Specifically, the first voice feature and the second voice feature can be weighted fused, or the feature information in the first voice feature can be replaced based on the second voice feature, which can be selected according to the actual design requirements. After fusing the first voice feature and the second voice feature, the obtained voice feature is inversely transformed to obtain an enhanced voice signal.
[0072] It can be understood that by utilizing the qualified posture angle and qualified number of the headset, the microphone array model of the headset under voice call is determined, and then the target weight parameters that meet the current usage scenario are quickly determined, and then the first voice feature and the second voice feature are extracted and fused from the audio signal collected based on the target weight parameters to obtain an enhanced voice signal. This can realize adaptive matching of the corresponding microphone array weight parameters according to the specific circumstances of the headset wearing and usage scenario, and can effectively solve the problem of relatively low matching between the microphone array enhancement direction and the actual usage situation, and improve the accuracy of directional acquisition. At the same time, based on the characteristics of the human voice and the auditory characteristics of the human ear, the first voice feature and the second voice feature are fused, so that the user's voice signal can be extracted relatively more accurately, and the quality of the obtained voice signal can be improved, thereby improving the call quality of the user using the headset.
[0073] Further, in one embodiment, in response to triggering the voice call mode, step S1 of obtaining the qualified attitude angle and qualified number of the earphone includes:
[0074] S11, obtaining a wearing posture angle of the headset in a wearing state;
[0075] S12, determining whether the wearing posture angle falls within a preset qualified posture angle interval;
[0076] S13: If yes, marking the wearing posture angle as a qualified posture angle;
[0077] S14, determining the number of qualified attitude angles to obtain a qualified number.
[0078] Among them, the qualified posture angle interval is the angle range of the wearing posture angle that can effectively collect voice signals after the headset is worn.
[0079] After the voice call mode is triggered, the wearing state and wearing attitude angle of the headset are determined by the wearing detection sensor and attitude sensor inside the headset, respectively, and then the wearing attitude angle of the headset in the wearing state can be obtained. Specifically, the wearing detection sensor and the attitude detection sensor can be operated at the same time, and then the wearing attitude angle of the headset in the wearing state is determined based on the obtained information; or the wearing detection sensor can be used to determine the headset in the wearing state first, and then the attitude sensor of the headset in the wearing state is used for detection to determine the wearing attitude angle of the headset in the wearing state. The selection can be made according to the actual design requirements.
[0080] After obtaining the wearing posture angle, by determining the size relationship between the wearing posture angle and the two endpoints of the qualified posture angle interval, it can be determined whether the wearing posture angle falls within the qualified posture angle interval. When the wearing posture angle does not fall within the qualified posture angle interval, it means that the wearing posture of the headset has a more significant impact on the user's voice collection; when the wearing posture angle falls within the qualified posture angle interval, the wearing posture of the headset can relatively effectively collect the user's voice signal, and the wearing posture angle can be marked as a qualified posture angle. After determining the qualified posture angle, record the number of qualified posture angles to get the required qualified number.
[0081] It can be understood that by comparing the wearing posture angle with the qualified posture angle range, it is possible to quickly determine whether the current wearing of the headset meets the requirements for voice signal collection, so as to preliminarily determine the wearing condition and usage of the headset, and then prepare for the subsequent classification of users using the headset for calls based on the qualified posture angle and qualified number of the headset.
[0082] For further information, please refer to Figure 2 In one embodiment, the step S2 of determining the array model of the microphone array according to the qualified attitude angle and the qualified number includes:
[0083] S21, determining the corresponding triggered acquisition mode according to the qualified number;
[0084] S22, obtaining a model database corresponding to the aforementioned acquisition mode and a number of corresponding determination intervals;
[0085] S23. Determine the array model of the corresponding microphone array based on the qualified attitude angle, a plurality of determination intervals and a model database.
[0086] The acquisition mode is a method for determining the array model of the microphone array. The model database is a data collection of several array models related to the acquisition mode. The determination interval is a classification interval of qualified attitude angles.
[0087] Specifically, the acquisition mode includes a first acquisition mode and a second acquisition mode, the model database includes a first database and a second database, and the determination interval includes a first determination interval and a second determination interval. The first database and the first determination interval are respectively associated with the first acquisition mode, and the second database and the second determination interval are respectively associated with the second acquisition mode.
[0088] Specifically, if the qualified number is 2, the first acquisition mode is triggered, and the corresponding first database and first determination interval are called; if the qualified number is 1, the second acquisition mode is triggered, and the corresponding second database and second determination interval are called.
[0089] After determining the triggered acquisition mode according to the qualified number, call the model database corresponding to the acquisition mode and the corresponding several determination intervals; then determine the relationship between the qualified posture angle and the several determination intervals. Specifically, determine the size relationship between the qualified posture angle and the endpoints of each determination interval in turn, and then determine the determination interval into which the qualified posture angle falls. After obtaining the determination interval into which the qualified posture angle falls, the corresponding array model can be directly determined according to the determination interval, or the corresponding characteristic value can be obtained based on the determination interval, and then the corresponding array model can be determined according to the mapping relationship.
[0090] In one embodiment, the step S23 of determining the array model of the corresponding microphone array based on the qualified attitude angle, the plurality of determination intervals and the model database includes:
[0091] S231, obtaining a determination interval into which a qualified posture angle falls;
[0092] S232, obtaining a characteristic value corresponding to the determination interval;
[0093] S233. Determine a corresponding preset array model from a model database according to the characteristic value, and use the preset array model as the array model of the microphone array.
[0094] Specifically, when the qualified number is 2, the first acquisition mode is triggered, and the first database and a plurality of first determination intervals corresponding to the first acquisition mode are called. The falling relationship between the two qualified posture angles and the first determination interval is determined respectively, and then the first determination interval in which the two qualified posture angles fall is obtained. Then, the characteristic values corresponding to the two first determination intervals are obtained respectively, and then the corresponding array model is obtained from the first database according to the two characteristic values and the preset first mapping relationship.
[0095] When the qualified number is 1, it means that the user may only wear one earphone or only one of the two earphones meets the wearing requirements, then the second acquisition mode is triggered, and the second database and several second determination intervals corresponding to the second acquisition mode are called. The relationship between the qualified posture angle and the second determination interval is determined, and then the second determination interval in which each posture angle falls is obtained, and then the characteristic value corresponding to the second determination interval is obtained, and then the corresponding array model is obtained from the second database according to the characteristic value and the preset second mapping relationship.
[0096] It can be understood that by using the judgment interval to determine the array model, the efficiency of obtaining the array model of the microphone array can be improved, thereby improving the efficiency of subsequently obtaining the weight parameters of the microphone array, and reducing the response time for voice collection to a certain extent. In the actual use of headphones by users, there are a variety of wearing postures. By using the judgment interval method, the various wearing postures can be summarized and simplified, which can reduce the system memory usage and the amount of calculation to a certain extent, and obtain the required array model of the microphone array in a relatively lower load manner.
[0097] Further, in one embodiment, step S6 of obtaining the first speech feature and the second speech feature from the STFT spectrum and the Mel spectrum includes:
[0098] S61, obtaining a preset first extraction mode and a second extraction mode;
[0099] S62, obtaining a first speech feature according to the STFT spectrum and the first extraction mode;
[0100] S63. Obtain a second speech feature according to the Mel spectrum and the second extraction mode.
[0101] Among them, the first extraction mode is a method preset for extracting the first speech feature for the STFT spectrum, and the first extraction mode includes a first extraction range and a first extraction bandwidth. The second extraction mode is a method preset for extracting the second speech feature for the Mel spectrum, and the second extraction mode includes a second extraction range and a second extraction bandwidth. The first extraction range is the range of the full frequency band, and the second extraction range is a specific frequency band where the human ear is relatively sensitive.
[0102] Specifically, after obtaining the STFT spectrum and the Mel spectrum, the first extraction mode and the second extraction mode are obtained. According to the first extraction range and the first extraction bandwidth corresponding to the first extraction mode, a number of first speech features are extracted from the STFT spectrum; according to the second extraction range and the second extraction bandwidth corresponding to the second extraction mode, a number of second speech features are extracted from the Mel spectrum.
[0103] Furthermore, after obtaining the first speech feature and the second speech feature, step S7 of fusing the first speech feature and the second speech feature to obtain an enhanced speech signal includes:
[0104] S71, based on the second extraction mode, integrating the second speech feature into the first speech feature to obtain a specific speech feature;
[0105] S72. Inversely transform the specific speech feature to obtain an enhanced speech signal.
[0106] Specifically, after obtaining the first voice feature and the second voice feature, according to the second extraction range corresponding to the second extraction mode, the second voice feature is integrated into the first voice feature with a preset weight parameter, thereby obtaining a specific voice feature, that is, the voice feature extracted from the specific frequency band of the Mel spectrum is integrated into the corresponding frequency band of the voice feature extracted from the STFT spectrum, thereby obtaining a specific voice feature containing more comprehensive voice information. Then, the specific voice feature is inversely transformed to obtain an enhanced voice signal.
[0107] It can be understood that by extracting the first voice feature over the full frequency band of the STFT spectrum, extracting the second voice feature over the specific frequency band of the Mel spectrum, and obtaining the specific voice feature using the first voice feature and the second voice feature, the refinement of the first voice feature and the human ear perception characterization of the second voice feature can be combined, the user's voice signal can be extracted more accurately, the quality of the obtained voice signal can be improved, and the call quality of the user using headphones can be achieved.
[0108] Furthermore, in a preferred embodiment, after obtaining the weight parameter corresponding to the array model as the target weight parameter step S3, the following steps are further included:
[0109] S10, obtaining the wearing state and wearing posture angle of the headset in real time;
[0110] S20, based on the wearing state and the wearing posture angle, detecting whether to obtain a new target weight parameter of the microphone array;
[0111] S30: If yes, obtain the qualified attitude angle and qualified number of the headset.
[0112] When a user uses headphones to make a voice call, the user may adjust the headphones, take off a single headphone, or put on the headphones again. To ensure the quality of the user's voice call, it is necessary to detect in real time whether a new target weight parameter of the microphone array needs to be obtained during the user's voice call.
[0113] Specifically, during a user's voice call, the wearing detection sensor and the posture detection sensor provided inside the headset respectively detect the wearing state of the headset and the wearing posture angle of the headset in real time. After obtaining the wearing state and the wearing posture angle of the headset, it is possible to first determine whether the number of headsets worn has changed based on the wearing state of the headset to determine whether it is in the process of triggering the acquisition of the target weight parameters of the new microphone array. When it is determined that the number of microphones worn has changed, the acquisition of the target weight parameters of the new microphone array is triggered. After determining that the number of microphones worn has not changed, it is then determined whether the change in the wearing posture angle of the headset exceeds the preset permitted range. If so, the acquisition of the target weight parameters of the new microphone array is triggered; if not, the acquisition of the target weight parameters of the new microphone array is not triggered.
[0114] It can be understood that by real-time detection of whether to trigger the acquisition of the target weight parameters of the new microphone array, it is possible to better cope with the changing usage scenarios during user voice calls, to ensure the quality of the user's voice signal collected in a directionally manner, and thus to improve the user's experience of using headphones for voice calls.
[0115] In summary, the present invention determines the microphone array model of the headset under voice call by utilizing the qualified posture angle and qualified number of the headset, and then quickly determines the target weight parameters that meet the current usage scenario, and then extracts and fuses the first voice feature and the second voice feature from the audio signal collected based on the target weight parameters to obtain an enhanced voice signal. It can realize adaptive matching of the corresponding microphone array weight parameters according to the specific circumstances of the headset wearing and usage scenario, and can effectively solve the problem of relatively low matching between the microphone array enhancement direction and the actual usage situation, and improve the accuracy of directional acquisition. At the same time, according to the characteristics of the human voice and the auditory characteristics of the human ear, the first voice feature and the second voice feature are fused, which can realize relatively more accurate extraction of the user's voice signal, improve the quality of the obtained voice signal, and thus improve the call quality of the user using the headset for calls.
[0116] The following is the content of the second aspect of the present invention:
[0117] The present invention provides an earphone, such as Figure 3As shown, the headset includes a memory 10, a processor 20, and a method program instruction 30 for headset speech enhancement stored in the memory 10 and executable on the processor 20. When the method program instruction 30 for headset speech enhancement is executed by the processor 20, the aforementioned method for headset speech enhancement is implemented.
[0118] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the headset. In this embodiment, the processor is used to run program codes stored in a readable storage medium or process data.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a readable storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0120] The following is the content of the third aspect of the present invention:
[0121] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned headphone speech enhancement method are implemented.
[0122] The above is only a specific implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for headphone speech enhancement, characterized in that: include: In response to triggering the voice call mode, obtaining a qualified posture angle and a qualified number of the earphone; Determining an array model of the microphone array according to the qualified attitude angle and the qualified number; Obtaining a weight parameter corresponding to the array model as a target weight parameter; Based on the target weight parameter, the microphone array on the headset collects an audio signal; According to the audio signal, a STFT spectrum and a Mel spectrum are obtained respectively; Obtaining a first speech feature and a second speech feature from the STFT spectrum and the Mel spectrum respectively; The first speech feature and the second speech feature are fused to obtain an enhanced speech signal.
2. The method for headphone speech enhancement according to claim 1, characterized in that: In response to triggering the voice call mode, the step of obtaining the qualified attitude angle and qualified number of the earphone comprises: Obtaining the wearing posture angle of the headset in the wearing state; Determining whether the wearing posture angle falls within a preset qualified posture angle interval; If yes, marking the wearing posture angle as a qualified posture angle; The number of qualified attitude angles is determined to obtain a qualified number.
3. The method for headphone speech enhancement according to claim 1, characterized in that: The step of determining the array model of the microphone array according to the qualified attitude angle and the qualified number comprises: Determine a corresponding triggered acquisition mode according to the qualified number; Acquire a model database corresponding to the acquisition mode and a number of corresponding determination intervals; Based on the qualified attitude angle, the plurality of determination intervals and the model database, an array model of a corresponding microphone array is determined.
4. The method for headphone speech enhancement according to claim 3, characterized in that: The acquisition mode includes a first acquisition mode and a second acquisition mode; the step of determining the acquisition mode corresponding to the trigger according to the qualified quantity includes: If the qualified number is 2, the first acquisition mode is triggered; If the qualified number is 1, the second acquisition mode is triggered.
5. The method for headphone speech enhancement according to claim 3, characterized in that: The step of determining the array model of the corresponding microphone array based on the qualified attitude angle, the plurality of determination intervals and the model database comprises: Obtaining a determination interval into which the qualified posture angle falls; Obtaining a characteristic value corresponding to the determination interval; According to the characteristic value, a corresponding preset array model is determined from the first database, and the preset array model is used as the array model of the microphone array.
6. The method for headphone speech enhancement according to claim 1, characterized in that: The step of obtaining the first speech feature and the second speech feature from the STFT spectrum and the Mel spectrum comprises: Obtaining a preset first extraction mode and a second extraction mode; Obtaining a first speech feature according to the STFT spectrum and the first extraction mode; A second speech feature is obtained according to the Mel spectrum and the second extraction mode.
7. The method for headphone speech enhancement according to claim 6, characterized in that: The step of fusing the first speech feature and the second speech feature to obtain an enhanced speech signal comprises: Based on the second extraction mode, the second speech feature is integrated into the first speech feature to obtain a specific speech feature; The specific speech feature is inversely transformed to obtain an enhanced speech signal.
8. The method for headphone speech enhancement according to claim 1, characterized in that: After the step of obtaining the weight parameter corresponding to the array model as the target weight parameter, the following step further comprises: Obtain the wearing status and wearing posture angle of the headset in real time; Based on the wearing state and the wearing posture angle, detecting whether to obtain a new target weight parameter of the microphone array; If yes, obtain the qualified attitude angle and qualified number of the headset.
9. A headset comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the headphone speech enhancement method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the headphone speech enhancement method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Intelligent device adaptive sound pickup methods, equipment, media and software products
CN122579025A