Voice information determination method and device, playing equipment, electronic equipment and storage medium
By obtaining the weights of the voice signal and the feedforward voice signal when the wind noise pollution level of the voice signal meets specific conditions, and combining the reference voice signal to determine the voice information, the problem of low accuracy in voice signal determination in outdoor winds and sports scenes of playback equipment is solved, and the effect of improving call quality is achieved.
Patent Information
- Application Number
- CN202311618267.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
In outdoor strong winds and sports scenarios, the microphone in the playback device generates a lot of wind noise, resulting in low accuracy in determining voice information in the voice signal and reducing call quality.
A voice information determination method is adopted to obtain the weights of the voice signal and the feedforward voice signal when the wind noise pollution level of the voice signal meets specific conditions, and combine the reference voice signal to determine the voice information. The method includes using a call microphone, a feedforward microphone and an error microphone to collect voice signals, and through specific weight configuration and processing, improving the analysis accuracy of voice information.
By flexibly configuring the proportion of voice signals and feedforward voice signals in the voice information analysis process, it is suitable for personalized wind noise scenarios, and supports accurate determination of the voice information carried in the voice signals received by the playback device, improving call quality.
Smart Images

Figure CN120075667A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of electronic devices, and particularly to a method and apparatus for determining voice information, a playback device, an electronic device, and a storage medium. Background Art
[0002] With the continuous development of playback device technologies (e.g., Bluetooth headsets, True Wireless Stereo Earbuds (TWS)), people use playback devices in more and more scenarios. As one of the main functions of playback devices, improving the effect of calls is crucial for user experience.
[0003] In related technologies, in outdoor windy and sports scenarios, the microphone in the playback device usually generates a large amount of wind noise, resulting in low accuracy in determining the voice information in the voice signal received by the playback device and reducing the call quality. Summary of the Invention
[0004] The present disclosure aims to solve at least one of the technical problems in related technologies to some extent.
[0005] To this end, the present disclosure provides a method and apparatus for determining voice information, a playback device, an electronic device, and a storage medium, which can accurately determine the voice information carried in the voice signal received by the playback device and improve the call quality.
[0006] The method for determining voice information according to the first aspect embodiment of the present disclosure is applied to a playback device, where the playback device includes: a call microphone, a feedforward microphone, and an error microphone. The call microphone is used to collect call voice signals, the feedforward microphone is used to collect feedforward voice signals, and the error microphone is used to collect reference voice signals. The method includes: when the wind noise pollution degree of the call voice signal satisfies a first condition, obtaining a first weight corresponding to the call voice signal and a second weight corresponding to the feedforward voice signal; determining the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal.
[0007] The apparatus for determining voice information according to the second aspect embodiment of the present disclosure includes: a first determination module, configured to obtain a first weight corresponding to the call voice signal collected by the call microphone and a second weight corresponding to the feedforward voice signal collected by the feedforward microphone when the wind noise pollution degree of the call voice signal collected by the call microphone satisfies a first condition; a second determination module, configured to determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal collected by the error microphone.
[0008] The playback device provided in the third aspect of the present disclosure includes: a first part, where the first part includes a first sound inlet hole and an error microphone; a second part, where the second part includes a second sound inlet hole, a feedforward microphone, and a call microphone, and the wind noise cancellation structure of the feedforward microphone includes a first channel, and the wind noise cancellation structure of the call microphone includes a second channel; wherein, the first channel communicates with the first sound inlet hole; the second channel communicates with the second sound inlet hole; the types of the first channel and the second channel are different.
[0009] The electronic device provided in the fourth aspect of the present disclosure includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method provided in the above embodiments of the present disclosure.
[0010] The fifth aspect of the present disclosure provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method provided in the above embodiments of the present disclosure.
[0011] The sixth aspect of the present disclosure provides a computer program product, and when the instructions in the computer program product are executed by a processor, they execute the method provided in the above embodiments of the present disclosure.
[0012] The voice information determination method, device, playback device, electronic device, and storage medium provided by the present disclosure obtain a first weight corresponding to the call voice signal and a second weight corresponding to the feedforward voice signal when the wind noise pollution degree of the call voice signal satisfies a first condition; and determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal. Since the proportion of the call voice signal and the feedforward voice signal in the voice information analysis process can be flexibly configured, the analysis of the voice information is more applicable to personalized wind noise scenarios, thereby supporting the accurate determination of the voice information carried in the voice signal received by the playback device and improving the call quality.
[0013] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. Description of the Drawings
[0014] The above and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0015] Figure 1 is a flowchart of the voice information determination method provided in an embodiment of the present disclosure;
[0016] Figure 2It is a schematic flowchart of a voice information determination method proposed in another embodiment of the present disclosure;
[0017] Figure 3 It is a schematic structural diagram of a voice information determination device proposed in an embodiment of the present disclosure;
[0018] Figure 4 It is a schematic structural diagram of a playback device proposed in an embodiment of the present disclosure;
[0019] Figure 5 It shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure. Detailed Embodiments
[0020] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present disclosure and should not be construed as limiting the present disclosure. On the contrary, the embodiments of the present disclosure include all changes, modifications, and equivalents falling within the spirit and scope of the appended claims.
[0021] Figure 1 It is a schematic flowchart of a voice information determination method proposed in an embodiment of the present disclosure.
[0022] In this embodiment, the voice information determination method is exemplified as being configured in a voice information determination device. In this embodiment, the voice information determination method can be configured in the voice information determination device, and the voice information determination device can be set in a terminal. The terminal includes, for example, at least one of a mobile phone, a wearable device, an Internet of Things device, an automobile with communication function, a smart automobile, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, and a wireless terminal device in a smart home, but is not limited thereto.
[0023] It should be noted that the execution subject of the embodiments of the present disclosure can be, for example, the central processing unit (CPU) of the terminal in terms of hardware, and can be, for example, the relevant background service of the terminal in terms of software, which is not limited herein. This embodiment can be specifically applied to a playback device, for example, for processing various voice signals collected by the playback device to obtain voice information.
[0024] The playback device in the embodiments of the present disclosure may include: a call microphone, a feedforward microphone, and an error microphone. The call microphone is used to collect call voice signals, the feedforward microphone is used to collect feedforward voice signals, and the error microphone is used to collect reference voice signals. Among them, the call microphone can be denoted as Talk mic, the feedforward microphone can be denoted as FF mic, and the error microphone can be denoted as FB mic. For the descriptions of Talk mic, FF mic, and FB mic, reference can be made to the related art and will not be elaborated herein.
[0025] As Figure 1 shown, the method for determining the voice information includes:
[0026] S101: When the wind noise pollution degree of the call voice signal meets the first condition, obtain the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal.
[0027] Among them, the call voice signal can refer to the voice signal collected by the call microphone.
[0028] In some embodiments, the wind noise pollution degree of the call voice signal can be detected. The wind noise pollution degree is used to describe the situation of the call voice signal being polluted by wind noise. If the call voice signal is severely polluted by wind noise, the wind noise pollution degree is relatively severe; if the call voice signal is slightly polluted by wind noise, the wind noise pollution degree is relatively light.
[0029] Among them, the first condition can be a threshold condition for determining that the wind noise pollution degree is relatively light. The first condition can be, for example, that the wind noise pollution degree is less than the pollution degree threshold, or can be, for example, that the signal quality of the call voice signal determined based on the wind noise pollution degree is greater than the quality threshold (i.e., the signal quality is good), or can also be, based on the wind noise pollution degree, determining that effective voice information can be effectively recognized based on the call voice signal. At this time, it can all indicate that the wind noise pollution degree of the call voice signal meets the first condition.
[0030] In some scenarios, if the wind noise pollution degree of the call voice signal meets the first condition, the call voice signal quality is better, and it can support the recognition of valid voice information based on the call voice signal. In other scenarios, if the wind noise pollution degree of the call voice signal does not meet the first condition, the call voice signal is severely affected by noise, and at this time, the call voice signal quality is poor and cannot support the recognition of valid voice information based on the call voice signal.
[0031] In the embodiments of the present disclosure, it is possible to first detect whether the wind noise pollution degree of the call voice signal meets the first condition. When it is determined that the wind noise pollution degree meets the first condition, it is possible to obtain the first weight to which the call voice signal is referred during the determination of the voice information, and the second weight to which the feedforward voice signal is referred, so as to implement considering at least part of the call voice signal during the determination of the voice information.
[0032] In some embodiments, during the process of detecting whether the wind noise pollution degree of the call voice signal meets the first condition, it is possible to analyze the situation of the call voice signal being affected by wind noise. If the call voice signal is severely affected by wind noise, it is determined that the wind noise pollution degree does not meet the first condition. If the call voice signal is slightly affected by wind noise, it is determined that the wind noise pollution degree meets the first condition; alternatively, the call voice signal can also be input into an analysis model, and based on the output of the analysis model, it is determined whether the wind noise pollution degree of the call voice signal meets the first condition. The analysis model can be an artificial intelligence model, and the analysis model has learned the mapping relationship between the call voice signal and whether its wind noise pollution degree meets the first condition; of course, it is also possible to implement the detection of whether the wind noise pollution degree of the call voice signal meets the first condition based on any other possible method.
[0033] In the embodiments of the present disclosure, it is possible to obtain the call signal characteristics of the call voice signal, and detect the noise characteristic value from the call signal characteristics. If the noise characteristic value is less than the noise threshold, it is determined that the wind noise pollution degree of the call voice signal meets the first condition; if the noise characteristic value is greater than or equal to the noise threshold, it is determined that the wind noise pollution degree of the call voice signal does not meet the first condition. This realizes accurately and quickly detecting whether the wind noise pollution degree of the call voice signal meets the first condition.
[0034] Among them, the signal characteristics extracted from the call voice signal can be referred to as call signal characteristics.
[0035] Exemplarily, to detect the noise eigenvalue from the call signal features, the call signal features can be input into a classification model, and the output noise eigenvalue of the classification model can be obtained. The classification model can be, for example, a support vector machine, a decision tree, a logistic regression, and a neural network, etc. When establishing the classification model, the data set can be divided into a training set and a test set. The model can be trained using the training set, and then the performance of the model can be evaluated using the test set. During the training and testing process of the classification model, the model parameters and feature selection strategies can be continuously adjusted to improve the accuracy and stability of the model. This classification model can achieve noise detection. For example, for the actual call signal features, noise detection and classification can be performed to determine the noise eigenvalue contained in the call signal features.
[0036] Among them, the noise threshold is the threshold value corresponding to the noise eigenvalue when it is determined that the wind noise pollution degree of the call voice signal to which the corresponding noise eigenvalue belongs is relatively serious. If the noise eigenvalue is greater than or equal to the noise threshold, it means that the wind noise pollution degree of the call voice signal is relatively serious and does not meet the first condition. If the noise eigenvalue is less than the noise threshold, it means that the wind noise pollution degree of the call voice signal is relatively light or there is no pollution, and the first condition is met.
[0037] In some embodiments, if it is determined that the wind noise pollution degree of the call voice signal meets the first condition, the weight situation of the call voice signal being referenced during the determination of the voice information can be obtained, and the first weight can be used to describe this weight situation. The weight situation of the feedforward voice signal being referenced during the determination of the voice information can also be obtained, and the second weight can be used to describe this weight situation.
[0038] Among them, the first weight can be, for example, the proportion of the call signal features of the call voice signal being referenced. The second weight can be, for example, the proportion of the feedforward signal features of the feedforward voice signal (the signal features of the feedforward voice signal can be referred to as feedforward signal features) being referenced.
[0039] In some embodiments, some preset reference weights can be configured for the call microphone, and then a suitable reference weight can be selected as the first weight based on the actual call voice signal collected by the call microphone in this call. Some preset reference weights can also be configured for the feedforward microphone, and then a suitable reference weight can be selected as the second weight based on the actual call voice signal collected by the feedforward microphone in this call. Or the reference weights corresponding to the call microphone and the reference weights corresponding to the feedforward microphone can be jointly configured. For example, the reference weights of different microphones are configured in a combined form. After determining the first weight, the second weight corresponding to the first weight can be directly determined.
[0040] S102: Determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal.
[0041] After determining the first weight and the second weight, the voice information may be determined according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal.
[0042] In some embodiments, determining the voice information may refer to the process of combining various voice signals to identify the voice information contained therein. The voice information may specifically be, for example, information such as the valid voice, speech segments, voice content, semantics, etc. contained in the voice signal. The process of determining the voice information may specifically be, for example, restoring the voice, restoring the voice and extracting the valid information in the voice, etc., and this is not limited thereto.
[0043] In some embodiments, the reference voice signal may be used as an aid to jointly process the call voice signal and the feedforward voice signal to extract the effective voice information. By way of example, a voice recognition model may be established in advance, and this model is usually constructed using a deep neural network or a recurrent neural network. The voice recognition model usually includes multiple hidden layers for extracting high-level representations of the input features and then mapping these representations to the text output. The call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal may be input into the voice recognition model, and the voice recognition model is used to perform model processing on the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal, and the voice information output by the voice recognition model is obtained.
[0044] In some embodiments, the call voice signal, the feedforward voice signal, and the reference voice signal may be preprocessed respectively. For example, various voice signals are filtered, gain-adjusted, noise-reduced, etc. Then, the preprocessed various voice signals are divided into several frames, and usually the duration of each frame is 10 - 30 milliseconds. The frame processing can retain the time-domain information of the signal. Features are extracted from each frame of the voice signal. The features are, for example, Mel frequency cepstral coefficients, linear predictive coding, etc., and the voice information is analyzed based on various signal features.
[0045] Of course, other arbitrary possible ways may also be adopted to determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal.
[0046] In some embodiments, call signal features can be extracted from a call voice signal, feedforward signal features can be extracted from a feedforward voice signal, reference signal features can be extracted from a reference voice signal, the call signal features can be processed based on a first weight to obtain target call signal features, the feedforward signal features can be processed based on a second weight to obtain target feedforward signal features, and voice information can be determined according to the target call signal features, the target feedforward signal features, and the reference signal features. Since the call signal features are processed with reference to the first weight, the feedforward signal features are processed with reference to the second weight, and the voice information is determined with the assistance of the reference signal features based on the obtained target call signal features and target feedforward signal features, the accuracy and convenience of voice information determination can be greatly improved.
[0047] Among them, processing the call signal features based on the first weight can be, for example, weighting the call signal features with the first weight and using the weighted call signal features as the target call signal features. Processing the feedforward signal features based on the second weight can be, for example, weighting the feedforward signal features with the second weight and using the weighted feedforward signal features as the target feedforward signal features.
[0048] In this embodiment, when the wind noise pollution degree of the call voice signal satisfies a first condition, the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal are obtained; the voice information is determined according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal. Since the proportion of the call voice signal and the feedforward voice signal in the voice information analysis process can be flexibly configured, the analysis of the voice information is more applicable to personalized wind noise scenarios, thereby supporting the accurate determination of the voice information carried in the voice signal received by the playback device and improving the call quality.
[0049] In some embodiments, when the wind noise pollution degree of the call voice signal does not satisfy the first condition, the voice information can be determined according to the feedforward voice signal and the reference voice signal. Thus, when determining the voice information, the call voice signal severely polluted by wind noise can be not considered, and the feedforward voice signal and the reference voice signal can be combined to determine the voice information, thereby effectively reducing the wind noise introduced in the voice information determination process and supporting the improvement of the recognition accuracy of the voice information.
[0050] In some embodiments, when performing the step of determining the voice information according to the feedforward voice signal and the reference voice signal, the feedforward signal features can also be extracted from the feedforward voice signal, the reference signal features can be extracted from the reference voice signal, and the voice information can be determined according to the feedforward signal features and the reference signal features. Thus, it supports correctly combining various voice signals to determine the voice information.
[0051] In the embodiments of the present disclosure, in the process of extracting the feedforward signal features from the feedforward voice signal, the feedforward voice signal may be input into a noise reduction model to obtain the target feedforward voice signal output by the noise reduction model. The noise reduction model is used to eliminate the first type of noise in the feedforward voice signal. The noise reduction model has learned the mapping relationship between the feedforward voice signal and the target feedforward voice signal. The target feedforward voice signal does not contain the first type of noise, and the feedforward signal features are extracted from the target feedforward voice signal. This can further reduce the influence of wind noise, thereby supporting the further improvement of the accuracy of voice signal determination.
[0052] Among them, the first type of noise may be the wind noise introduced by the wind noise reduction structure of the feedforward microphone.
[0053] That is to say, a noise reduction model can be pre-trained. The noise reduction model has the function of eliminating the first type of noise in the feedforward voice signal collected by the feedforward microphone. The noise reduction model can be an artificial intelligence model and can be trained based on a deep learning dataset. It has learned the mapping relationship between the feedforward voice signal and the target feedforward voice signal. The target feedforward voice signal does not contain the first type of noise. That is to say, the target feedforward voice signal is the feedforward voice signal from which the first type of noise has been eliminated.
[0054] The deep learning dataset may include the sample feedforward voice signal and the labeled target feedforward voice signal. The voice signal of the sample can be collected in a scenario with better passive wind resistance effect. Then, the sample feedforward voice signal is input into the artificial intelligence model, and the predicted target feedforward voice signal output by the artificial intelligence model is obtained, and multiple rounds of iterative training are performed to support the convergence condition between the predicted target feedforward voice signal and the labeled target feedforward voice signal. The trained artificial intelligence model is used as the noise reduction model.
[0055] Among them, the first type of noise may be, for example, the whistling sound generated by the T-shaped through hole. The T-shaped through hole may be, for example, the through hole of the wind noise reduction structure of the feedforward microphone. Specifically, reference can be made to the structure description of the playback device in the following embodiments.
[0056] Figure 2 It is a schematic flowchart of a voice information determination method proposed in another embodiment of the present disclosure.
[0057] As Figure 2 shown, the voice information determination method includes:
[0058] S201: Determine that the wind noise pollution degree of the call voice signal meets the first condition.
[0059] For the description of S201, specific reference can be made to the above embodiments and will not be elaborated here.
[0060] S202: Determine multiple reference suppression frequency bands, where a reference suppression frequency band represents the frequency range corresponding to the wind noise signal that can be suppressed by the call microphone or the feedforward microphone.
[0061] Among them, one reference suppression frequency band can represent the frequency range corresponding to the wind noise signal that can be suppressed by the call microphone, and another reference suppression frequency band can represent the frequency range corresponding to the wind noise signal that can be suppressed by the feedforward microphone. Different reference suppression frequency bands can represent the frequency ranges corresponding to the wind noise signals that can be suppressed by different microphones, and different reference suppression frequency bands can also represent the frequency ranges corresponding to the wind noise signals that can be suppressed by the same microphone. Within the same reference suppression frequency band, it can include the frequency part corresponding to the call microphone or the frequency part corresponding to the feedforward microphone, and there is no limitation on this.
[0062] Exemplarily, the optimal wind noise suppression frequency band of the feedforward microphone (for example, a1 Hz (Hertz) to a2 Hz) and the optimal wind noise suppression frequency band of the call microphone (for example, b1 Hz to b2 Hz) can be confirmed in advance based on experimental detection. Moreover, the suppression of the through holes of the feedforward microphone on the wind in all directions is more obvious in the frequency band of 100 Hz to 1000 Hz, while the through holes of the call microphone have a suppression effect on the frequency band above 600 Hz in the 0° and 180° directions of the wind. Generally, a1 Hz < b1 Hz < a2 Hz < b2 Hz. Then, the wind noise suppression frequency band can be divided into 100 Hz to a1 Hz, a1 Hz to b1 Hz, b1 Hz to a2 Hz, a2 Hz to b2 Hz, and b2 to 8 kHz. These frequency ranges can all be referred to as reference suppression frequency bands.
[0063] In some embodiments, the multiple determined reference suppression frequency bands can be used to assist in determining the first weight and the second weight.
[0064] S203: Obtain the first weight and the second weight according to the call voice signal and the multiple reference suppression frequency bands.
[0065] It can be understood that when the frequency of the wind noise signal carried by the call voice signal (since different microphones may collect voice signals based on the same scenario, the frequency of the wind noise signal in the call voice signal can be the same as that of the wind noise signal in the feedforward microphone, and here the frequency of the wind noise signal carried by the feedforward voice signal can also be used) is within a certain reference suppression frequency band mentioned above, if this certain reference suppression frequency band corresponds to the call microphone, it means that the call microphone can effectively suppress this wind noise signal. At this time, a relatively large first weight can be configured for the call voice signal (since the call microphone can already effectively suppress this wind noise signal, the call voice signal is relatively less affected by this type of wind noise signal), so as to increase the proportion of the call voice signal being referenced in the process of determining the voice information. If this certain reference suppression frequency band corresponds to the feedforward microphone, it means that the feedforward microphone can effectively suppress this wind noise signal. At this time, a relatively large second weight can be configured for the feedforward voice signal (since the feedforward microphone can already effectively suppress this wind noise signal, the feedforward voice signal is relatively less affected by this type of wind noise signal), so as to increase the proportion of the feedforward voice signal being referenced in the process of determining the voice information.
[0066] Therefore, the frequency of the wind noise signal in the call voice signal can be compared with the above various reference suppression frequency bands, so as to adaptively allocate the first weight to the call voice signal and allocate the second weight to the feedforward voice signal, thereby improving the flexibility and applicability of the allocation of the first weight and the second weight. For any type of wind noise signal, a better noise reduction effect can be achieved, and thus the accuracy of voice information determination can be greatly improved.
[0067] In some embodiments, in order to effectively improve the efficiency and convenience of obtaining the first weight and the second weight, and to support the improvement of the efficiency of voice signal extraction, corresponding first reference weights and second reference weights can also be pre-allocated for each reference suppression frequency band. The first reference weight is the reference weight pre-configured for the call microphone, and the second reference weight is the reference weight pre-configured for the feedforward microphone. Then, in the process of obtaining the first weight and the second weight according to the call voice signal and multiple reference suppression frequency bands, it can be to determine the frequency of the wind noise signal included in the call voice signal, determine the first reference suppression frequency band to which the frequency of the wind noise signal belongs from the multiple reference suppression frequency bands, and use the first reference weight corresponding to the first reference suppression frequency band as the first weight, and use the second reference weight corresponding to the first reference suppression frequency band as the second weight.
[0068] In some embodiments, the frequency ranges of the multiple reference suppression frequency bands increase gradually from small to large.
[0069] In some embodiments, the first weight corresponding to the reference suppression frequency band with a smaller frequency range is less than the first weight corresponding to the reference suppression frequency band with a larger frequency range; the second weight corresponding to the reference suppression frequency band with a smaller frequency range is greater than the second weight corresponding to the reference suppression frequency band with a larger frequency range. It can effectively be applicable to the optimal wind noise suppression frequency band of the call microphone and the feedforward microphone, and effectively improve the acquisition accuracy of the first weight and the second weight.
[0070] The setting method for multiple reference suppression frequency bands can be illustrated as follows:
[0071] (1) Obtain the measured results. The optimal wind noise suppression frequency band (a1 Hz to a2 Hz) of the feedforward microphone FF MIC and the optimal wind noise suppression frequency band (b1 Hz to b2 Hz) of the call microphone TALK MIC can be confirmed respectively. In the embodiments of the present disclosure, the through hole of the wind noise resistance structure of the feedforward microphone FF MIC can be a T-shaped through hole, and the suppression of the wind in all directions by the T-shaped through hole is more obvious in the frequency band of 100 Hz to 1000 Hz, while the through hole of the wind noise resistance structure of the call microphone TALK MIC is an L-shaped through hole, which mainly has a suppression effect on the frequency band above 600 Hz in the 0° and 180° wind directions. Generally, a1 Hz < b1 Hz < a2 Hz < b2 Hz.
[0072] (2) For the measured results, the wind noise suppression frequency band is divided into 100 Hz to a1 Hz, a1 Hz to b1 Hz, b1 Hz to a2 Hz, a2 Hz to b2 Hz, b2 Hz to 8 kHz. If the call voice signal is obtained, it can be judged whether the call voice signal is contaminated by wind noise. If the contamination is relatively serious and the call voice signal does not support the recognition of language information, it can be judged as lateral wind, and the call voice signal is not used, but the feedforward voice signal is used, and the reference voice signal of the error microphone FB MIC is used as a reference for voice restoration (i.e., recognition of language information). If the call voice signal supports the recognition of language information, the voice signal weights of FF MIC / TALK MIC can be set to 95% / 5%, 75% / 25%, 55% / 45%, 35% / 65%, 25% / 75% respectively, and the reference voice signal of the error microphone FB MIC is used as a reference for voice restoration.
[0073] Among them, the first reference weight is, for example, 5%, 25%, 45%, 65%, 75%. The second reference weight is, for example, 95%, 75%, 55%, 35%, 25%. The voice signal of the FF MIC, i.e., an optional example of the feedforward voice signal. The voice signal of the TALK MIC, i.e., an optional example of the voice signal of the call microphone. 100 Hz to a1 Hz, a1 Hz to b1 Hz, b1 Hz to a2 Hz, a2 Hz to b2 Hz, b2 Hz to 8 kHz can be some optional examples of multiple reference suppression frequency bands. The first reference weight corresponding to the reference suppression frequency band 100 Hz to a1 Hz is 5%, and the second reference weight is 95%; the first reference weight corresponding to the reference suppression frequency band a1 Hz to b1 Hz is 25%, and the second reference weight is 75%; the first reference weight corresponding to the reference suppression frequency band b1 Hz to a2 Hz is 45%, and the second reference weight is 55%; the first reference weight corresponding to the reference suppression frequency band a2 Hz to b2 Hz is 65%, and the second reference weight is 35%; the first reference weight corresponding to the reference suppression frequency band b2 Hz to 8 kHz is 75%, and the second reference weight is 25%.
[0074] It can be seen that the first weight corresponding to the reference suppression frequency band with a smaller frequency range is less than the first weight corresponding to the reference suppression frequency band with a larger frequency range; the second weight corresponding to the reference suppression frequency band with a smaller frequency range is greater than the second weight corresponding to the reference suppression frequency band with a larger frequency range.
[0075] S204: Determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal.
[0076] For the description of S204, specific reference can be made to the above embodiments, which will not be elaborated here.
[0077] In this embodiment, when the wind noise pollution degree of the call voice signal meets the first condition, the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal are obtained; the voice information is determined according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal. Since the proportion of the call voice signal and the feedforward voice signal in the voice information analysis process can be flexibly configured, the analysis of the voice information is more applicable to personalized wind noise scenarios, thereby supporting accurately determining the voice information carried in the voice signal received by the playback device and improving the call quality. The frequency of the wind noise signal in the call voice signal can be compared with the above various reference suppression frequency bands to adaptively assign the first weight to the call voice signal and the second weight to the feedforward voice signal, so as to improve the flexibility and applicability of the assignment of the first weight and the second weight. For any type of wind noise signal, a good noise reduction effect can be achieved. Therefore, the accuracy of determining the voice information is greatly improved. It is also possible to restore more details of the voice information and suppress the wind noise in a wider frequency band. And it is possible to well eliminate the whistling sound of the playback device under the condition that the data scale remains basically unchanged.
[0078] Figure 3 FIG. 4 is a schematic structural diagram of a voice information determination device according to an embodiment of the present disclosure.
[0079] As Figure 3 shown, the voice information determination device 30 includes:
[0080] A first determination module 301, configured to obtain a first weight corresponding to the call voice signal and a second weight corresponding to the feedforward voice signal collected by the feedforward microphone when the wind noise pollution degree of the call voice signal collected by the call microphone meets the first condition.
[0081] A second determination module 302, configured to determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal collected by the error microphone.
[0082] It should be noted that the foregoing explanation of the voice information determination method is also applicable to the voice information determination device in this embodiment, and will not be repeated here.
[0083] In this embodiment, when the degree of wind noise pollution of the call voice signal meets the first condition, the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal are obtained; the voice information is determined according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal. Since the proportion of the call voice signal and the feedforward voice signal in the voice information analysis process can be flexibly configured, the analysis of the voice information is more applicable to personalized wind noise scenarios, thereby supporting the accurate determination of the voice information carried in the voice signal received by the playback device and improving the call quality.
[0084] Figure 4 It is a schematic structural diagram of a playback device proposed in an embodiment of the present disclosure.
[0085] As Figure 4 shown, the playback device 40 includes:
[0086] The first part 401, where the first part 401 includes a first sound inlet hole 4011 and an error microphone 4012.
[0087] The second part 402, where the second part 402 includes a second sound inlet hole 4021, a feedforward microphone 4022, and a call microphone 4023. The wind noise resistance structure of the feedforward microphone 4022 includes a first channel 40221, and the wind noise resistance structure of the call microphone 4023 includes a second channel 40231; where
[0088] The first channel 40221 communicates with the first sound inlet 4011 hole.
[0089] The second channel 40231 communicates with the second sound inlet hole 4021.
[0090] The types of the first channel 40221 and the second channel 40231 are different.
[0091] In some embodiments, the type of the first channel 40221 is T-shaped; the type of the second channel 40231 is L-shaped.
[0092] It should be noted that the foregoing explanation of the voice information determination method also applies to the playback device of this embodiment, and will not be repeated here.
[0093] In this embodiment, when the wind noise pollution degree of the call voice signal meets the first condition, a first weight corresponding to the call voice signal and a second weight corresponding to the feedforward voice signal are obtained; the voice information is determined according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal. Since the proportion of the call voice signal and the feedforward voice signal in the voice information analysis process can be flexibly configured, the analysis of the voice information is more applicable to personalized wind noise scenarios, thereby supporting accurately determining the voice information carried in the voice signal received by the playback device and improving the call quality.
[0094] Figure 5 FIG. shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure. Figure 5 The shown electronic device 12 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0095] As Figure 5 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a memory 28, and a bus 18 connecting different system components (including the memory 28 and the processing unit 16).
[0096] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.
[0097] The electronic device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0098] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 5 not shown, commonly referred to as a "hard disk drive").
[0099] Although Figure 5 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as compact disc read only memory (CD-ROM), digital video disc read only memory (DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 through one or more data media interfaces. The memory 28 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.
[0100] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present disclosure.
[0101] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a human body to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0102] The processing unit 16 executes various functional applications and data processing by running programs stored in the memory 28, such as implementing the methods mentioned in the foregoing embodiments.
[0103] To implement the foregoing embodiments, the present disclosure also proposes a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method proposed in the foregoing embodiments of the present disclosure.
[0104] To implement the foregoing embodiments, the present disclosure also proposes a computer program product, and when the instructions in the computer program product are executed by a processor, they execute the method proposed in the foregoing embodiments of the present disclosure.
[0105] It should be noted that in the description of the present disclosure, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present disclosure, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0106] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, and this should be understood by those skilled in the technical field to which the embodiments of the present disclosure belong.
[0107] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0108] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0109] In addition, in each embodiment of the present disclosure, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0110] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0111] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0112] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for determining voice information, characterized in that, it is applied to a playback device, and the playback device includes: a call microphone, a feedforward microphone, and an error microphone. The call microphone is used to collect call voice signals, the feedforward microphone is used to collect feedforward voice signals, and the error microphone is used to collect reference voice signals; wherein, the method includes: When the wind noise pollution degree of the call voice signal meets the first condition, obtain the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal; Determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal.
2. The method according to claim 1, characterized in that, the method further includes: When the wind noise pollution degree of the call voice signal does not meet the first condition, determine the voice information according to the feedforward voice signal and the reference voice signal.
3. The method according to any one of claims 1-2, characterized in that, the method further includes: Obtain the call signal characteristics of the call voice signal; Detect the noise characteristic value from the call signal characteristics; If the noise characteristic value is less than the noise threshold, it is determined that the wind noise pollution degree of the call voice signal meets the first condition; If the noise characteristic value is greater than or equal to the noise threshold, it is determined that the wind noise pollution degree of the call voice signal does not meet the first condition.
4. The method according to claim 1, characterized in that, the obtaining of the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal includes: Determine a plurality of reference suppression frequency bands, wherein the reference suppression frequency bands represent the frequency ranges corresponding to the wind noise signals that the call microphone or the feedforward microphone supports suppressing; Obtain the first weight and the second weight according to the call voice signal and the plurality of reference suppression frequency bands.
5. The method according to claim 4, characterized in that, Each reference suppression frequency band corresponds to a first reference weight and a second reference weight. The first reference weight is a reference weight pre-configured for the call microphone, and the second reference weight is a reference weight pre-configured for the feedforward microphone; wherein, the obtaining of the first weight and the second weight according to the call voice signal and the plurality of reference suppression frequency bands includes: Determine the frequency of the wind noise signal included in the call voice signal; Determine the first reference suppression frequency band to which the frequency of the wind noise signal belongs from the plurality of reference suppression frequency bands; Use the first reference weight corresponding to the first reference suppression frequency band as the first weight, and use the second reference weight corresponding to the first reference suppression frequency band as the second weight.
6. The method according to claim 4, characterized in that, the frequency ranges of the plurality of reference suppression frequency bands increase in ascending order from small to large.
7. The method according to claim 6, characterized in that, wherein, The first weight corresponding to the reference suppression frequency band with a smaller frequency range is less than the first weight corresponding to the reference suppression frequency band with a larger frequency range; The second weight corresponding to the reference suppression frequency band with a smaller frequency range is greater than the second weight corresponding to the reference suppression frequency band with a larger frequency range.
8. The method according to claim 1, wherein, the determining the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal includes: extracting call signal features from the call voice signal; extracting feedforward signal features from the feedforward voice signal; extracting reference signal features from the reference voice signal; processing the call signal features based on the first weight to obtain target call signal features, and processing the feedforward signal features based on the second weight to obtain target feedforward signal features; determining the voice information according to the target call signal features, the target feedforward signal features, and the reference signal features.
9. The method according to claim 2, wherein, the determining the voice information according to the feedforward voice signal and the reference voice signal includes: extracting feedforward signal features from the feedforward voice signal; extracting reference signal features from the reference voice signal; determining the voice information according to the feedforward signal features and the reference signal features.
10. The method according to any one of claims 8-9, wherein, the extracting feedforward signal features from the feedforward voice signal includes: inputting the feedforward voice signal into a noise reduction model to obtain a target feedforward voice signal output by the noise reduction model, wherein the noise reduction model is used to eliminate the first type of noise in the feedforward voice signal, the noise reduction model has learned the mapping relationship between the feedforward voice signal and the target feedforward voice signal, and the target feedforward voice signal does not contain the first type of noise; extracting the feedforward signal features from the target feedforward voice signal.
11. A voice information determining device, wherein, the device includes: a first determining module, configured to obtain the first weight corresponding to the call voice signal and the second weight corresponding to the feedforward voice signal collected by a feedforward microphone when the wind noise pollution degree of the call voice signal collected by a call microphone satisfies a first condition; a second determining module, configured to determine the voice information according to the call voice signal, the feedforward voice signal, the first weight, the second weight, and the reference voice signal collected by an error microphone.
12. A playback device, wherein, the playback device includes: a first part, wherein the first part includes a first sound inlet hole and an error microphone; a second part, wherein the second part includes a second sound inlet hole, a feedforward microphone, and a call microphone, and the wind noise resistance structure of the feedforward microphone includes a first channel, and the wind noise resistance structure of the call microphone includes a second channel; wherein, the first channel is communicated with the first sound inlet hole; the second channel is communicated with the second sound inlet hole; The types of the first channel and the second channel are different.
13. The playback device according to claim 12, characterized in that, wherein, the type of the first channel is T-shaped; the type of the second channel is L-shaped.
14. An electronic device, characterized in that, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-10.
15. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, wherein, the computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.