Human voice transparent transmission method and device, earphone, storage medium and program product

By incorporating a voice recognition module and a lightweight voice separation network into the headphones, the human voice signal in the external audio signal is first identified and then separated, thus solving the problems of poor headphone transmission effect and high power consumption, and enabling the function of hearing human voices while reducing noise.

CN116343756BActive Publication Date: 2026-01-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111582502.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2026-01-13
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

Existing headphones cannot effectively distinguish between the target sound source and other sound sources in the pass-through function, resulting in poor pass-through performance and high power consumption.

Method used

By incorporating a voice recognition module and a lightweight voice separation network into the headphones, the system first identifies the voice signal in the external audio signal, then separates and mixes it to generate a signal that drives the speaker to produce sound, thus achieving voice transmission.

Benefits of technology

It improves the headphone's transparency, allowing users to easily hear surrounding voices while enjoying noise cancellation, and reduces the headphone's power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343756B_ABST
    Figure CN116343756B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a human voice transparent transmission method and device, earphones, a storage medium and a program product, and belongs to the technical field of audio processing. The method is used for earphones, and the method comprises the following steps: performing human voice identification on collected external audio signals; in the case that it is identified that the external audio signals contain human voice signals, separating the human voice signals from the external audio signals; performing mixing processing on the separated human voice signals and noise reduction signals to obtain mixed signals, wherein the noise reduction signals are used for active noise reduction; and driving a loudspeaker to sound based on the mixed signals. The scheme of the embodiment of the application can improve the human voice transparent transmission effect of the earphones, and simultaneously reduce the power consumption of the human voice transparent transmission system of the earphones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a method, apparatus, earphone, storage medium, and program product for human voice transmission. Background Technology

[0002] With the improvement of living standards, headphones have become an indispensable part of people's lives. In noisy environments such as airports, subways, and restaurants, the noise cancellation function of headphones can eliminate external noise interference to the greatest extent. However, in scenarios where users need to receive external voices and ambient noise, headphones also need to have a pass-through function to transmit external sound signals to the user, so that the user can hear external sounds without removing the headphones.

[0003] In related technologies, the pass-through function of headphones transmits the target sound source signal and other sound source signals that the user needs to hear to the user. Therefore, the sound heard by the user includes the target sound source and other sound sources, which reduces the pass-through effect. Summary of the Invention

[0004] This application provides a method, apparatus, earphone, storage medium, and program product for transmitting human voice, the technical solution of which is as follows:

[0005] On one hand, embodiments of this application provide a method for transmitting human voice, the method being used in headphones, the method comprising:

[0006] Human voice recognition is performed on the collected external audio signals;

[0007] If the external audio signal is found to contain a human voice signal, the human voice signal is separated from the external audio signal;

[0008] The separated human voice signal and the noise-reduced signal are mixed to obtain a mixed signal, and the noise-reduced signal is used for active noise reduction.

[0009] The speaker is driven to produce sound based on the mixed signal.

[0010] On the other hand, embodiments of this application provide a voice transmission device for headphones, the device comprising:

[0011] The voice recognition module is used to recognize human voices from the collected external audio signals;

[0012] A separation module is used to separate the human voice signal from the external audio signal when the external audio signal is identified to contain a human voice signal;

[0013] The mixing module is used to mix the separated human voice signal and the noise reduction signal to obtain a mixed signal, and the noise reduction signal is used for active noise reduction;

[0014] A driver module is used to drive the speaker to produce sound based on the mixed signal.

[0015] On the other hand, embodiments of this application provide an earphone, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set, or an instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the voice pass-through method as described above.

[0016] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the voice pass-through method as described above.

[0017] On the other hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions to perform the voice pass-through method provided in the above aspects.

[0018] The technical solution provided in this application may include the following beneficial effects:

[0019] In this embodiment, the headphones first perform voice recognition on the acquired external audio signal. If a voice signal is detected within the external audio signal, it is then separated from the external audio signal. The separated voice signal is then mixed with a noise-reduced signal to generate a single signal that drives the speaker, thus achieving the headphone's pass-through function. In this embodiment, the headphones only pass through the voice signal, allowing users to enjoy noise reduction while still easily hearing surrounding voices, improving the headphone's pass-through effect. Furthermore, by first recognizing the voice signal from the external audio signal and then separating it only after detection, the headphones' power consumption is reduced. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] Figure 1A schematic diagram illustrating the headphone pass-through principle provided in an exemplary embodiment of this application is shown;

[0022] Figure 2 A schematic diagram of a DSP-based headphone pass-through principle provided in an exemplary embodiment of this application is shown;

[0023] Figure 3 A schematic diagram of an NPU-based headphone pass-through principle provided in an exemplary embodiment of this application is shown;

[0024] Figure 4 A flowchart illustrating a voice pass-through method provided in an exemplary embodiment of this application is shown;

[0025] Figure 5 A flowchart of a voice pass-through method provided in another exemplary embodiment of this application is shown;

[0026] Figure 6 A schematic diagram illustrating the VAD classifier training process provided in an exemplary embodiment of this application is shown.

[0027] Figure 7 A flowchart of a voice separation method provided in an exemplary embodiment of this application is shown;

[0028] Figure 8 This illustration shows a schematic diagram of the process of obtaining a human voice probability matrix using U-net according to an exemplary embodiment of this application;

[0029] Figure 9 This illustration shows a schematic diagram of a voice separation process provided in an exemplary embodiment of this application;

[0030] Figure 10 This invention provides a structural block diagram of a voice transmission device according to an exemplary embodiment of the present application.

[0031] Figure 11 A structural block diagram of an earphone provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0033] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0034] In related technologies, headphones not only have Active Noise Cancellation (ANC) mode but also Hear Through (HT) mode. When active noise cancellation is enabled, users can enjoy a comfortable noise-canceling experience in various noisy environments such as airports, subways, and restaurants. However, in some situations, users still need to receive external speech or ambient noise. For example, on the subway, users need to hear subway announcements to avoid missing their stop, or they need to hear conversations around them. In these cases, headphones need to activate the Hear Through mode, which transmits the target sound signal to the user's ear. Figure 1 The diagram illustrates a schematic of the headphone pass-through principle provided in an exemplary embodiment of this application. It includes a Playback path 110, an ANC path 120, a pass-through path 130, and a mixing circuit 140. The Playback path 110 plays music or downlink call data. The ANC path 120 is used for active noise cancellation, which works by generating an inverse sound wave equal to the external noise signal to neutralize the noise and achieve noise reduction. Optionally, the ANC path 120 can be a feedforward structure, a feedback structure, or a hybrid structure, i.e., combining both feedforward and feedback structures. The pass-through path 130 processes the external audio signal that needs to be passed through. The mixing circuit 140 mixes the music signal or downlink call signal from the Playback path 110, the noise-reduced signal from the ANC path 120, and the target sound signal from the pass-through path 130 into a single signal, which drives the headphone speaker to produce sound.

[0035] In related technologies, such as Figure 2As shown, the transmission path 130 is implemented based on a DSP (Digital Signal Processor). The headphones contain a DSP that performs EQ (Equalizer) processing on external audio signals, adjusting the gain or attenuation of one or more frequency bands to adjust the timbre and make the sound more comfortable for the human ear. However, this method does not differentiate between external audio signals; it transmits both speech and environmental noise signals to the ear, resulting in a noisy, mixed-up sound and poor transmission quality. Optionally, the DSP can be an FIR (Finite Impulse Response) filter, an IIR (Infinite Impulse Response) filter, etc., but this embodiment does not limit the specific type of filter used.

[0036] For example, Figure 2 The system uses FIR or IIR filters to perform EQ processing on external audio signals. Where x k (n) is used to represent the input signal, i.e., the external audio signal, y k (n) is used to represent the output signal, that is, the processed external audio signal, b k0 b k1 b k2 Used to represent filter coefficients, a k1 a k2 z is used to represent the feedback coefficient. -1 Used to characterize the Z-transform.

[0037] In the embodiments of this application, in order to improve the transmission effect, such as Figure 3 As shown, the pass-through path 130 is implemented based on an NPU (Neural Processing Unit). The earphone incorporates an NPU, which separates human voices from external audio signals to obtain clean human voices, improving the earphone's pass-through effect. However, since current earphone batteries generally have small capacities (approximately 50mAh), separating human voices from external audio signals requires complex calculations, resulting in significant power consumption for the earphone. Therefore, to reduce the overall power consumption of the earphone's human voice pass-through system, in this embodiment, the earphone first performs human voice recognition on the external audio signal. When a human voice signal is detected within the external audio signal, human voice separation is then performed, thereby reducing the computational cost of human voice separation and consequently lowering the earphone's power consumption. The human voice pass-through method in this embodiment is described below.

[0038] Please refer to Figure 4 The diagram illustrates a flowchart of a voice pass-through method provided in an exemplary embodiment of this application, the method comprising:

[0039] Step 410: Perform human voice recognition on the collected external audio signals.

[0040] The headphones use a microphone to collect external audio signals in real time and perform human voice recognition on the external audio signals.

[0041] Optionally, the headphones can be wireless headphones, wired headphones, etc., and this application embodiment does not limit this.

[0042] Optionally, the external audio signal may include human voice signals, environmental noise signals, music signals, etc., and this application embodiment does not limit this.

[0043] In one possible implementation, the device used for voice recognition in the earphone can be an NPU, DSP, MCU (Microcontroller Unit), VAD (Voice Activity Detection) hardware, etc., and this application embodiment does not limit this. For example, the VAD is used to detect the start position of the voice signal, separate the voice end and the non-voice end, thereby achieving the purpose of voice recognition.

[0044] Step 420: If the external audio signal is found to contain a human voice signal, separate the human voice signal from the external audio signal.

[0045] In one possible implementation, when the headphones detect that the external audio signal contains a human voice signal, the human voice signal is separated from the external audio signal.

[0046] In another possible implementation, when the headphones detect that the external audio signal does not contain human voice signals, they stop separating human voices from the external audio signal in order to reduce the power consumption of the headphones.

[0047] Optionally, the headphones are equipped with an NPU, which performs voice separation.

[0048] Optionally, the human voice signal can be a human speaking voice signal, a speaking voice signal in a broadcast, etc., and this application embodiment does not limit it.

[0049] Step 430: Mix the separated human voice signal and the noise-reduced signal to obtain a mixed signal. The noise-reduced signal is used for active noise reduction.

[0050] In one possible implementation, when the headphones activate noise cancellation mode, the noise cancellation component inside the headphones generates a noise cancellation signal. This noise cancellation signal is an inverse sound wave signal with the opposite phase and the same amplitude as the external noise signal. By neutralizing the external noise signal collected by the headphone microphone with the noise cancellation signal, the external noise signal is canceled out, thereby achieving the active noise cancellation function of the headphones. However, when users wear headphones, they may need to hear announcements or conversations in certain scenarios. Therefore, without removing the headphones, the headphones also need to separate the human voice signal from the external audio signal and transmit it to the user's ear. In this case, there are two signals inside the headphones: one is the separated human voice signal, and the other is the noise cancellation signal. To convert multiple signals into one signal, the headphones are equipped with a mixing circuit. The headphones transmit the separated human voice signal and the noise cancellation signal to the mixing circuit, which converts the two signals into a single mixed signal.

[0051] Step 440: Drive the speaker to produce sound based on the mixed signal.

[0052] The speakers in the headphones convert the mixed electrical signals into sound signals, which in turn produce sound. Therefore, in noisy environments such as airports and subways, users wearing headphones can enjoy noise cancellation while still being able to hear other people talking or announcements on the near-field side of the headphones without removing them. Compared to related technologies that transmit all external sound signals to the ear, the transmission effect is significantly improved.

[0053] In summary, in this embodiment, the headphones first perform voice recognition on the acquired external audio signal. If a voice signal is detected within the external audio signal, it is then separated from the external audio signal. The separated voice signal is then mixed with a noise-canceling signal to generate a single signal that drives the speaker, thus achieving the headphone's pass-through function. In this embodiment, the headphones only pass through the voice signal, allowing users to enjoy noise cancellation while still easily hearing surrounding voices, improving the headphone's pass-through effect. Furthermore, by first recognizing the voice signal from the external audio signal and then separating it only after detection, the headphones' power consumption is reduced.

[0054] In one possible implementation, the headphones first use a low-power VAD classifier to identify human voice signals in the external audio signal. When the headphones detect that the external audio signal contains human voice signals, they then use a high-power human voice separation network to separate the human voice signals from the external audio signal, reducing the frequency of use of the human voice separation network and thus reducing the power consumption of the headphones. Please refer to [reference needed]. Figure 5 The diagram illustrates a flowchart of a voice pass-through method provided in another exemplary embodiment of this application, the method comprising:

[0055] Step 510: Extract features from the collected external audio signals to obtain audio features.

[0056] In one possible implementation, the headphones extract features from the external audio signal to obtain the audio features of the external audio signal.

[0057] Optionally, the audio features can be energy features, frequency domain features, cepstral features, harmonic features, or long-time features; however, this application does not limit these features.

[0058] Different methods exist for extracting various audio features. Optionally, the energy features of the external audio signal can be obtained based on its signal strength; the frequency domain features can be obtained by performing STFT (Short-Time Fourier Transform) on the external audio signal; the cepstral features of the external audio signal can be obtained by performing cepstral analysis, or the MFCC (Mel Frequency Cepstral Coefficients) of the external audio signal can be used as its cepstral features. Since speech contains a fundamental frequency and multiple harmonic frequencies, harmonic features exist even in strong noise environments. Therefore, the fundamental frequency can be found using autocorrelation methods, thereby determining the harmonic features of the external audio signal. Speech is a non-stationary signal, while most everyday noise is a stationary signal. Given that long-term features can analyze the non-stationarity of speech, long-term features of the external audio signal can be extracted.

[0059] Step 520: Classify the audio features using the VAD classifier to obtain the classification result. The classification result is used to characterize the signal type of the signal contained in the external audio signal.

[0060] In one possible implementation, the headphones incorporate VAD (Voice Amplifier) ​​hardware. The headphones input the audio features of extracted external audio signals into the VAD hardware, and a trained VAD classifier classifies these audio features to obtain a classification result. This classification result can be one of three types: spoken voice signals, ambient noise signals, or musical vocal signals. Therefore, this VAD classifier can identify spoken voice signals and musical vocal signals contained within external audio signals, improving the headphone's transmission performance.

[0061] In one possible implementation, the VAD classifier is trained based on sample audio signals that include sample signal type labels.

[0062] Optionally, the sample audio signal can be obtained by mixing at least two of the following: sample spoken voice signal, sample ambient noise signal, and sample musical voice signal. For example, the sample audio signal can be a mixture of sample spoken voice signal and sample ambient noise signal, a mixture of sample spoken voice signal and sample musical voice signal, or a combination of sample spoken voice signal, sample ambient noise signal, and sample musical voice signal, etc.

[0063] Regarding the training process of the VAD classifier, an example is as follows: Figure 6 As shown, the sample audio signal is a mixture of three signals: a spoken voice signal, an ambient noise signal, and a musical voice signal. First, different signal types in the sample audio signal are automatically or manually labeled; for example, a spoken voice signal is labeled as 0, and a musical voice signal is labeled as 1. Audio features are then extracted from the sample audio signal. These features can be energy features, frequency domain features, cepstral features, harmonic features, etc. The audio features are then input into the network model for training, resulting in a trained network model that can distinguish between spoken and musical voice signals.

[0064] The audio features of the sample audio signal are related to the network model, and different network models correspond to different audio signal features. Optionally, the network model can be a GMM (Gaussian Mixed Model), an SVM (Support Vector Machine) model, or a DNN (Deep Neural Networks) model, etc., and this application does not limit it.

[0065] During the model testing phase, audio features of the external audio signal are manually extracted and input into the trained network model to obtain classification results. Post-processing is then performed on the classification results to determine whether they belong to a speaking voice signal. The purpose of post-processing is to more accurately determine whether the classification result is a speaking voice signal. Post-processing can include removing residual echoes and background noise, or controlling clarity, etc., and this embodiment does not limit the specific post-processing steps.

[0066] Step 530: If the classification result indicates that the external audio signal contains a human voice signal, the human voice signal is separated from the external audio signal by a human voice separation network.

[0067] In one possible implementation, if the classification result indicates that the external audio signal contains a speaking voice signal, the speaking voice signal is separated from the external audio signal by a voice separation network.

[0068] In this embodiment of the application, in order to reduce the power consumption of the headphones and extend the standby time of the headphones, a lightweight voice separation network is selected.

[0069] Optionally, the voice separation network can be U-net, etc., and this application embodiment does not limit it.

[0070] Step 540: Mix the separated human voice signal and the noise-reduced signal to obtain a mixed signal. The noise-reduced signal is used for active noise reduction.

[0071] Please refer to step 430 for the implementation method of this step; this application embodiment will not repeat the details.

[0072] Step 550: Drive the speaker to produce sound based on the mixed signal.

[0073] Please refer to step 440 for the implementation method of this step; this application embodiment will not repeat the details.

[0074] In this embodiment, when the headphones identify that the external audio signal contains a human voice signal through a low-power VAD classifier, they then separate the human voice signal from the external audio signal through a high-power human voice separation network, thus avoiding the increase in power consumption of the headphones caused by directly separating the human voice signal from the external audio signal through the human voice separation network.

[0075] Given the relatively small battery capacity of headphones, in order to separate the human voice signal from the external audio signal with low power consumption and high speed, a lightweight human voice separation network, such as U-net, is selected for human voice separation in this embodiment. The human voice separation method provided in this embodiment is described below; please refer to [link / reference]. Figure 7 The diagram illustrates a flowchart of a voice separation method provided in an exemplary embodiment of this application, the method comprising:

[0076] Step 710: Perform time-frequency transformation on the external audio signal to obtain the amplitude spectrum and phase spectrum of the external audio signal.

[0077] In one possible implementation, the headphones include an NPU (Network Processing Unit) for performing voice separation. After the headphones' microphone picks up an external audio signal, it stores the signal in a buffer. When the headphones need to perform voice separation on the external audio signal using the NPU, the NPU reads the signal from the buffer and performs time-frequency transformation on the external audio signal using an STFT (Simultaneous Transformation Theory).

[0078] For example, the headphones collect an external audio signal x through the microphone, and perform a time-frequency transformation on the external audio signal using an STFT to obtain the amplitude spectrum X and phase spectrum Y of the external audio signal:

[0079] X = abs(STFT(x))

[0080] Y = angle(STFT(x))

[0081] Step 720: Perform human voice probability prediction on the amplitude spectrum using a human voice separation network to obtain the human voice probability matrix.

[0082] Furthermore, the headphones input the amplitude spectrum of the external audio signal into the voice separation network to obtain the voice probability matrix. This voice separation network is a pre-trained voice separation network.

[0083] For example, regarding the training process of the human voice separation network, the sample human voice signal s is first processed using STFT. i Mixed acoustic signal x with sample i Performing time-frequency transformation, the amplitude spectra of the sample human voice signal and the sample mixed voice signal are obtained as follows:

[0084] X i =abs(STFT(x) i ))

[0085] S i =abs(STFT(s) i ))

[0086] Among them, X i S represents the amplitude spectrum of the sample mixed acoustic signal. i This represents the amplitude spectrum of the sample human voice signal. It should be noted that the sample human voice signal is a pure human voice signal, while the sample mixed sound signal is a human voice signal mixed with ambient sound. The ambient sound can be environmental noise, music, etc., and this application embodiment does not limit this.

[0087] Furthermore, the amplitude spectrum X of the sample mixed acoustic signal is... i Input Voice Separation Network (Net) i In the process, the feature map m of the human voice probability is obtained. i for:

[0088] m i =Net i (X i )

[0089] Furthermore, the feature map m of human voice probability i Amplitude spectrum X of the mixed acoustic signal with the sample i Multiplication to extract the amplitude spectrum of the human voice signal for:

[0090]

[0091] The loss function L during the training process is:

[0092]

[0093] In one possible implementation, the human voice separation network Net is obtained by performing gradient backpropagation on the loss function L. i The weight update values ​​are used to update the voice separation network Net. i The weights in the network are updated so that the loss function L gradually converges to the human voice separation network Net. i The optimal value is obtained, and then the trained human voice separation network Net is obtained.

[0094] Considering that headphones can accurately separate human voice signals from external audio signals, improving voice transmission, the separation accuracy metric is used to guide the optimization of the network parameters, specifically the weights, of the aforementioned voice separation network during training. Furthermore, according to human auditory psychology models, when the voice transmission delay exceeds 20ms, the human ear can perceive the delay, thus affecting the user experience. Therefore, during the training of the voice separation network, the separation speed metric is also used to guide the optimization of the network architecture, requiring the voice separation network model to have a fast computation speed.

[0095] Optionally, the voice separation network can be a convolutional neural network, a recurrent neural network, or a combination of convolutional neural networks and recurrent neural networks, etc., and the embodiments of this application do not limit this.

[0096] The trained voice separation network model is pre-configured in the NPU of the headphones, and the headphones use the NPU to separate voices from external audio signals.

[0097] For example, by inputting the amplitude spectrum X of the aforementioned external audio signal into the trained human voice separation network Net, the human voice probability matrix m is obtained as follows:

[0098] m = Net(X)

[0099] Net is the trained human voice separation network.

[0100] In one possible implementation, in order to reduce the power consumption of the headphone voice pass-through system and improve the calculation speed of the voice separation network while avoiding delay, the voice separation network adopts U-net. The U-net is used to predict the voice probability of the amplitude spectrum X to obtain the voice probability matrix.

[0101] First, the amplitude spectrum is feature extracted through the n-layer feature extraction layers of U-net to obtain the downsampled feature maps output by each feature extraction layer.

[0102] n is selected based on the separation speed index.

[0103] Secondly, the upsampled feature map and the downsampled feature map are fused by the n-layer feature fusion layer of U-net to obtain the target feature map. The upsampled feature map is obtained by upsampling the downsampled feature map by the feature fusion layer.

[0104] Optionally, the fusion method can be concatenation or addition, and the embodiments of this application are not limited to this.

[0105] Finally, activation processing is performed on the target feature map to obtain the human voice probability matrix.

[0106] The target feature map is processed by an activation function to obtain a human voice probability matrix, which ranges from 0 to 1.

[0107] Alternatively, commonly used activation functions may be sigmoid, tanh, etc., and this application does not limit them.

[0108] For example, such as Figure 8 As shown, it illustrates a schematic diagram of the process of obtaining a human voice probability matrix using U-net according to an exemplary embodiment of this application.

[0109] The amplitude spectrum of the external audio signal is input into the U-net. After the first convolutional layer, 2×2 max pooling is used to downsample the feature map in both the time and frequency domains through three consecutive downsampling convolutional layers until the bottleneck layer is reached. Subsequently, 3×3 convolution and ReLU are used to upsample the feature map through the same number of upsampling convolutional layers. To recover the details lost during downsampling, the feature map of each downsampling layer is connected to the feature map of the corresponding upsampling layer; that is, the upsampling and downsampling feature maps are fused together. The fusion method can be concatenation or superposition. Finally, 1×1 convolution and ReLU are used to pass the signal through an output layer to obtain the amplitude spectrum mask of the output audio, i.e., the target feature map. The target feature map is activated using the sigmoid function to obtain the human voice probability matrix.

[0110] Step 730: Generate the human voice amplitude spectrum of the human voice signal based on the human voice probability matrix and amplitude spectrum.

[0111] For example, the human voice probability matrix m is multiplied by the amplitude spectrum X of the external audio signal x to obtain the human voice amplitude spectrum.

[0112]

[0113] Step 740: Perform inverse time-frequency transformation based on the amplitude spectrum and phase spectrum of the human voice to obtain the human voice signal.

[0114] For example, the amplitude spectrum of the human voice and the phase spectrum of the external audio signal are combined to form the complex spectrum of the output human voice signal. The complex spectrum is subjected to inverse time-frequency transformation using ISTFT (Inverse Short-Time Fourier Transform) to obtain the human voice signal S:

[0115]

[0116] For example, such as Figure 9 The diagram illustrates a voice separation process provided in an exemplary embodiment of this application. The headset includes a VAD hardware 910, an NPU 920, and a buffer 930. The headset stores external audio signals collected by the microphone in the buffer 930 and simultaneously transmits these signals to the VAD hardware 910. The VAD classifier in the VAD hardware 910 performs voice recognition on the external audio signals. When it detects that the external audio signals contain speaking voice signals, it triggers the NPU 920. The NPU 920 then uses the external audio signals stored in the buffer 930 to perform voice separation, thereby obtaining the speaking voice signal.

[0117] In this embodiment, a lightweight voice separation network, such as U-net, is used for voice separation, which reduces the power consumption of the headphones while ensuring the speed of voice separation.

[0118] Please refer to Figure 10 This illustration shows a structural block diagram of a voice transmission device provided in an exemplary embodiment of this application. The device is used for headphones and includes:

[0119] The voice recognition module 1001 is used to recognize human voices from the collected external audio signals;

[0120] The separation module 1002 is used to separate the human voice signal from the external audio signal when the external audio signal is identified to contain a human voice signal;

[0121] The mixing module 1003 is used to mix the separated human voice signal and noise reduction signal to obtain a mixed signal, and the noise reduction signal is used for active noise reduction.

[0122] The driver module 1004 is used to drive the speaker to produce sound based on the mixed signal.

[0123] Optionally, the voice recognition module 1001 is used for:

[0124] The acquired external audio signals are subjected to feature extraction to obtain audio features;

[0125] The audio features are classified using a VAD classifier to obtain a classification result, which is used to characterize the signal type of the signal contained in the external audio signal.

[0126] The separation module 1002 includes:

[0127] A separation unit is used to separate the human voice signal from the external audio signal through a human voice separation network when the classification result indicates that the external audio signal contains the human voice signal.

[0128] The power consumption of classification by the VAD classifier is lower than that of voice separation by the voice separation network.

[0129] Optionally, the signal type includes at least one of a speaker's voice signal, an ambient noise signal, and a musical voice signal;

[0130] The separation unit is used for:

[0131] If the classification result indicates that the external audio signal contains the speaking voice signal, the speaking voice signal is separated from the external audio signal by the voice separation network.

[0132] Optionally, the VAD classifier is trained based on sample audio signals containing sample signal type labels, wherein the sample audio signals are obtained by mixing at least two of the following: sample speaking voice signals, sample ambient noise signals, and sample musical voice signals.

[0133] Optionally, the separation unit is used for:

[0134] The external audio signal is subjected to time-frequency transformation to obtain the amplitude spectrum and phase spectrum of the external audio signal;

[0135] The human voice probability matrix is ​​obtained by performing human voice probability prediction on the amplitude spectrum through the human voice separation network.

[0136] The human voice amplitude spectrum of the human voice signal is generated based on the human voice probability matrix and the amplitude spectrum.

[0137] The human voice signal is obtained by performing an inverse time-frequency transformation based on the amplitude spectrum and phase spectrum of the human voice.

[0138] Optionally, the voice separation network adopts U-net;

[0139] Optionally, the separation unit is further configured to:

[0140] The amplitude spectrum is feature extracted through the n-layer feature extraction layer of the human voice separation network to obtain the downsampled feature map output by each feature extraction layer;

[0141] The upsampled feature map and the downsampled feature map are fused by the n-layer feature fusion layer of the human voice separation network to obtain the target feature map. The upsampled feature map is obtained by upsampling the downsampled feature map by the feature fusion layer.

[0142] The target feature map is activated to obtain the human voice probability matrix.

[0143] Optionally, the training metrics used during the training of the voice separation network include a separation accuracy metric and a separation speed metric, wherein the separation accuracy metric is used to guide the optimization of the network parameters of the voice separation network, and the separation speed metric is used to guide the optimization of the network architecture of the voice separation network.

[0144] Optionally, the device further includes:

[0145] The stop module is used to stop the human voice separation of the external audio signal when it is determined that the external audio signal does not contain a human voice signal.

[0146] Optionally, the headset is equipped with VAD hardware and NPU, wherein voice recognition is performed by the VAD and voice separation is performed by the NPU; or, the headset is equipped with an NPU, wherein voice recognition and voice separation are performed by the NPU.

[0147] In summary, in this embodiment, the headphones first perform voice recognition on the acquired external audio signal. If a voice signal is detected within the external audio signal, it is then separated from the external audio signal. The separated voice signal is then mixed with a noise-canceling signal to generate a single signal that drives the speaker, thus achieving the headphone's pass-through function. In this embodiment, the headphones only pass through the voice signal, allowing users to enjoy noise cancellation while still easily hearing surrounding voices, improving the headphone's pass-through effect. Furthermore, by first recognizing the voice signal from the external audio signal and then separating it only after detection, the headphones' power consumption is reduced.

[0148] It should be noted that the apparatus provided in the above embodiments is only illustrative of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0149] Please refer to Figure 11The diagram illustrates a structural block diagram of an earphone 1100 provided in an exemplary embodiment of this application. The earphone in this application may include one or more of the following components: a processor 1110, a memory 1120, a microphone 1130, a speaker 1140, and VAD hardware 1150, wherein the processor 1110 is electrically connected to the memory 1120, the microphone 1130, the speaker 1140, and the VAD hardware 1150, respectively.

[0150] The processor 1110 may include an NPU for voice separation and an MCU (Micro Controller Unit) for implementing other functions of the headset. The processor 1110 connects to various parts inside the headset 1100 using various interfaces and lines, and performs various functions of the headset 1100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in memory 1120 and calling data stored in memory 1120.

[0151] The memory 1120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1120 may include non-transitory computer-readable storage medium. The memory 1120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing at least one function (such as a sound playback function), and the data storage area may also store data such as external audio signals collected by the microphone 1130.

[0152] Microphone 1130 is a transducer that converts sound signals into electrical signals and is used to collect external audio signals.

[0153] The loudspeaker 1140 is a transducer that converts electrical signals into sound signals, used to play out the separated human voice signal.

[0154] The VAD hardware 1150 is used to identify human voice signals contained in external audio signals.

[0155] In addition, those skilled in the art will understand that the structure shown in the above figures does not constitute a limitation on the headphone 1100. The headphone 1100 may include more or fewer components than shown, or combine certain components or arrange different components. For example, the headphone 1100 may also include components such as sensors, audio circuits, control circuits, and power supplies, which will not be described in detail in this application.

[0156] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the voice transmission method provided in the above embodiments.

[0157] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an earphone reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the earphone to perform the voice pass-through method provided in various alternative implementations of the above aspect.

[0158] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0159] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A human voice transparent transmission method, characterized by, The method is used for a headset, and the method comprises: extracting features of the collected external audio signal to obtain audio features; classifying the audio features by a VAD classifier to obtain a classification result, the classification result being used to represent a signal type of a signal contained in the external audio signal, and a power consumption of the classification by the VAD classifier being lower than a power consumption of human voice separation by a human voice separation network, the human voice separation network adopting a U-net; in a case where the classification result indicates that the external audio signal contains a human voice signal, performing time-frequency conversion on the external audio signal to obtain an amplitude spectrum and a phase spectrum of the external audio signal; extracting features of the amplitude spectrum by n layers of feature extraction layers of the human voice separation network to obtain down-sampling feature maps output by each feature extraction layer; performing feature fusion on up-sampling feature maps and the down-sampling feature maps by n layers of feature fusion layers of the human voice separation network to obtain a target feature map, the up-sampling feature maps being obtained by up-sampling the down-sampling feature maps by the feature fusion layers; performing activation processing on the target feature map to obtain a human voice probability matrix; generating a human voice amplitude spectrum of the human voice signal based on the human voice probability matrix and the amplitude spectrum; and performing inverse time-frequency conversion based on the human voice amplitude spectrum and the phase spectrum to obtain the human voice signal; performing mixing processing on the separated human voice signal and a noise reduction signal to obtain a mixed signal, the noise reduction signal being used for active noise reduction; driving a loudspeaker to emit sound based on the mixed signal.

2. The method of claim 1, wherein, The signal type comprises at least one of a speaker voice signal, an environmental noise signal, and a music human voice signal. The method further comprises: in a case where the classification result indicates that the external audio signal contains the speaker voice signal, separating the human voice signal from the external audio signal by the human voice separation network.

3. The method of claim 2, wherein, The VAD classifier is trained based on sample audio signals containing sample signal type labels, the sample audio signals being obtained by mixing at least two of a sample speaker voice signal, a sample environmental noise signal, and a sample music human voice signal.

4. The method of claim 1, wherein, Training indicators used in a training process of the human voice separation network comprise a separation accuracy indicator and a separation speed indicator, wherein the separation accuracy indicator is used to guide optimization of network parameters of the human voice separation network, and the separation speed indicator is used to guide optimization of a network architecture of the human voice separation network.

5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: in a case where it is identified that the external audio signal does not contain a human voice signal, stopping human voice separation on the external audio signal.

6. The method of any one of claims 1 to 4, wherein the headset is provided with a VAD hardware and an NPU, wherein human voice recognition is performed by the VAD, and human voice separation is performed by the NPU, or the headset is provided with an NPU, wherein human voice recognition and human voice separation are performed by the NPU.

7. A human voice see-through device, characterized by, The device is used for a headset, and the device comprises: a human voice recognition module configured to extract features of a collected external audio signal to obtain audio features; The audio features are classified by a VAD classifier to obtain a classification result, the classification result is used to represent a signal type of a signal contained in the external audio signal, and power consumption of classification by the VAD classifier is lower than power consumption of human voice separation by a human voice separation network, and the human voice separation network adopts a U-net; The separation module is configured to, in a case where the classification result indicates that the external audio signal contains a human voice signal, perform time-frequency transformation on the external audio signal to obtain an amplitude spectrum and a phase spectrum of the external audio signal; perform feature extraction on the amplitude spectrum by n layers of feature extraction layers of the human voice separation network to obtain down-sampling feature maps output by each feature extraction layer; perform feature fusion on up-sampling feature maps and the down-sampling feature maps by n layers of feature fusion layers of the human voice separation network to obtain a target feature map, the up-sampling feature maps being obtained by up-sampling the down-sampling feature maps by the feature fusion layers; perform activation processing on the target feature map to obtain a human voice probability matrix; generate a human voice amplitude spectrum of the human voice signal based on the human voice probability matrix and the amplitude spectrum; and perform inverse time-frequency transformation based on the human voice amplitude spectrum and the phase spectrum to obtain the human voice signal. The mixing module is configured to perform mixing processing on the separated human voice signal and a noise reduction signal to obtain a mixed signal, the noise reduction signal being used for active noise reduction. The driving module is configured to drive a loudspeaker to emit sound based on the mixed signal.

8. An earphone, characterized by The earphone includes a processor and a memory, the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the human voice transparent transmission method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the human voice transparent transmission method according to any one of claims 1 to 6.

10. A computer program product, characterised in that, The computer program product includes computer instructions, and the computer instructions are executed by the processor to implement the human voice transparent transmission method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Mode control method and device of Bluetooth headset and computer readable storage medium

    CN113573195A

  • Sound processing circuit, electroacoustic device, and sound processing system

    CN214226506U