Target human voice separation system and method based on bone conduction and air conduction
By combining bone conduction and air conduction speech acquisition modules, along with sub-band segmentation and feature extraction modules, the problem of target voice separation in multi-person speaking scenarios in translation headphones has been solved, achieving efficient and accurate speech recognition and translation.
Patent Information
- Application Number
- CN202511469376.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-02-13
AI Technical Summary
Existing translation headsets struggle to accurately capture and translate the speech of specific speakers in multi-person conversation scenarios, limited by computing resources and storage space, and their highly complex algorithms are difficult to run efficiently in embedded devices.
By employing the synergistic action of bone conduction and air conduction speech acquisition modules, combined with sub-band segmentation and feature extraction modules, the target human voice signal is separated through a segmentation strategy that is refined in low frequencies and sparse in high frequencies. Furthermore, the bone conduction signal is used to guide the speech separation, reducing the computational burden and improving the accuracy of speech recognition and translation.
It significantly improves the accuracy of speech recognition and translation, reduces the computational burden, adapts to the computing load of embedded devices, and meets the requirements of low latency and high real-time performance.
Smart Images

Figure CN121528236A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing, and more specifically, to a target human voice separation system and method based on bone conduction and air conduction. Background Technology
[0002] In complex scenarios such as multi-person meetings or the United Nations General Assembly, multiple people often speak simultaneously. In such situations, the main challenge for traditional simultaneous interpretation equipment is accurately capturing and translating the speech of a specific speaker. Because air microphones capture a mixture of multiple human voices, the backend speech recognition and translation functions struggle to function properly. Furthermore, extracting the speech of a specific speaker from multiple voices requires addressing numerous interference factors in the complex environment, such as noise and reverberation, which often results in unsatisfactory speech extraction results, further increasing the difficulty of technical implementation.
[0003] Currently, translation headsets are favored by many due to their convenience and real-time translation capabilities. However, existing methods for extracting the target speaker's voice often require substantial computational resources and storage space. The high computational demands, similar to solving the "cocktail party" problem, significantly hinder the practical application of this technology in embedded devices. This conflict with the limited computing power and RAM capacity of the chips integrated into the headsets. Furthermore, the low latency and high real-time requirements of simultaneous interpretation devices make it difficult for highly complex algorithms to run efficiently in practical applications. Therefore, designing a low-complexity yet highly effective target speaker extraction method that can withstand the computational load of embedded devices has become a pressing technical challenge. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings and deficiencies of existing technologies and provide a target voice separation system and method based on bone conduction and air conduction. The system of this invention segments the frequency domain signal of the mixed audio signal into several sub-band frequency domain mixed signals through a sub-band segmentation module, and adopts a segmentation strategy of "fine low-frequency, sparse high-frequency," that is, the signal is subdivided into multiple frequency bands, especially with fine segmentation in the low-frequency part, while the high-frequency part uses coarser segmentation. This strategy allows the model to retain important low-frequency speech information while significantly reducing the computational burden. Furthermore, this invention is the first to introduce the synergistic effect of the bone conduction speech acquisition module and the air conduction speech acquisition module into the target voice separation system. Based on the temporal characteristics of the bone conduction speech signal, the voice separation module is guided to extract the wearer's target speaker from the speech signal, thereby significantly improving the accuracy of speech recognition and translation.
[0005] To achieve the above objectives, the present invention provides a target human voice separation system based on bone conduction and air conduction, characterized in that the system comprises: Bone conduction speech acquisition module, used to acquire the wearer's bone conduction speech signal; The air conduction voice acquisition module is used to acquire the wearer's air conduction voice signal, as well as other voice signals and background noise; An air conduction input signal processing module is used to process the mixed audio signal acquired by the air conduction voice acquisition module, and convert the time domain signal of the mixed audio signal into a frequency domain signal through short-time Fourier transform; The feature extraction module is used to extract spectral features from the frequency domain signal of the mixed audio signal; The voice separation module inputs the spectral features of the mixed audio signal extracted by the feature extraction module into the voice separation module, separates the individual voice spectral signals of all voices, and performs a short-time inverse Fourier transform on each individual voice spectral signal to finally obtain multiple clean individual voice time-domain signals. The channel selection module is used to output multiple clean single human voice time-domain signals separated by the human voice separation module through multiple signal channels, and to evaluate the confidence of each clean single human voice time-domain signal with the wearer's bone conduction speech signal. Finally, the single human voice time-domain signal with the highest confidence evaluation score is selected as the target human voice audio signal.
[0006] According to one embodiment of the present invention, the system further includes a sub-band segmentation module for segmenting the frequency domain signal of the mixed audio signal into a plurality of sub-band frequency domain mixed signals; Before the feature extraction module extracts spectral features from the frequency domain signal of the mixed audio signal, the frequency domain signal of the mixed audio signal is first divided into several sub-band frequency domain mixed signals, and then the feature extraction module extracts sub-band frequency domain features from each sub-band frequency domain mixed signal. All obtained sub-band frequency domain features are input into the human voice separation module for single human voice spectrum signal separation. Each sub-band frequency domain mixed signal is separated into sub-band single human voice spectrum signals of all human voices. Then, the sub-band single human voice spectrum signals belonging to the same human voice in all sub-band frequency domain mixed signals are merged to finally obtain all complete single human voice spectrum signals. Each single human voice spectrum signal is then subjected to short-time inverse Fourier transform to obtain all clean single human voice time domain signals.
[0007] According to one embodiment of the present invention, the subband segmentation module adopts a segmentation strategy of "fine low frequency and sparse high frequency" when segmenting into several subband frequency domain mixed signals; the segmentation strategy of "fine low frequency and sparse high frequency" means that the low frequency part of the frequency domain signal of the mixed audio signal is finely divided with a higher resolution, while the high frequency part is coarsely divided with a lower resolution.
[0008] According to one embodiment of the present invention, the bandwidth of the frequency domain signal of the mixed audio signal is (0, 8kHz), and the bandwidth (0, 8kHz) is divided into 54 sub-bands, as follows: 0~1kHz: divided into 20 segments, each 50Hz, for a total of 20 sub-bands; 1k~3kHz: divided into 20 bands, each 100Hz, for a total of 20 sub-bands; 3k~6kHz: divided into 10 segments, each 300Hz, for a total of 10 sub-bands; 6kHz~8kHz: Divided into 4 segments, each 500Hz, for a total of 4 sub-bands. The frequency domain signal of the mixed audio signal is divided into 54 sub-band frequency domain mixed signals after passing through 54 sub-bands.
[0009] According to one embodiment of the present invention, the channel selection module further includes a synchronization processing unit, which is used to compensate for the actual system delay, such that the bone conduction speech signal and the air conduction speech signal are kept in time aligned before each clean single human voice time-domain signal output is evaluated with the wearer's bone conduction speech signal for confidence.
[0010] According to one embodiment of the present invention, the system further includes a bone conduction input signal processing module, which is used to process the bone conduction speech signal acquired by the bone conduction speech acquisition module, and convert the time domain signal of the bone conduction speech signal into a bone conduction frequency domain signal through short-time Fourier transform; When the voice separation module separates the air-conducted mixed audio signal into multiple clean single voice spectrum signals, each single voice spectrum signal is output through multiple signal channels of the channel selection module. Each clean single voice spectrum signal is then compared with the wearer's bone conduction frequency domain signal for confidence evaluation. The single voice spectrum signal with the highest confidence evaluation score is selected. Finally, a short-time inverse Fourier transform is performed on the selected single voice spectrum signal to obtain the audio signal of the target voice.
[0011] Another objective of this invention is to provide a target human voice separation method based on bone conduction and air conduction, the target human voice separation method comprising the following steps: S1. The bone conduction speech acquisition module is used to acquire the wearer's bone conduction speech signal; the air conduction speech acquisition module is used to acquire the wearer's air conduction speech signal, other human voices around the wearer, and background noise. S2. The mixed audio signal acquired by the air conduction speech acquisition module in step S1 is processed into frames according to a predetermined frame length and frame shift. The sampling rate of the mixed audio signal is 16KHz. A fast Fourier transform is performed on each frame of the mixed audio signal to convert the time domain signal of the mixed audio signal into a frequency domain signal. S3. The frequency domain signal of the mixed audio signal is divided into several sub-band frequency domain mixed signals by the sub-band segmentation module, and then the sub-band frequency domain features are extracted for each sub-band frequency domain mixed signal by the feature extraction module. S4. Input all the obtained sub-band frequency domain features into the human voice separation module to separate the single human voice spectrum signal. Each sub-band frequency domain mixed signal will be separated into the sub-band single human voice spectrum signal of all human voices. Then, the sub-band single human voice spectrum signals belonging to the same human voice in all sub-band frequency domain mixed signals will be merged to finally obtain all complete single human voice spectrum signals. Then, each single human voice spectrum signal will be subjected to short-time inverse Fourier transform to obtain all clean single human voice time domain signals. S5. All clean single human voice time-domain signals separated by the channel selection module in step S4 are output through multiple signal channels, and each clean single human voice time-domain signal is compared with the wearer's bone conduction speech signal for confidence evaluation. Finally, the single human voice time-domain signal with the highest confidence evaluation score is selected as the target human voice audio signal.
[0012] According to an embodiment of the present invention, in step S5, before performing confidence evaluation on each clean single human voice time-domain signal output and the wearer's bone conduction speech signal respectively, the actual system delay is compensated by the synchronization processing unit to ensure that the bone conduction speech signal and the air conduction speech signal are aligned in time. Then, each clean single human voice time-domain signal output is evaluated on confidence with the wearer's bone conduction speech signal respectively, and finally the single human voice time-domain signal with the highest confidence evaluation score is selected as the audio signal of the target human voice.
[0013] According to an embodiment of the present invention, step S2 further includes performing a fast Fourier transform on each frame of bone conduction audio signal to convert the time-domain signal of the bone conduction audio signal into a bone conduction frequency-domain signal; when all complete single human voice spectrum signals are finally obtained in step 4, each complete single human voice spectrum signal is output through multiple signal channels of the channel selection module, and each complete single human voice spectrum signal is evaluated with the bone conduction frequency-domain signal to obtain the confidence level, the single human voice spectrum signal with the highest confidence level is selected, and finally a short-time inverse Fourier transform is performed on the selected single human voice spectrum signal to obtain the audio signal of the target human voice.
[0014] Compared with existing technologies, this invention has the following advantages: The system of this invention segments the frequency domain signal of the mixed audio signal into several sub-band frequency domain mixed signals through a sub-band segmentation module, and adopts a segmentation strategy of "fine low-frequency, sparse high-frequency," meaning the signal is subdivided into multiple frequency bands, especially with fine segmentation in the low-frequency part, while the high-frequency part uses coarser segmentation. This strategy allows the model to retain important low-frequency speech information while significantly reducing the computational burden. Furthermore, this invention, for the first time, introduces the synergistic effect of the bone conduction speech acquisition module and the air conduction speech acquisition module into the target human voice separation system. Based on the temporal characteristics of the bone conduction speech signal, the human voice separation module is guided to extract the wearer's target speaker from the speech signal, thereby significantly improving the accuracy of speech recognition and translation. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the target human voice separation system of the present invention; Figure 2 This is another principle block diagram of the target human voice separation system of the present invention; Figure 3 This is a schematic diagram of an embodiment of the present invention; Figure 4 This is a schematic diagram of another embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the steps of the target human voice separation method of the present invention. Detailed Implementation
[0016] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.
[0017] like Figure 1 The diagram shown is a schematic block diagram of the target human voice separation system based on bone conduction and air conduction provided by the present invention. The system includes: Bone conduction speech acquisition module, used to acquire the wearer's bone conduction speech signal; The air conduction voice acquisition module is used to acquire the air conduction voice signal of the wearer, and also to acquire other voice signals and background noise. In this embodiment of the invention, for ease of description, all signals acquired by the air conduction voice acquisition module are collectively referred to as mixed audio signals. The air conduction input signal processing module processes the mixed audio signal acquired by the air conduction speech acquisition module. This processing involves: first, dividing the continuous air conduction mixed audio signal into frames according to a predetermined frame length (32ms) and frame shift (16ms), with a sampling rate of 16kHz. Specifically, the entire mixed audio signal is divided into multiple overlapping short-time frames, each 32 milliseconds long with a frame shift of 16 milliseconds. This signal processing is based on the assumption that the speech signal is stationary over a short period, effectively reflecting local temporal changes and providing an appropriate time window for subsequent frequency domain analysis, thereby improving the accuracy of signal feature capture. The processed mixed audio signal is then converted from its time domain signal to its frequency domain signal using a short-time Fourier transform. The feature extraction module is used to extract spectral features from the frequency domain signal of the mixed audio signal; The voice separation module takes the spectral features of the mixed audio signal extracted by the feature extraction module and inputs them into the voice separation module to separate the individual voice spectral signals of all voices. Then, it performs a short-time inverse Fourier transform on each individual voice spectral signal to finally obtain multiple clean individual voice time-domain signals. The channel selection module is used to output multiple clean, single-voice time-domain signals separated by the voice separation module through multiple signal channels. Each output clean, single-voice time-domain signal is compared with the wearer's bone conduction speech signal for confidence assessment. Finally, the single-voice time-domain signal with the highest confidence assessment score is selected as the target voice audio signal. In this embodiment, the confidence assessment method for the two signals is a conventional technique in the art. Since this confidence assessment method is not the core innovation of this invention, it is not described in detail here. In this embodiment, the target voice separation system also includes other necessary modules, such as a processor module, a power supply module, and a storage module. These are essential functional modules for speech signal processing and are well-known to those skilled in the art, so they are not described in detail in the schematic diagram and specification.
[0018] In this embodiment of the invention, to improve the efficiency and anti-interference capability of the feature extraction module in extracting spectral features from the frequency domain signal of mixed audio signals, a sub-band segmentation strategy is introduced after performing a short-time Fourier transform on the air-conducted speech signal, such as... Figure 2 As shown, the target human voice separation system of the present invention also includes a sub-band segmentation module, which is used to segment the frequency domain signal of the mixed audio signal into several sub-band frequency domain mixed signals; Before the feature extraction module extracts spectral features from the frequency domain signal of the mixed audio signal, the frequency domain signal of the mixed audio signal is first divided into several sub-band frequency domain mixed signals. Then, the feature extraction module extracts sub-band frequency domain features from each sub-band frequency domain mixed signal. All obtained sub-band frequency domain features are input into the voice separation module to separate the single voice spectrum signal. Each sub-band frequency domain mixed signal is separated into the sub-band single voice spectrum signal of all voices. Then, the sub-band single voice spectrum signals belonging to the same voice in all sub-band frequency domain mixed signals are merged to finally obtain all complete single voice spectrum signals. Each single voice spectrum signal is then subjected to short-time inverse Fourier transform to obtain all clean single voice time domain signals.
[0019] In this embodiment of the invention, considering that human voice energy is mainly distributed in the lower frequency range (0~3kHz), and that the high-frequency part (above 3kHz) contributes relatively little to speech intelligibility, this invention adopts a "refined low-frequency, sparse high-frequency" banding strategy. This allows the network to focus more learning resources on the key low-frequency range, achieving low-latency and high-precision speech extraction. Specifically, this "refined low-frequency, sparse high-frequency" segmentation strategy involves: finely dividing the low-frequency part of the mixed audio signal's frequency domain signal with a higher resolution; while the high-frequency part is coarsely divided with a lower resolution. In this embodiment of the invention, the bandwidth of the mixed audio signal's frequency domain signal is (0, 8kHz), and the bandwidth (0, 8kHz) is divided into four intervals and 54 sub-bands, as follows: 0~1kHz: divided into 20 segments, each 50Hz, for a total of 20 sub-bands; 1k~3kHz: divided into 20 bands, each 100Hz, for a total of 20 sub-bands; 3k~6kHz: divided into 10 segments, each 300Hz, for a total of 10 sub-bands; 6kHz~8kHz: Divided into 4 segments, each 500Hz, for a total of 4 sub-bands. The frequency domain signal of the mixed audio signal is divided into 54 sub-bands after passing through 54 sub-bands, such as... Figure 3As shown, the 54 sub-band frequency domain mixed signals are sub-band 1 frequency domain mixed signal, sub-band 2 frequency domain mixed signal, ..., sub-band 54 frequency domain mixed signal. First, the sub-band 1 frequency domain mixed signal, sub-band 2 frequency domain mixed signal, ..., sub-band 54 frequency domain mixed signal are respectively processed by the feature extraction module to extract their respective sub-band frequency domain features, and then their respective sub-band frequency domain features are input to the voice separation module for single voice separation. In this embodiment of the invention, it is assumed that the mixed audio signal contains three voice signals, so the sub-band 1 frequency domain mixed signal is separated into the first voice frequency domain signal, the second voice frequency domain signal, and the third voice frequency domain signal of sub-band 1. Similarly, the sub-band 2 frequency domain mixed signal is separated into the first human voice frequency domain signal, the second human voice frequency domain signal, and the third human voice frequency domain signal of sub-band 2; ...; the sub-band 54 frequency domain mixed signal is separated into the first human voice frequency domain signal, the second human voice frequency domain signal, and the third human voice frequency domain signal of sub-band 54. To recover the complete and clean spectral signals obtained from the first, second, and third voices, the first voice frequency domain signals of sub-band 1, sub-band 2, ..., and sub-band 54 are merged to obtain the complete first voice frequency domain signal; the second voice frequency domain signals of sub-band 1, sub-band 2, ..., and sub-band 54 are merged to obtain the complete second voice frequency domain signal; and the third voice frequency domain signals of sub-band 1, sub-band 2, ..., and sub-band 54 are merged to obtain the complete third voice frequency domain signal. Then, the complete first, second, and third human voice frequency domain signals are subjected to short-time inverse Fourier transform to obtain the complete first, second, and third human voice time domain signals, respectively. Finally, the bone conduction time domain signal is used to evaluate the signal confidence of the complete first, second, and third human voice time domain signals. The human voice time domain signal with the highest evaluation score is determined as the target human voice time domain signal. The target human voice separated by the method of this invention can be used for speech translation, which can improve the accuracy of translation.
[0020] In this embodiment of the invention, since there will inevitably be a time delay when the mixed audio signal obtained by the air conduction voice acquisition module is processed, in order to ensure that the bone conduction signal and the air conduction signal are highly consistent in time, the channel selection module of the present invention also includes a synchronization processing unit. The synchronization processing unit is used to supplement the actual system delay, so that before each clean single human voice time domain signal output is compared with the wearer's bone conduction voice signal for confidence evaluation, the bone conduction voice signal and the air conduction voice signal are kept aligned in time.
[0021] In this embodiment of the invention, the target voice separation system further includes a bone conduction input signal processing module, which is used to process the bone conduction speech signal acquired by the bone conduction speech acquisition module, and convert the time domain signal of the bone conduction speech signal into a bone conduction frequency domain signal through short-time Fourier transform. This step is the same as the air conduction speech signal processing.
[0022] Accordingly, when the voice separation module separates the mixed audio signal conducted through the air into multiple clean single voice spectrum signals, each single voice spectrum signal is output through multiple signal channels of the channel selection module. Each clean single voice spectrum signal is then compared with the wearer's bone conduction frequency domain signal for confidence assessment. The single voice spectrum signal with the highest confidence assessment score is selected. Finally, a short-time inverse Fourier transform is performed on the selected single voice spectrum signal to obtain the audio signal of the target voice, such as... Figure 4 As shown, the obtained complete first human voice frequency domain signal, second human voice frequency domain signal and third human voice frequency domain signal are compared with the bone conduction frequency domain signal for confidence evaluation. The single human voice spectrum signal with the highest confidence evaluation score is selected. Finally, the selected single human voice spectrum signal is subjected to short-time inverse Fourier transform to obtain the audio signal of the target human voice.
[0023] Another object of the present invention is to provide a target human voice separation method based on bone conduction and air conduction, the target human voice separation method comprising the following steps, such as Figure 5 As shown, S1. The bone conduction speech acquisition module is used to acquire the wearer's bone conduction speech signal; the air conduction speech acquisition module is used to acquire the wearer's air conduction speech signal, other human voices around the wearer, and background noise. S2. The mixed audio signal acquired by the air conduction speech acquisition module in step S1 is processed into frames according to a predetermined frame length and frame shift. The sampling rate of the mixed audio signal is 16KHz. A fast Fourier transform is performed on each frame of the mixed audio signal to convert the time domain signal of the mixed audio signal into a frequency domain signal. S3. The frequency domain signal of the mixed audio signal is divided into several sub-band frequency domain mixed signals by the sub-band segmentation module, and then the sub-band frequency domain features are extracted for each sub-band frequency domain mixed signal by the feature extraction module. S4. Input all the obtained sub-band frequency domain features into the human voice separation module to separate the single human voice spectrum signal. Each sub-band frequency domain mixed signal will be separated into the sub-band single human voice spectrum signal of all human voices. Then, the sub-band single human voice spectrum signals belonging to the same human voice in all sub-band frequency domain mixed signals will be merged to finally obtain all complete single human voice spectrum signals. Then, each single human voice spectrum signal will be subjected to short-time inverse Fourier transform to obtain all clean single human voice time domain signals. S5. All clean single human voice time-domain signals separated by the channel selection module in step S4 are output through multiple signal channels. Each clean single human voice time-domain signal is compared with the wearer's bone conduction speech signal for confidence evaluation. Finally, the single human voice time-domain signal with the highest confidence evaluation score is selected as the target human voice audio signal.
[0024] In step S5 of the target human voice separation method of the present invention, before performing confidence evaluation on each clean single human voice time-domain signal output and the wearer's bone conduction speech signal, the actual system delay is compensated by the synchronization processing unit to ensure that the bone conduction speech signal and the air conduction speech signal are aligned in time. Then, each clean single human voice time-domain signal output is evaluated on confidence with the wearer's bone conduction speech signal, and finally the single human voice time-domain signal with the highest confidence evaluation score is selected as the audio signal of the target human voice.
[0025] In this embodiment of the invention, step S2 further includes performing a fast Fourier transform on each frame of bone conduction audio signal to convert the time-domain signal of the bone conduction audio signal into a bone conduction frequency-domain signal; when all complete single human voice spectrum signals are finally obtained in step 4, each complete single human voice spectrum signal is output through multiple signal channels of the channel selection module, and each complete single human voice spectrum signal is evaluated with the bone conduction frequency-domain signal to select the single human voice spectrum signal with the highest confidence evaluation score. Finally, a short-time inverse Fourier transform is performed on the selected single human voice spectrum signal to obtain the audio signal of the target human voice.
[0026] In this embodiment of the invention, the target human voice separation system based on bone conduction and air conduction can be applied to translation headphones. The translation headphones are equipped with bone conduction microphones and air conduction microphones. The system and method of this invention filter out the target human voice and send the filtered voice to a translation engine for recognition and translation, thereby improving translation accuracy. Of course, the system and method of this invention can also be used in other target speech acquisition devices or apparatuses.
[0027] As mentioned repeatedly in this invention, after obtaining the spectral characteristics of the mixed audio signal obtained through air conduction, inputting these spectral characteristics into the voice separation module can separate multiple clean voice signals. Similarly, inputting sub-band frequency domain characteristics into the voice separation module can also separate multiple single voice spectral signals. This voice separation method first requires establishing a voice separation model, training the model, and then inputting it into the voice separation module of this invention. The construction and training of the voice separation model can employ conventional training methods in the field. For example, Chinese Patent Publication No. CN109461447B provides a method for training a two-voice separation model on pages 6 to 9 of its specification. Similarly, this method can also be used to construct and train three or more voice separation models. Therefore, this invention does not provide a detailed description of the voice separation module technology.
[0028] In summary, this invention provides a target voice separation system and method based on bone conduction and air conduction. The system and method of this invention segment the frequency domain signal of a mixed audio signal into several sub-band frequency domain mixed signals through a sub-band segmentation module, employing a "fine low-frequency, sparse high-frequency" segmentation strategy. That is, the signal is subdivided into multiple frequency bands, with particularly fine segmentation in the low-frequency portion, while the high-frequency portion uses coarser segmentation. This strategy allows the model to retain important low-frequency speech information while significantly reducing the computational burden. Furthermore, this invention is the first to introduce the synergistic effect of a bone conduction speech acquisition module and an air conduction speech acquisition module into a target voice separation system. The temporal characteristics of the bone conduction speech signal guide the voice separation module to extract the wearer's target speaker from the speech signal, thereby significantly improving the accuracy of speech recognition and translation.
[0029] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A target human voice separation system based on bone conduction and air conduction, characterized in that, The system includes: Bone conduction speech acquisition module, used to acquire the wearer's bone conduction speech signal; The air conduction voice acquisition module is used to acquire the wearer's air conduction voice signal, as well as other voice signals and background noise; An air conduction input signal processing module is used to process the mixed audio signal acquired by the air conduction voice acquisition module, and convert the time domain signal of the mixed audio signal into a frequency domain signal through short-time Fourier transform; The feature extraction module is used to extract spectral features from the frequency domain signal of the mixed audio signal; The voice separation module inputs the spectral features of the mixed audio signal extracted by the feature extraction module into the voice separation module, separates the individual voice spectral signals of all voices, and performs a short-time inverse Fourier transform on each individual voice spectral signal to finally obtain multiple clean individual voice time-domain signals. The channel selection module is used to output multiple clean single human voice time-domain signals separated by the human voice separation module through multiple signal channels, and to evaluate the confidence of each clean single human voice time-domain signal with the wearer's bone conduction speech signal. Finally, the single human voice time-domain signal with the highest confidence evaluation score is selected as the target human voice audio signal.
2. The target human voice separation system based on bone conduction and air conduction according to claim 1, characterized in that, The system also includes a sub-band segmentation module, used to segment the frequency domain signal of the mixed audio signal into several sub-band frequency domain mixed signals; Before the feature extraction module extracts spectral features from the frequency domain signal of the mixed audio signal, the frequency domain signal of the mixed audio signal is first divided into several sub-band frequency domain mixed signals, and then the feature extraction module extracts sub-band frequency domain features from each sub-band frequency domain mixed signal. All obtained sub-band frequency domain features are input into the human voice separation module for single human voice spectrum signal separation. Each sub-band frequency domain mixed signal is separated into sub-band single human voice spectrum signals of all human voices. Then, the sub-band single human voice spectrum signals belonging to the same human voice in all sub-band frequency domain mixed signals are merged to finally obtain all complete single human voice spectrum signals. Each single human voice spectrum signal is then subjected to short-time inverse Fourier transform to obtain all clean single human voice time domain signals.
3. A target human voice separation system based on bone conduction and air conduction according to claim 2, characterized in that, The subband segmentation module adopts a "fine low-frequency, sparse high-frequency" segmentation strategy when segmenting into several subband frequency domain mixed signals. The "fine low-frequency, sparse high-frequency" segmentation strategy means that the low-frequency part of the frequency domain signal of the mixed audio signal is finely divided with a higher resolution, while the high-frequency part is coarsely divided with a lower resolution.
4. A target human voice separation system based on bone conduction and air conduction according to claim 3, characterized in that, The bandwidth of the frequency domain signal of the mixed audio signal is (0, 8kHz), and the bandwidth (0, 8kHz) is divided into 54 sub-bands, as follows: 0~1kHz: divided into 20 segments, each 50Hz, for a total of 20 sub-bands; 1k~3kHz: divided into 20 bands, each 100Hz, for a total of 20 sub-bands; 3k~6kHz: divided into 10 segments, each 300Hz, for a total of 10 sub-bands; 6kHz~8kHz: Divided into 4 segments, each 500Hz, for a total of 4 sub-bands. The frequency domain signal of the mixed audio signal is divided into 54 sub-band frequency domain mixed signals after passing through 54 sub-bands.
5. A target human voice separation system based on bone conduction and air conduction according to claim 1, characterized in that, The channel selection module also includes a synchronization processing unit, which is used to compensate for the actual system delay, so as to ensure that the bone conduction speech signal and the air conduction speech signal are time-aligned before each clean single human voice time-domain signal output is evaluated with the wearer's bone conduction speech signal.
6. A target human voice separation system based on bone conduction and air conduction according to claim 1, characterized in that, The system also includes a bone conduction input signal processing module, which is used to process the bone conduction speech signal acquired by the bone conduction speech acquisition module, and convert the time domain signal of the bone conduction speech signal into the bone conduction frequency domain signal through short-time Fourier transform. When the voice separation module separates the air-conducted mixed audio signal into multiple clean single voice spectrum signals, each single voice spectrum signal is output through multiple signal channels of the channel selection module. Each clean single voice spectrum signal is then compared with the wearer's bone conduction frequency domain signal for confidence evaluation. The single voice spectrum signal with the highest confidence evaluation score is selected. Finally, a short-time inverse Fourier transform is performed on the selected single voice spectrum signal to obtain the audio signal of the target voice.
7. A method for separating target human voices based on bone conduction and air conduction, characterized in that, The target human voice separation method includes the following steps: S1. The bone conduction speech acquisition module is used to acquire the wearer's bone conduction speech signal; the air conduction speech acquisition module is used to acquire the wearer's air conduction speech signal, other human voices around the wearer, and background noise. S2. The mixed audio signal acquired by the air conduction speech acquisition module in step S1 is processed into frames according to a predetermined frame length and frame shift. The sampling rate of the mixed audio signal is 16KHz. A fast Fourier transform is performed on each frame of the mixed audio signal to convert the time domain signal of the mixed audio signal into a frequency domain signal. S3. The frequency domain signal of the mixed audio signal is divided into several sub-band frequency domain mixed signals by the sub-band segmentation module, and then the sub-band frequency domain features are extracted for each sub-band frequency domain mixed signal by the feature extraction module. S4. Input all the obtained sub-band frequency domain features into the human voice separation module to separate the single human voice spectrum signal. Each sub-band frequency domain mixed signal will be separated into the sub-band single human voice spectrum signal of all human voices. Then, the sub-band single human voice spectrum signals belonging to the same human voice in all sub-band frequency domain mixed signals will be merged to finally obtain all complete single human voice spectrum signals. Then, each single human voice spectrum signal will be subjected to short-time inverse Fourier transform to obtain all clean single human voice time domain signals. S5. All clean single human voice time-domain signals separated by the channel selection module in step S4 are output through multiple signal channels, and each clean single human voice time-domain signal is compared with the wearer's bone conduction speech signal for confidence evaluation. Finally, the single human voice time-domain signal with the highest confidence evaluation score is selected as the target human voice audio signal.
8. The target human voice separation method based on bone conduction and air conduction according to claim 6, characterized in that, In step S5, before performing confidence assessment on each clean single voice time-domain signal output and the wearer's bone conduction speech signal, the actual system delay is compensated by the synchronization processing unit to ensure that the bone conduction speech signal and the air conduction speech signal are aligned in time. Then, each clean single voice time-domain signal output and the wearer's bone conduction speech signal are evaluated for confidence. Finally, the single voice time-domain signal with the highest confidence assessment score is selected as the audio signal of the target voice.
9. The target human voice separation method based on bone conduction and air conduction according to claim 6, characterized in that, Step S2 further includes performing a Fast Fourier Transform on each frame of bone conduction audio signal to convert the time-domain signal of the bone conduction audio signal into a bone conduction frequency-domain signal. When all complete single human voice spectrum signals are finally obtained in step 4, each complete single human voice spectrum signal is output through multiple signal channels of the channel selection module, and each complete single human voice spectrum signal is evaluated with the bone conduction frequency-domain signal to obtain the confidence level. The single human voice spectrum signal with the highest confidence evaluation score is selected, and finally a Short-Time Inverse Fourier Transform is performed on the selected single human voice spectrum signal to obtain the audio signal of the target human voice.
Citation Information
Patent Citations
An end-to-end speaker segmentation method and system based on deep learning
CN109461447B