Refined bone conduction voice recovery method and system using Wav2Vec 2.0 embedding and key value memory network

Through the combination of Wav2Vec 2.0 embedding and key-value memory network, the problem of high-frequency components recovery in bone-conducted speech is solved, and the high-quality enhancement of bone-conducted speech is achieved in noisy environments, improving the reconstruction effect of speech clarity and spectrum details.

CN120279928APending Publication Date: 2025-07-08CHINESE PEOPLES LIBERATION ARMY KET FORCE SERGEANT SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391343.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively restore high-frequency components in bone-conducted speech, especially in noisy environments, and traditional models are difficult to capture the complex nonlinear relationship between bone-conducted speech and air-conducted speech, resulting in speech blurring and limited intelligibility.

Method used

Using the refined bone conduction speech recovery method of Wav2Vec 2.0 embedding and key-value memory network, the time domain enhancement of bone conduction speech is achieved by building an improved Wave-U-Net model and key-value memory network, and using Wav2Vec 2.0 embedding as a recovery prompt in the latent space, combining the multi-head cross attention module and filtering unit to achieve time domain enhancement of bone conduction speech.

Benefits of technology

It significantly improves the enhanced quality of bone-conducted speech, can restore high-frequency components without air-conducted speech reference, improves the clarity and intelligibility of speech, and enhances the reconstruction effect of spectrum details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279928A_ABST
    Figure CN120279928A_ABST
Patent Text Reader

Abstract

The invention discloses a refined bone conduction voice recovery method and a refined bone conduction voice recovery system using a Wav2Vec 2.0 embedding and key value memory network. A novel BC speech enhancement time domain model is provided, and the model uses Wav2Vec 2.0 embedding as a recovery prompt in a potential space to improve the enhancement quality. During training, the dimension of air conduction (AC) speech embedding is extracted from a Wav2Vec 2.0 model and adjusted by using linear interpolation. And dynamically fusing the adjusted embedding into a bottleneck layer of a main network through a cross attention mechanism to serve as the prior of a potential space. Meanwhile, a key value memory network is introduced to bridge the relationship between BC features and AC embedding, so that cross-modal embedding can be retrieved without AC voice reference during reasoning. Due to the fact that embedding is pre-trained on large-scale data, the method has enhanced generalization ability, and due to the fact that the key value memory network can effectively bridge the relation between BC features and AC voice embedding, cross-modal voice clues are provided, and therefore fine recovery is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of voice communication, and particularly relates to a refined bone conduction voice restoration method and system using Wav2Vec 2.0 embedding and key-value memory network. Background Art

[0002] Bone conduction (BC) sensors detect vibrations from body tissues (such as the throat and face), enabling speech capture to be performed in noisy environments (T. Dekens and W. Verhelst, “Body conducted speech enhancement by equalization and signal fusion,” IEEE transactions on audio, speech, and language processing, vol. 21, no. 12, pp. 2481–2492, 2013.). This makes them highly valuable in noise-critical scenarios such as military communication and firefighting. However, body tissues cause significant high-frequency attenuation, resulting in blurred speech and limited intelligibility.

[0003] Traditional models, including Gaussian mixture models (M. Nilsson, H. Gustaftson, S. V. Andersen, and W. B. Kleijn, “Gaussian mixture model based mutual information estimation between frequency bands in speech,” in 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 1. IEEE, 2002, pp. I–525; M. T. Turan and E. Erzin, “Source and filter estimation for throat microphone speech enhancement,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 24, no. 2, pp. 265–275, 2015.) and linear prediction coding (T. T. Vu, M. Unoki, and M. Akagi, “A blind restoration model for bone-conducted speech based on a linear prediction scheme,” IEICE Proceedings Series, vol. 41, no. 19 AM2-C-5, 2007; P. N. Trung, M. Unoki, and M. Akagi, “A study on restoration of bone-conducted speech in noisy environments with lp-based model and gaussian mixture model,” Journal of Signal Processing, vol. 16, no. 5, pp. 409–417, 2012.), have been used to recover the lost high-frequency components in bone-conducted speech. Although effective spectral mapping enhancement has been achieved, they have difficulty capturing the complex non-linear relationship between BC and air-conducted (AC) speech. Deep neural networks are able to better handle these non-linearities.Architectures such as DAE (H.-P. Liu, Y. Tsao, and C.-S. Fuh, “Bone-conducted speech enhancement using deep denoising autoencoder,” Speech Communication, vol. 104, pp. 106–112, 2018.), LSTM (C. Zheng, T. Cao, J. Yang, X. Zhang, and M. Sun, “Spectrum restoration of bone-conducted speech via attention-based contextual information and spectro-temporal structure constraint,” IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, vol. 102, no. 12, pp. 2001–2007, 2019.) and CNN (Y. Li, Y. Wang, X. Liu, Y. Shi, S. Patel, and S.-F. Shih, “Enabling realtime on-chip audio super resolution for bone-conduction microphones,” Sensors, vol. 23, no. 1, p. 35, 2022.) have shown the ability to enhance the amplitude spectrum. However, the persistent phase distortion in these models fundamentally limits the reconstruction quality.

[0004] Recently, time-domain methods such as DPT-EGNet (C. Zheng, L. Xu, X. Fan, J. Yang, J. Fan, and X. Huang, “Dualpath transformer-based network with equalization-generation components prediction for flexible vibrational sensor speech enhancement in the time domain,” The Journal of the Acoustical Society of America, vol. 151, no. 5, pp. 2814–2825, 2022.), EBEN (H. Julien, J. Thomas, Z. Veronique, and B.′Eric, “Configurable eben: Extreme bandwidth extension network to enhance body-conducted speech capture,” IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2023.), and U-Net-Like models (C. Li, F. Yang, and J. Yang, “Restoration of bone-conducted speech with u-net-like model and energy distance loss,” IEEE Signal Processing Letters, 2023.) have demonstrated promising performance through joint amplitude-phase optimization. However, it remains a challenge to recover a large amount of missing high-frequency components from highly restricted low-frequency information (usually 1 - 2 kHz). Summary of the Invention

[0005] An object of the present invention is to provide a refined bone-conducted speech restoration method and system using Wav2Vec 2.0 embedding and key-value memory network for the problems existing in the above-mentioned prior art.

[0006] On the one hand, a technical solution for achieving the object of the present invention is provided: a refined bone-conducted speech restoration method using Wav2Vec 2.0 embedding and key-value memory network, which uses Wav2Vec 2.0 embedding as a restoration hint in the latent space to improve the speech enhancement quality, and specifically includes:

[0007] Step 1, collect the bone-conducted speech audio to be restored;

[0008] Step 2, construct a model for time-domain bone-conducted speech enhancement using Wav2Vec 2.0 embedding and key-value memory network, and train the model with multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library;

[0009] Step 3, input the bone-conducted speech audio to be restored into the model trained in Step 2 to obtain the enhanced bone-conducted speech audio.

[0010] Further, the model for time-domain bone-conducted speech enhancement using Wav2Vec 2.0 embedding and key-value memory network in Step 2 specifically includes: a mainstream module, an embedding extraction module, and a key-value memory module;

[0011] The mainstream module is used to encode the time-domain waveform of the bone-conducted speech, obtain the bottleneck features of the bone-conducted speech, and use the bottleneck features as the input of the key-value memory module; it is also used to fuse the bottleneck features and the output of the key-value memory module, and decode the fused features to obtain the enhanced speech;

[0012] The embedding extraction module is used to use the Wav2Vec 2.0 model to extract the speech embedding Z of the air-conducted speech align and transmit it to the key-value memory module;

[0013] The key-value memory module is used to obtain a speech embedding that mimics the speech embedding Z align according to the bottleneck features and the speech embedding Z align and denote it as the mimicked speech embedding and obtain a speech embedding after compression processing of the speech embedding Z align and denote it as the reconstructed speech embedding

[0014] Further, the mainstream module is an improved Wave-U-Net model, and the improved Wave-U-Net model includes: n downsampling modules, a multi-head cross-attention module, n upsampling modules, and an output layer;

[0015] Among them, each downsampling module includes a one-dimensional convolutional layer and a downsampling layer; each upsampling module includes an upsampling layer and a one-dimensional convolutional layer; the output layer includes a one-dimensional convolutional layer and an activation function layer;

[0016] The input speech audio is sequentially fed into each downsampling module, where the output of the i-th downsampling module is fed into the (i + 1)-th downsampling module. At the same time, the output of the i-th downsampling module is saved and will be subsequently fed into the (n - i + 1)-th upsampling module for processing; the output of the n-th downsampling module forms the bottleneck feature B; i = 1, 2,..., n - 1;

[0017] Taking the bottleneck feature B as the query Q input of the multi-head cross-attention module, and taking the reconstructed speech embedding as the value and key inputs of the multi-head cross-attention module, forming a set of inputs with the bottleneck feature B. At the same time, taking the mimicked speech embedding as the value and key inputs of the multi-head cross-attention module, forming a set of inputs with the bottleneck feature B, respectively obtaining two outputs of the two sets of inputs, feeding both outputs into the j-th upsampling module, concatenating the output obtained by the j-th upsampling module and the output of the (n - j + 1)-th downsampling module, and feeding the concatenated result into the (j + 1)-th upsampling module. The output of the n-th upsampling module is concatenated with the input speech audio and fed into the output layer to obtain the enhanced speech output, where j = 1, 2,..., n - 1.

[0018] Furthermore, the embedding extraction module includes:

[0019] A filtering unit for filtering the air-conducted speech signal;

[0020] A speech embedding unit for processing the filtered speech signal using the Wav2Vec 2.0 model to obtain the embedding Z;

[0021] An embedding dimension adjustment unit for adjusting Z along the time and channel dimensions through a dimension adjuster to obtain an aligned embedding Z align .

[0022] Furthermore, the key-value memory module includes a key memory K and a value memory V, and a correlation between the key memory K and the value memory V is established through an address vector, allowing the retrieval of the embedding in the absence of air-conducted speech, i.e., AC speech;

[0023] The value memory V is used to obtain the mimicked speech embedding according to the bottleneck feature and the speech embedding Z align , and obtain the reconstructed speech embedding and

[0024] The key memory K is used to establish the correlation between the key memory K and the value memory V through the address vector, that is, to bridge the two.

[0025] Furthermore, the value memory V obtains the mimicked speech embedding and the reconstructed speech embedding Specifically, it includes:

[0026] (1) Calculate the address vector of the value memory V

[0027]

[0028] Among them,

[0029]

[0030]

[0031] In the formula, represents the instantaneous embedding in the speech embedding Z align for the time step j, and the cosine similarity between the i-th slot feature v i in V, N is the number of slots in V, and τ is the temperature constant that controls the sparsity of the Softmax distribution;

[0032] (2) Calculate the address vector of the key memory K

[0033]

[0034] Among them,

[0035]

[0036]

[0037]

[0038] In the formula, represents the instantaneous feature b of the bottleneck feature B j for the time step j, i and the cosine similarity between the i-th slot feature k

[0039] (3) Calculate the reconstructed speech embedding

[0040]

[0041] Among them,

[0042]

[0043] In the formula, represents the reconstructed speech embedding corresponding to the time step j, and L represents the number of time steps;

[0044] (4) Calculate the imitated speech embedding

[0045]

[0046] Among them,

[0047]

[0048] In the formula, represents the imitated speech embedding corresponding to the time step j.

[0049] Furthermore, in step 2, training the model specifically includes:

[0050] (1) Using multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library as training samples;

[0051] (2) Constructing the total loss function:

[0052]

[0053] Among them,

[0054]

[0055] In the formula, is the task loss function, used to evaluate the similarity between the restored speech and the target speech, h(*) represents the fusion and decoding operations, g(*) is the multi-scale spectral loss function, and y represents the target speech, that is, the ideal air-conducted speech; is the reconstruction loss function, E j [*] represents summation according to the time step j, represents squaring after calculating the second norm, is the bridging loss function, D KL (·) represents the Kullback-Leibler divergence; λ1, λ2, λ3 are respectively weights;

[0056] (3) Training the model based on the training samples and the total loss function to minimize the total loss function value, thereby obtaining a preliminarily trained model;

[0057] (4) Removing the embedding extraction module in the preliminarily trained model as the finally trained model.

[0058] On the other hand, a refined bone-conducted speech restoration system using Wav2Vec 2.0 embedding and key-value memory network is provided. The system includes:

[0059] The first module is used to collect the bone-conducted speech audio to be restored;

[0060] The second module is used to construct a model for time-domain bone-conducted speech enhancement using Wav2Vec 2.0 embedding and key-value memory network, and train the model with multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library;

[0061] The third module is used to input the bone-conducted speech audio to be restored into the trained model to obtain the enhanced bone-conducted speech audio.

[0062] Compared with the prior art, the remarkable advantages of the present invention are:

[0063] (1) The present invention is a method for bone-conducted speech enhancement processing. By using an improved Wave-U-Net model as the basic bone-conducted speech time-domain enhancement model, the modeling structure of multi-scale features and the optimized upsampling and deconvolution strategies enable the constructed basic model to obtain good bone-conducted speech enhancement effects.

[0064] (2) Extract the AC speech embedding of the Wav2Vec 2.0 model as the recovery clue for the missing frequency components of the bone-conducted speech. Since the Wav2Vec 2.0 model has been pre-trained with large-scale data, the extracted embedding has enhanced generalization ability, which can further implicitly improve the speech enhancement effect.

[0065] (3) In addition, the method uses a key-value memory network to bridge the association between BC speech features and AC speech embeddings, so that in the absence of AC speech, the cross-modal speech clue of AC speech embedding can be recalled only relying on BC speech information.

[0066] (4) During the process of extracting the embedding of the Wav2Vec 2.0 model, a low-frequency filtering unit and a dimension adjustment unit are designed. The low-frequency filtering unit filters the low-frequency speech information so that the extracted embedding contains the high-frequency components most lacking in the bone-conducted speech. This compressed embedding information will make it easier for the key-value network to recall, thereby improving the clue quality. The dimension adjustment unit changes the embedding dimension using linear interpolation, which can flexibly adapt to the basic framework while better retaining the original embedding information.

[0067] (5) In addition, the Wav2Vec 2.0 model only needs to exist during the training process and does not need to be involved in the actual bone-conducted speech enhancement process. Therefore, on the basis of the basic bone-conducted speech time-domain enhancement model, the proposed model only adds a small number of parameters (such as the sizes of the memory network K and V matrices), but the speech enhancement quality is greatly improved, that is, the actual enhancement process does not require the participation of the Wav2Vec 2.0 model, and the BC speech enhancement quality is significantly improved on the premise of minimally increasing the number of parameters.

[0068] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0069] Figure 1 It is a schematic diagram of a model framework for time-domain bone-conducted speech enhancement using Wav2Vec 2.0 embedding and key-value memory network in an embodiment.

[0070] Figure 2 It is a schematic diagram of an embedding extraction module in an embodiment.

[0071] Figure 3 It is a schematic diagram of a mainstream module in an embodiment, where Figure 3 Figure (a) in it is a schematic diagram of the framework of the mainstream module, Figure 3 Figure (b) in it is a schematic diagram of the column representation of the mainstream module. Detailed Embodiment

[0072] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0073] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0074] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0075] In an embodiment, a refined bone-conducted speech restoration method using Wav2Vec 2.0 embedding and key-value memory network is provided. The method uses Wav2Vec 2.0 embedding as a restoration hint in the latent space to improve the speech enhancement quality, and specifically includes:

[0076] Step 1, collect the bone-conducted speech audio to be restored;

[0077] Step 2: Construct a model for time-domain bone-conduction speech enhancement using Wav2Vec 2.0 embedding and key-value memory network, and train the model with multiple pairs of bone-conduction speech audio and air-conduction speech audio in the speech library;

[0078] Step 3: Input the bone-conduction speech audio to be restored into the model trained in Step 2 to obtain the enhanced bone-conduction speech audio.

[0079] Further, in one embodiment, combined with Figure 1 , the model for time-domain bone-conduction speech enhancement using Wav2Vec 2.0 embedding and key-value memory network in Step 2 specifically includes: a mainstream module, an embedding extraction module, and a key-value memory module;

[0080] The mainstream module is used to encode the time-domain waveform of bone-conduction speech to obtain bone-conduction speech bottleneck features, and use the bottleneck features as the input of the key-value memory module; it is also used to fuse the bottleneck features and the output of the key-value memory module, and decode the fused features to obtain enhanced speech;

[0081] The embedding extraction module is used to use the Wav2Vec 2.0 model to extract the speech embedding Z align of air-conduction speech and transmit it to the key-value memory module;

[0082] The key-value memory module is used to obtain a speech embedding that imitates the speech embedding Z align according to the bottleneck features and the speech embedding Z align , denoted as the imitated speech embedding and obtain a speech embedding after compression processing of the speech embedding Z align , denoted as the reconstructed speech embedding

[0083] Further, in one embodiment, the mainstream module is an improved Wave-U-Net model, and the improved Wave-U-Net model includes: n downsampling modules, a multi-head cross-attention module, n upsampling modules, and an output layer;

[0084] Among them, each downsampling module includes a one-dimensional convolutional layer and a downsampling layer; each upsampling module includes an upsampling layer and a one-dimensional convolutional layer; the output layer includes a one-dimensional convolutional layer and an activation function layer;

[0085] Here, the one-dimensional convolutional layer includes a one-dimensional convolution operation, a batch normalization layer, and an activation function layer;

[0086] The input speech audio is sequentially fed into each downsampling module. The output of the i-th downsampling module is fed into the (i + 1)-th downsampling module. Meanwhile, the output of the i-th downsampling module is saved and will be subsequently fed into the (n - i + 1)-th upsampling module for processing. The output of the n-th downsampling module forms the bottleneck feature B, where i = 1, 2,..., n - 1.

[0087] The bottleneck feature B is used as the query Q input of the multi-head cross-attention module, and the reconstructed speech embedding is used as the value and key inputs of the multi-head cross-attention module, forming a set of inputs with the bottleneck feature B. Meanwhile, the imitated speech embedding Z align is used as the value and key inputs of the multi-head cross-attention module, forming a set of inputs with the bottleneck feature B. Two outputs of the two sets of inputs are obtained respectively. Both outputs are fed into the j-th upsampling module. The output obtained by the j-th upsampling module and the output of the (n - j + 1)-th downsampling module are concatenated, and the concatenated result is fed into the (j + 1)-th upsampling module. The output of the n-th upsampling module is concatenated with the input speech audio and fed into the output layer to obtain the enhanced speech output, where j = 1, 2,..., n - 1.

[0088] Exemplarily, preferably, in some embodiments, in combination with Figure 3 , n is taken as 8. When inputting 1s of speech with a dimension of (16384 * 1), the dimension of the enhanced speech output by the mainstream module is (16384 * 1).

[0089] Furthermore, in one of the embodiments, in combination with Figure 2 , the embedding extraction module includes:

[0090] A filtering unit for filtering the air-conducted speech signal.

[0091] A speech embedding unit for processing the filtered speech signal using the Wav2Vec 2.0 model to obtain the embedding Z.

[0092] An embedding dimension adjustment unit for adjusting Z along the time and channel dimensions through a dimension adjuster to obtain an aligned embedding Z align .

[0093] Here, the filtering unit filters the air-conducted speech signal through a high-pass filter. Preferably, the high-pass filter uses a Butterworth high-pass filter. The designed cut-off frequency of the filter is 4 kHz. This filter has a smooth amplitude-frequency characteristic, small signal distortion in the passband, a relatively smooth transition band, and can better retain the high-frequency components of the speech signal.

[0094] Furthermore, in one of the embodiments, the key-value memory module includes a key memory K and a value memory V, and the correlation between the key memory K and the value memory V is established through an address vector, allowing retrieval of the embedding in the absence of air-conducted speech, i.e., AC speech;

[0095] The value memory V is used to obtain the mimicked speech embedding align and the reconstructed speech embedding and the reconstructed speech embedding

[0096] The key memory K is used to establish the correlation between the key memory K and the value memory V through the address vector, that is, to bridge the two.

[0097] Preferably, in some embodiments, the value memory V obtains the mimicked speech embedding and the reconstructed speech embedding Specifically, it includes:

[0098] (1) Calculate the address vector of the value memory V

[0099]

[0100] where

[0101]

[0102]

[0103] In the formula, represents the cosine similarity between the instantaneous embedding align in the speech embedding Z at time step j and the i-th slot feature v i in V, N is the number of slots in V, and τ is the temperature constant that controls the sparsity of the Softmax distribution;

[0104] (2) Calculate the address vector of the key memory K

[0105]

[0106] where

[0107]

[0108]

[0109]

[0110] In the formula, Denote the cosine similarity between the bottleneck feature B and the i-th slot feature k in K for time step j, M is the number of slots in K, and τ is the temperature constant that controls the sparsity of the Softmax distribution; i where,

[0111] (3) Calculate the reconstructed speech embedding

[0112]

[0113] wherein,

[0114]

[0115] In the formula, denotes the reconstructed speech embedding corresponding to time step j, and L denotes the number of time steps;

[0116] (4) Calculate the imitated speech embedding

[0117]

[0118] wherein,

[0119]

[0120] In the formula, denotes the imitated speech embedding corresponding to time step j.

[0121] Furthermore, in one of the embodiments, training the model in step 2 specifically includes:

[0122] (1) Using multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library as training samples;

[0123] (2) Constructing the total loss function:

[0124]

[0125] wherein,

[0126]

[0127]

[0128]

[0129] In the formula, is the task loss function, which is used to evaluate the similarity between the recovered speech and the target speech, h(*) represents the fusion and decoding operations, g(*) is the multi-scale spectral loss function, and y represents the target speech, i.e., the ideal air-conducted speech; is the reconstruction loss function, Ej [*] denotes summation according to time step j, denotes squaring after calculating the two - norm, is the bridging loss function, D KL (·) represents the Kullback - Leibler divergence; λ1, λ2, λ3 are respectively the weights of;

[0130] (3) Train the model based on the training samples and the total loss function to minimize the value of the total loss function, thereby obtaining a preliminarily trained model;

[0131] (4) Remove the embedding extraction module in the preliminarily trained model as the finally trained model.

[0132] Preferably, λ1 = 1, λ2 = 10 4 and λ3 = 10 5 .

[0133] In one embodiment, a refined bone - conduction speech restoration system using Wav2Vec 2.0 embeddings and key - value memory networks is provided. The system includes:

[0134] The first module is used to collect the bone - conduction speech audio to be restored;

[0135] The second module is used to construct a model for enhancing time - domain bone - conduction speech using Wav2Vec 2.0 embeddings and key - value memory networks, and train the model with multiple pairs of bone - conduction speech audio and air - conduction speech audio in the speech library;

[0136] The third module is used to input the bone - conduction speech audio to be restored into the trained model to obtain enhanced bone - conduction speech audio.

[0137] For the specific limitations of the refined bone - conduction speech restoration system using Wav2Vec 2.0 embeddings and key - value memory networks, reference can be made to the limitations of the refined bone - conduction speech restoration method using Wav2Vec 2.0 embeddings and key - value memory networks in the above text, which will not be elaborated here. Each module in the above - mentioned refined bone - conduction speech restoration system using Wav2Vec 2.0 embeddings and key - value memory networks can be implemented in whole or in part by software, hardware, and their combination. The above - mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above - mentioned modules.

[0138] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following is implemented:

[0139] Step 1, collect the bone-conducted speech audio to be restored;

[0140] Step 2, construct a model for time-domain bone-conducted speech enhancement using Wav2Vec 2.0 embedding and key-value memory network, and train the model with multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library;

[0141] Step 3, input the bone-conducted speech audio to be restored into the model trained in Step 2 to obtain the enhanced bone-conducted speech audio.

[0142] For the specific limitations of each step, reference can be made to the limitations of the refined bone-conducted speech restoration method using Wav2Vec 2.0 embedding and key-value memory network in the above text, which will not be elaborated here.

[0143] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following is implemented:

[0144] Step 1, collect the bone-conducted speech audio to be restored;

[0145] Step 2, construct a model for time-domain bone-conducted speech enhancement using Wav2Vec 2.0 embedding and key-value memory network, and train the model with multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library;

[0146] Step 3, input the bone-conducted speech audio to be restored into the model trained in Step 2 to obtain the enhanced bone-conducted speech audio.

[0147] For the specific limitations of each step, reference can be made to the limitations of the refined bone-conducted speech restoration method using Wav2Vec 2.0 embedding and key-value memory network in the above text, which will not be elaborated here.

[0148] In summary, a time-domain BC speech enhancement model proposed by the present invention. It is challenging for this model to achieve higher similarity by using key-value pairs with Wav2Vec2.0 embedding. The representation in the pre-trained Wav2Vec 2.0 model is used as a clue for restoration in the latent space. The key-value memory establishes a cross-modal mapping, so that no AC reference signal is required during the inference process. The supplementary cross-modal clues in the latent space significantly improve the performance, especially in terms of spectral details.

[0149] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the description in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network, characterized in that, The method uses Wav2Vec 2.0 embeddings as recovery hints in the latent space to improve the quality of voice enhancement, specifically including: Step 1, collect the bone-conducted voice audio to be recovered; Step 2, construct a model for time-domain bone-conducted voice enhancement using Wav2Vec 2.0 embeddings and a key-value memory network, and train the model with multiple pairs of bone-conducted voice audio and air-conducted voice audio in the voice library; Step 3, input the bone-conducted voice audio to be recovered into the model trained in Step 2 to obtain the enhanced bone-conducted voice audio.

2. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 1, characterized in that, The model for time-domain bone-conducted voice enhancement using Wav2Vec 2.0 embeddings and a key-value memory network in Step 2 specifically includes: a mainstream module, an embedding extraction module, and a key-value memory module; The mainstream module is used to encode the time-domain waveform of the bone-conducted voice, obtain the bottleneck features of the bone-conducted voice, and use the bottleneck features as the input of the key-value memory module; it is also used to fuse the bottleneck features with the output of the key-value memory module and decode the fused features to obtain the enhanced voice; The embedding extraction module is used to implement the extraction of the speech embedding Z of the air-conducted speech by using the Wav2Vec 2.0 model and transmit it to the key-value memory module; align and extract and transmit it to the key-value memory module; The key-value memory module is used to obtain a voice embedding that mimics the voice embedding Z according to the bottleneck feature and the voice embedding Z align , and obtain a voice embedding that mimics the voice embedding Z align , denoted as the mimicked voice embedding and obtain the voice embedding after compression processing of the voice embedding Z align , denoted as the reconstructed voice embedding 3. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 2, wherein The mainstream module is an improved Wave-U-Net model, and the improved Wave-U-Net model includes: n downsampling modules, a multi-head cross-attention module, n upsampling modules, and an output layer; Among them, each downsampling module includes a one-dimensional convolutional layer and a downsampling layer; each upsampling module includes an upsampling layer and a one-dimensional convolutional layer; the output layer includes a one-dimensional convolutional layer and an activation function layer; The input voice audio is sequentially sent into each downsampling module, where the output of the i-th downsampling module is sent into the (i + 1)-th downsampling module, and at the same time, the output of the i-th downsampling module is saved and will be sent into the (n - i + 1)-th upsampling module for processing later; the output of the n-th downsampling module forms the bottleneck feature B, i = 1, 2,..., n - 1; Input the bottleneck feature B as the query Q of the multi-head cross-attention module, and input the reconstructed speech embedding as the value and key of the multi-head cross-attention module, forming a set of inputs with the bottleneck feature B. At the same time, input the mimicked speech embedding as the value and key of the multi-head cross-attention module, forming a set of inputs with the bottleneck feature B, and obtaining two outputs of the two sets of inputs respectively. Send the two outputs into the j-th upsampling module, concatenate the output obtained by the j-th upsampling module and the output of the (n - j + 1)-th downsampling module, and send the concatenated result into the (j + 1)-th upsampling module. Concatenate the output of the n-th upsampling module with the input speech audio, send it into the output layer, and obtain the enhanced speech output, where j = 1, 2,..., n - 1.

4. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 2, characterized in that The embedding extraction module includes: A filtering unit for filtering the air-conducted voice signal; A voice embedding unit for processing the filtered voice signal using the Wav2Vec 2.0 model to obtain the embedding Z; An embedding dimension adjustment unit for adjusting Z along the time and channel dimensions by a dimension adjuster to obtain an aligned embedding Z align .

5. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 4, characterized in that, The filtering unit filters the air-conducted voice signal through a high-pass filter.

6. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 2, wherein The key-value memory module includes a key memory K and a value memory V, and a correlation between the key memory K and the value memory V is established through an address vector, allowing the retrieval of embeddings in the absence of air-conducted voice, i.e., AC voice; The value memory V is used to obtain the embedded voice of the imitation align according to the bottleneck feature and the voice embedding Z and the embedded voice of the reconstruction The key memory K is used to establish the correlation between the key memory K and the value memory V through the address vector, that is, to bridge the two.

7. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 6, characterized in that, The value memory V obtains the imitated speech embedding Specifically, it includes: (1) Calculate the address vector of the key memory K Among them, where represents the instantaneous feature b of the bottleneck feature B for the time step j j and the cosine similarity between the i-th slot feature k in K i where M is the number of slots in K and τ is the temperature constant that controls the sparsity of the Softmax distribution; (2) Calculate the imitated speech embedding Among them, In the formula, represents the mimicked speech embedding corresponding to the time step j, and L represents the number of time steps.

8. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 6, wherein, The value memory V obtains the reconstructed speech embedding Specifically, it includes: (1) Calculate the address vector of the value memory V Among them, where denotes the instantaneous embedding in the speech embedding Z align for the time step j and the i-th slot feature v in V i The cosine similarity between them, N is the number of slots in V, and τ is the temperature constant that controls the sparsity of the Softmax distribution; (2) Calculate the reconstructed speech embedding Among them, In the formula, represents the reconstructed speech embedding corresponding to the time step j, and L represents the number of time steps.

9. The refined bone conduction speech restoration method using Wav2Vec 2.0 embedding and key-value memory network according to claim 7 or 8, characterized in that, The training of the model in Step 2 specifically includes: (1) Use multiple pairs of bone-conducted voice audio and air-conducted voice audio in the voice library as training samples; (2) Construct a total loss function: Among them, wherein, is the task loss function, used to evaluate the similarity between the restored speech and the target speech, h(*) represents the fusion and decoding operations, g(*) is the multi-scale spectral loss function, and y represents the target speech, i.e., the ideal air-conducted speech; is the reconstruction loss function, and E j [*] represents the summation according to the time step j, represents squaring after calculating the second norm, is the bridging loss function, and D KL (·) represents the Kullback-Leibler divergence; λ1, λ2, and λ3 are the weights of respectively; (3) Train the model based on the training samples and the total loss function to minimize the value of the total loss function, thereby obtaining the preliminarily trained model; (4) Remove the embedding extraction module in the preliminarily trained model as the finally trained model.

10. A refined bone conduction speech restoration system using Wav2Vec 2.0 embedding and key-value memory network based on the method according to any one of claims 1 to 9, characterized in that, The system includes: The first module is used to collect the bone-conducted speech audio to be restored; The second module is used to build a model for enhancing time-domain bone-conducted speech using Wav2Vec 2.0 embeddings and key-value memory networks, and train the model with multiple pairs of bone-conducted speech audio and air-conducted speech audio in the speech library; The third module is used to input the bone-conducted speech audio to be restored into the trained model to obtain the enhanced bone-conducted speech audio.