Speech processing method and system for bone conduction microphone

By employing spectral quality energy analysis, electromyography and inertial co-discrimination, and an improved deep learning model, the problem of pseudo-sound interference in noisy environments using bone conduction microphones was solved, achieving high-accuracy and high-purity speech recognition.

CN121171218BActive Publication Date: 2026-02-06BEIJING MILESTONE SCI & TECH DEV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511717881.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-02-06
Estimated Expiration
2045-11-21

AI Technical Summary

Technical Problem

Existing bone conduction microphones are susceptible to pseudo-sound interference in speech recognition in noisy environments, especially non-speech disturbances such as teeth clenching and swallowing, which occur frequently. Furthermore, existing methods are unable to effectively identify and eliminate pseudo-sound interference, leading to a decrease in recognition accuracy.

Method used

A deep modeling approach combining spectral mass energy analysis, electromyography and inertial co-discrimination, an improved RealFormer model, and the Wav2Vec 2.0 model is employed. Through a multimodal signal joint detection and elimination mechanism, pseudo-speech interference is identified and removed, thereby improving the robustness and purity of speech recognition.

Benefits of technology

It significantly improves the accuracy of pseudo-sound recognition and the quality of speech recognition in noisy environments, and achieves highly robust and pure voice interaction and command control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171218B_ABST
    Figure CN121171218B_ABST
Patent Text Reader

Abstract

The application discloses a voice processing method and system of a bone conduction microphone, comprising the following steps: step one: collecting the original voice signal of the bone conduction microphone and performing frame slicing to generate a spectrum quality energy spectrum; step two: identifying the time interval of the multi-band slope synchronous mutation as a pseudo-sound candidate section; step three: marking the pseudo-sound section; step four: constructing a frame state sequence; step five: generating a residual voice signal and a pseudo-sound frame sequence; step six: performing spectrum transformation on the pseudo-sound frame sequence to generate a pseudo-sound spectrum frame sequence, and inputting the pseudo-sound spectrum frame sequence into an improved RealFormer model to generate a frame-level pseudo-sound state label vector; and step seven: jointly inputting the residual voice signal and the frame-level pseudo-sound state label vector into a Wav2Vec 2.0 model to finally output a purified voice recognition result. The application realizes the purified recognition of bone conduction voice by fusing pseudo-sound detection and an improved RealFormer model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, in particular to a speech processing method and system of bone conduction microphone. BACKGROUND

[0002] With the wide application of voice interaction technology in man-machine communication in noisy environment, intelligent head-mounted devices and remote communication scenarios, bone conduction microphones are widely concerned in complex environment speech collection and recognition tasks due to their strong anti-airborne noise ability and compact structure. However, the existing speech recognition methods based on bone conduction signals have the following problems in practical application:

[0003] Non-speech disturbances (such as clenching teeth, swallowing, expression changes, etc.) frequently occur in bone conduction signals, and their time domain and frequency domain features are similar to those of effective speech, which are easily misidentified as speech signals, resulting in a significant decrease in recognition results; existing pseudo-sound detection methods mostly rely on single acoustic features or energy threshold judgment, lack modeling ability for multi-band coordinated mutation, and are difficult to cope with multi-source pseudo-sound interference, which is prone to false positives and false negatives; although some methods introduce electromyography or inertial signals, they do not fully utilize the mutation linkage characteristics between multi-modal signals, resulting in limited pseudo-sound judgment accuracy; in addition, mainstream speech recognition models mostly ignore the pseudo-sound structure in bone conduction signals, do not realize joint modeling of speech signals and pseudo-sound states, and cannot eliminate the influence of pseudo-sound interference on decoding results at the model level, which restricts the improvement of pure speech recognition capability.

[0004] Therefore, how to provide a speech processing method and system of bone conduction microphone is a problem that those skilled in the art need to solve. SUMMARY

[0005] One object of the present application is to provide a speech processing method and system of bone conduction microphone, which integrates spectral energy analysis, electromyography and inertial signal collaborative discrimination, improved RealFormer model and Wav2Vec 2.0 model deep modeling method, proposes a multi-modal joint detection and elimination mechanism to solve the problem of bone conduction microphone speech being easily disturbed by pseudo-sound such as clenching teeth and swallowing, significantly improves the accuracy of pseudo-sound recognition and removal, and finally realizes high robustness and high purity of speech recognition output, which is suitable for stable voice interaction and command control in noisy environments.

[0006] According to the speech processing method of the bone conduction microphone according to the embodiment of the present application, the following steps are included:

[0007] Step 1: Collect the original speech signal of the bone conduction microphone and perform frame slicing, use a band-pass filter bank to perform frequency band decomposition on each frame, calculate the spectral energy density value of each frequency band, and generate a spectral energy map;

[0008] Step two: perform first-order difference operation on the spectrum power spectrum to obtain a spectrum slope time sequence, calculate the spectrum slope covariance matrix between frequency bands, and identify the time interval of the multi-band slope synchronous mutation as a candidate segment of the artifact;

[0009] Step three: collect the electromyographic signal and inertial measurement signal corresponding to the time of the artifact candidate segment, and determine whether there is a muscle impulse mutation and an acceleration mutation, if both exist, mark the artifact candidate segment as a confirmed artifact segment;

[0010] Step four: based on the confirmed artifact segment, construct a frame state sequence, and mark the frames in the original speech signal as silent frames, speech frames or confirmed artifact frames;

[0011] Step five: delete the confirmed artifact frames in the frame state sequence to obtain a continuous speech segment composed of the remaining frames, which is defined as a residual speech signal; all deleted confirmed artifact frames form an artifact frame sequence;

[0012] Step six: perform spectrum transformation on the artifact frame sequence to generate an artifact spectrum frame sequence, and input it into the improved RealFormer model to generate a frame-level artifact state label vector;

[0013] Step seven: input the residual speech signal and the frame-level artifact state label vector into the Wav2Vec2.0 model, and finally output the purified speech recognition result.

[0014] Optionally, the step one is specifically:

[0015] Collect the bone conduction microphone original speech signal, and frame slice the collected bone conduction microphone original speech signal in a frame length of 8 milliseconds and a frame shift of 4 milliseconds to generate a frame sequence;

[0016] Input each frame in the frame sequence into a set of bandpass filter groups with a set center frequency and bandwidth, the bandpass filter group includes no less than 16 filter channels, the center frequency covers the bone conduction effective speech frequency band from 50Hz to 1500Hz, and outputs a frequency band signal set;

[0017] Calculate the instantaneous energy value of each frequency band signal, the instantaneous energy value is the average value of the square of the amplitude of the frequency band signal sampling points in the current frame;

[0018] Calculate the absolute mean of the first-order amplitude derivative in each frequency band signal as the mutation coefficient of the frequency band signal;

[0019] Multiply the instantaneous energy value of each frequency band signal with the mutation coefficient to obtain the spectrum quality energy density value of the corresponding frequency band;

[0020] All spectral energy density values of each frame are spliced into a spectral energy vector in frequency band order, and spectral energy vectors of all frames are spliced in time order to obtain a spectral energy atlas.

[0021] Optionally, the step two is specifically:

[0022] First-order difference operation is performed on spectral energy vectors of adjacent frames in the spectral energy atlas to obtain a spectral slope time sequence;

[0023] In the spectral slope time sequence, a spectral slope vector corresponding to each frame is extracted, and the spectral slope vector is composed of spectral energy density change values of each frequency band in the frame;

[0024] Based on the spectral slope vector of each frame, a covariance between spectral energy density change values of any two frequency bands in the frame is calculated to construct a spectral slope covariance matrix, and the spectral slope covariance matrix represents a coordination of spectral slope changes between frequency bands and is a key statistical feature for identifying multi-band synchronous disturbance (such as pseudo sound) and is used for automatically detecting sudden abnormal behaviors in a time period in a sliding window;

[0025] In a time sequence sliding window manner, statistical analysis is performed on spectral slope covariance matrices corresponding to consecutive frames, and a covariance average absolute value and a covariance change rate in each time sequence sliding window are calculated;

[0026] When the covariance average absolute value exceeds a set average threshold, the covariance change rate exceeds a set change rate threshold, and the number of frequency band pairs with the same change direction exceeds a preset proportion of the total number of frequency band pairs, it is determined that a time interval corresponding to the time sequence sliding window is a multi-band slope synchronous mutation period;

[0027] The time interval corresponding to the multi-band slope synchronous mutation period is output as a pseudo sound candidate period.

[0028] Optionally, the step three is specifically:

[0029] The electromyographic signal and the inertial measurement signal in the time interval corresponding to the pseudo sound candidate period are extracted, the electromyographic signal is an electrophysiological voltage sequence collected by a facial electromyographic sensor, and the inertial measurement signal is a three-axis acceleration time sequence collected by an acceleration sensor;

[0030] The electromyographic signal is calculated to obtain an electromyographic mean square energy value sequence in each sliding window;

[0031] When the electromyographic mean square energy value in the sliding window exceeds a preset energy mutation threshold compared with that in a previous sliding window, and the amplitude difference value signs of consecutive sampling points in the sliding window remain consistent, it is determined that there is an electromyographic pulse mutation in the sliding window.

[0032] Performing a sliding window maximum acceleration difference operation on the inertial measurement signal in each axis respectively, when the maximum acceleration difference of any axis is greater than a set acceleration mutation threshold, it is determined that there is an acceleration mutation in the sliding window;

[0033] When the pseudo sound candidate segment corresponds to any sliding window in which the muscle electric pulse mutation and the acceleration mutation are detected at the same time, the pseudo sound candidate segment is marked as a confirmed pseudo sound segment.

[0034] Optionally, the step four is specifically:

[0035] According to the time interval of the confirmed pseudo sound segment, mark the corresponding frame in the bone conduction microphone original speech signal as a confirmed pseudo sound frame;

[0036] Among the frames not marked as confirmed pseudo sound frames, mark the continuous frames with an energy value lower than a set mute threshold and a frame number not less than a preset mute frame number as mute frames, and mark the remaining unmarked frames as speech frames;

[0037] Construct a frame state sequence arranged in time sequence, which is composed of all mute frames, speech frames and confirmed pseudo sound frames.

[0038] Optionally, the pseudo sound frame sequence is subjected to a spectrum transformation to generate a pseudo sound atlas frame sequence, specifically:

[0039] Perform a short-time Fourier transform on the audio signal corresponding to each frame in the pseudo sound frame sequence to obtain the spectrum of each frame;

[0040] Divide the spectrum of each frame into several frequency band regions along the frequency axis, and calculate the spectral amplitude mean value in each frequency band region to form the spectral feature vector of the frame;

[0041] Arrange all the spectral feature vectors of the frames in time sequence to generate a pseudo sound atlas frame sequence.

[0042] Optionally, the improved RealFormer model specifically includes:

[0043] In each Transformer layer of the original RealFormer model, replace the standard feedforward network with a perturbation alignment double-channel residual module, which includes a main channel convolution submodule and a pseudo sound perturbation enhancement submodule;

[0044] The main channel convolution submodule includes two sequentially connected 1×1 convolution layers and a layer of grouped convolution layer with GELU activation function, which is used to extract the local spectral features of each frame in the pseudo sound atlas frame sequence and output a frame-level convolution feature vector;

[0045] The pseudo sound perturbation enhancement submodule includes an inter-frame perturbation aggregation channel and a frequency band mutation detection channel;

[0046] The inter-frame disturbance aggregation channel forms a symmetric three-frame window with the current frame and the adjacent frames before and after the current frame, respectively calculates the spectral amplitude vector difference between the current frame and the adjacent frames, and takes the average of the two spectral amplitude vector differences as the disturbance amplitude vector of the current frame;

[0047] The frequency band mutation detection channel performs a first-order difference operation on the spectral feature vector of each frame along the frequency dimension to obtain a frequency band gradient sequence corresponding to the frame; in the frequency band gradient sequence, a plurality of frequency band indexes with the largest gradient amplitudes and corresponding gradient values are sequentially extracted, and the extracted frequency band gradient values are spliced in order of the frequency band indexes to generate a frequency band mutation vector;

[0048] The disturbance amplitude vector and the frequency band mutation vector are spliced in the feature dimension to generate a disturbance enhanced feature vector corresponding to the frame;

[0049] The frame-level convolution feature vector and the disturbance enhanced feature vector are spliced along the channel dimension to generate a fusion representation vector;

[0050] The fusion representation vector is input into a frame-level classification network, and the frame-level classification network includes a full connection layer, a normalization layer and a Softmax classification layer, and outputs a frame-level artifact state label vector consistent with the number of frames of the artifact map frame sequence.

[0051] Optionally, the Wav2Vec 2.0 model includes a time convolution encoder, a linear embedding layer, a Transformer context network and a speech recognition decoding head;

[0052] The residual speech signal is input into the time convolution encoder to extract an acoustic feature vector of each frame and generate a frame-level acoustic feature vector sequence;

[0053] The frame-level artifact state label vector is input into the linear embedding layer to be mapped into a frame-level artifact state embedding vector sequence;

[0054] The frame-level acoustic feature vector sequence and the frame-level artifact state embedding vector sequence are spliced in the feature dimension frame by frame to construct a joint feature sequence;

[0055] The joint feature sequence is input into the Transformer context network to extract a frame-level context representation vector sequence with long-distance context dependency;

[0056] The frame-level context representation vector sequence is input into a speech recognition decoding head, and the speech recognition decoding head includes a full connection layer and a CTC decoding layer, and outputs a purified speech recognition result.

[0057] Optionally, the CTC decoding layer receives the frame-level character class prediction distribution output by the full connection layer, performs path compression and blank symbol removal operations on each frame prediction result, generates a continuous character sequence as a purified speech recognition result and outputs the purified speech recognition result.

[0058] According to the speech processing system of the bone conduction microphone, the following modules are included:

[0059] The speech signal acquisition and processing module is configured to acquire a raw speech signal of the bone conduction microphone, perform frame slicing and frequency band decomposition, and generate a spectrogram;

[0060] The pseudo-sound candidate segment identification module is configured to calculate a spectrum slope time sequence and a spectrum slope covariance matrix, identify a time interval of a multi-frequency band slope synchronous mutation, and output the time interval as a pseudo-sound candidate segment.

[0061] The pseudo-sound confirmation module is configured to acquire an electromyographic signal and an inertial measurement signal corresponding to a time of the pseudo-sound candidate segment, determine whether there is an electromyographic pulse mutation and an acceleration mutation, and if both exist, mark the pseudo-sound candidate segment as a confirmed pseudo-sound segment.

[0062] The frame state construction module is configured to mark a silent frame, a speech frame, and a confirmed pseudo-sound frame, and construct a frame state sequence.

[0063] The residual speech extraction module is configured to generate a residual speech signal and a pseudo-sound frame sequence.

[0064] The pseudo-sound modeling module is configured to perform spectrum transformation on the pseudo-sound frame sequence, generate a pseudo-sound spectrogram frame sequence, and input the pseudo-sound spectrogram frame sequence into an improved RealFormer model to generate a frame-level pseudo-sound state label vector.

[0065] The speech recognition module is configured to input the residual speech signal and the frame-level pseudo-sound state label vector into a Wav2Vec 2.0 model, and output a purified speech recognition result.

[0066] The speech processing system of the bone conduction microphone has the following advantages:

[0067] The application constructs a bone conduction artifact detection mechanism combining spectral power spectrum and spectral slope covariance analysis, proposes a systematic solution of multi-modal mutation discrimination and spectral disturbance modeling for the multi-band artifact interference and non-speech disturbance problems existing in bone conduction speech signals. Firstly, the spectral power spectrum is constructed on the basis of frame-level frequency band energy processing, and the stable identification of multi-band synchronous mutation segments is realized through the spectral slope time series and the covariance matrix, overcoming the problem of insufficient description of multi-band disturbance by the traditional energy threshold method; the dual-channel synchronous mutation detection mechanism of electromyographic signal and inertial measurement signal is introduced, and the artifact segments caused by non-speech actions such as clenching and swallowing are effectively identified through the joint judgment of electromyographic pulses and acceleration mutations. For the difficulty of artifact modeling, the application constructs a disturbance alignment dual-channel residual module to optimize the RealFormer structure, on the one hand, utilizes the main channel convolution to extract local spectral features, and on the other hand, models complex interference patterns through the inter-frame disturbance aggregation channel and the frequency band mutation detection channel, and improves the frame-level artifact state classification accuracy. In the speech recognition stage, the residual speech and frame-level artifact state label are jointly input into the Wav2Vec 2.0 model, the feature splicing and Transformer context network fusion modeling are used to effectively enhance the context modeling ability of pure speech information, improve the recognition accuracy under artifact interference, and finally realize the unified cooperation of accurate detection, disturbance modeling and recognition purification of the artifact segments in bone conduction speech signals, significantly improving the robustness and output quality of the speech recognition system in strong interference environment. BRIEF DESCRIPTION OF DRAWINGS

[0068] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:

[0069] Figure 1 The overall flowchart of the speech processing method of the bone conduction microphone proposed by the application is shown in the figure;

[0070] Figure 2 The structural schematic diagram of the speech processing system of the bone conduction microphone proposed by the application is shown in the figure. DETAILED DESCRIPTION

[0071] The application will now be described in further detail with reference to the drawings. These drawings are simplified schematic diagrams, and only illustrate the basic structure of the application in a schematic manner, and therefore only show the components related to the application.

[0072] Reference Figure 1 A speech processing method of a bone conduction microphone, comprising the following steps:

[0073] Step one: collect the original speech signal of the bone conduction microphone and perform frame slicing, use a band-pass filter set to perform frequency band decomposition on each frame, calculate the spectral power density value of each frequency band, and generate a spectral power spectrum;

[0074] Step two: perform first-order difference operation on the spectral power spectrum to obtain a spectral slope time sequence, calculate the spectral slope covariance matrix between frequency bands, identify the time interval of multi-band slope synchronous mutation as the pseudo-sound candidate segment;

[0075] Step three: collect the electromyographic signal and inertial measurement signal corresponding to the time of the pseudo-sound candidate segment, judge whether there is an electromyographic pulse mutation and an acceleration mutation, if both exist, mark the pseudo-sound candidate segment as a confirmed pseudo-sound segment;

[0076] Step four: based on the confirmed pseudo-sound segment, construct a frame state sequence, and mark the frames in the original speech signal as silent frames, speech frames or confirmed pseudo-sound frames;

[0077] Step five: delete the confirmed pseudo-sound frames in the frame state sequence to obtain a continuous speech segment composed of the remaining frames, which is defined as a residual speech signal; all deleted confirmed pseudo-sound frames form a pseudo-sound frame sequence;

[0078] Step six: perform frequency spectrum transformation on the pseudo-sound frame sequence to generate a pseudo-sound spectrum frame sequence, and input it into the improved RealFormer model to generate a frame-level pseudo-sound state label vector;

[0079] Step seven: input the residual speech signal and the frame-level pseudo-sound state label vector into the Wav2Vec2.0 model, and finally output the purified speech recognition result.

[0080] In the embodiment, the step one is specifically:

[0081] Collect the original speech signal of the bone conduction microphone, and perform frame slicing on the collected original speech signal of the bone conduction microphone in a frame length of 8 milliseconds and a frame shift of 4 milliseconds to generate a frame sequence;

[0082] Input each frame in the frame sequence into a set of band-pass filter sets with a set center frequency and bandwidth, the band-pass filter set includes no less than 16 filter channels, the center frequency covers the bone conduction effective speech frequency band from 50Hz to 1500Hz, and a frequency band signal set is output;

[0083] Calculate the instantaneous energy value of each frequency band signal, the instantaneous energy value is the average value of the amplitude square of the frequency band signal sampling points in the current frame;

[0084] Calculate the absolute mean value of the first-order amplitude derivative in each frequency band signal as the mutation coefficient of the frequency band signal;

[0085] The instantaneous energy value of each frequency band signal is multiplied by the mutation coefficient to obtain a spectral energy density value of the corresponding frequency band.

[0086] All spectral energy density values of each frame are spliced into a spectral energy vector in the order of frequency bands, and spectral energy vectors of all frames are spliced in the order of time to obtain a spectral energy atlas.

[0087] In the embodiment, the step two is specifically:

[0088] A first-order difference operation is performed on the spectral energy vectors of adjacent frames in the spectral energy atlas to obtain a spectral slope time sequence.

[0089] In the spectral slope time sequence, a spectral slope vector corresponding to each frame is extracted, and the spectral slope vector is composed of spectral energy density change values of each frequency band in the frame.

[0090] Based on the spectral slope vector of each frame, a covariance between spectral energy density change values of any two frequency bands in the frame is calculated to construct a spectral slope covariance matrix, and the spectral slope covariance matrix represents the cooperativity of spectral slope change between frequency bands and is a key statistical feature for identifying multi-band synchronous disturbance (such as pseudo sound), which is used for automatically detecting sudden abnormal behaviors in a time period in a sliding window.

[0091] In a time sequence sliding window manner, the spectral slope covariance matrix corresponding to a plurality of continuous frames is statistically analyzed, and the average absolute value of the covariance and the covariance change rate in each time sequence sliding window are calculated.

[0092] When the average absolute value of the covariance exceeds a set average threshold, the covariance change rate exceeds a set change rate threshold, and the number of frequency band pairs with the same change direction exceeds a preset proportion of the total number of frequency band pairs, it is determined that the time interval corresponding to the time sequence sliding window is a multi-band slope synchronous mutation section.

[0093] The time interval corresponding to the multi-band slope synchronous mutation section is output as a pseudo sound candidate section.

[0094] In the present application, step two is to perform first-order difference processing on the spectral quality-energy graph, construct a frame-level spectral slope time sequence, and further calculate the covariance relationship between frequency bands based on the spectral slope vector of each frame to form a spectral slope covariance matrix, which not only improves the ability to describe the synchronization of frequency band changes, but also enhances the accurate identification ability of the coordinated mutation characteristics of multiple frequency bands when the pseudo sound disturbance occurs. Compared with the pseudo sound detection method of the existing technology, the energy change or threshold judgment of a single frequency band, the present application quantifies the dynamic correlation between multiple frequency bands through the covariance matrix, and introduces a sliding window statistical method to jointly judge the average amplitude and change rate of the covariance, thereby realizing stable detection of the pseudo sound burst section in the time dimension, effectively reducing the false positive and false negative probabilities, enhancing the judgment accuracy and robustness of the pseudo sound candidate section, effectively identifying the multi-source pseudo sound interference situation, and improving the input purity and overall recognition accuracy of speech recognition.

[0095] In the present embodiment, step three is specifically:

[0096] Extracting the electromyographic signal and the inertial measurement signal in the time interval corresponding to the pseudo sound candidate section, the electromyographic signal being an electrophysiological voltage sequence collected by a facial electromyographic sensor, and the inertial measurement signal being a three-axis acceleration time sequence collected by an acceleration sensor;

[0097] Calculating the electromyographic mean square energy value in each sliding window for the electromyographic signal to obtain an electromyographic mean square energy value sequence;

[0098] When the electromyographic mean square energy value in the sliding window exceeds the preset energy mutation threshold compared with the previous sliding window, and the amplitude difference value signs of a plurality of consecutive sampling points in the sliding window remain consistent, it is determined that there is an electromyographic pulse mutation in the sliding window;

[0099] Performing sliding window maximum acceleration difference operation on the inertial measurement signal in each axis, respectively, and when the maximum acceleration difference of any axis is greater than the set acceleration mutation threshold, it is determined that there is an acceleration mutation in the sliding window;

[0100] When electromyographic pulse mutation and acceleration mutation are simultaneously detected in any sliding window corresponding to the pseudo sound candidate section, the pseudo sound candidate section is marked as a confirmed pseudo sound section;

[0101] In the present application, in view of the non-speech pseudo-sound interference problem existing in the bone conduction speech signal, by introducing the synchronization discrimination mechanism of the electromyography signal and the inertial measurement signal, the accuracy and robustness of the pseudo-sound detection are significantly improved. The electromyography signal reflects the electrophysiological changes generated by the user when the facial muscle group (such as the masseter muscle and the temporal muscle) is moved violently, which can effectively distinguish the physiological state difference between pseudo-sound and normal speech generation process; at the same time, the inertial measurement signal provides three-axis acceleration data, which can assist in judging the rapid movement behavior of the head or face, and further exclude the bone vibration interference caused by non-speech action.

[0102] In this step, the sliding window energy analysis and change rate detection method is used to quantify the transient characteristics of electromyography pulse mutation and acceleration mutation, the pseudo-sound characteristic window is judged using the set threshold and direction consistency rule, and the false detection rate is significantly reduced through the judgment mechanism that the double conditions are met at the same time. Compared with the traditional method which only relies on acoustic features or a single sensing channel, the present application can effectively eliminate the pseudo-sound segments caused by clenching teeth, swallowing and rapid expression changes and other non-speech movements, and provide a higher quality data basis for pseudo-sound state modeling and speech recognition.

[0103] In the present embodiment, the step four is specifically:

[0104] According to the time interval of the confirmed pseudo-sound segment, the corresponding frames in the bone conduction microphone original speech signal are marked as confirmed pseudo-sound frames;

[0105] Among the frames which are not marked as confirmed pseudo-sound frames, the continuous frames with energy value lower than the set silence threshold and frame number not less than the preset silence frame number are marked as silence frames, and the remaining unmarked frames are marked as speech frames;

[0106] A frame state sequence arranged in time sequence is constructed, which is composed of all silence frames, speech frames and confirmed pseudo-sound frames.

[0107] In the present embodiment, the pseudo-sound frame sequence is subjected to frequency spectrum transformation to generate a pseudo-sound atlas frame sequence, which is specifically:

[0108] The short-time Fourier transform is performed on the audio signal corresponding to each frame in the pseudo-sound frame sequence to obtain the frequency spectrum of each frame;

[0109] The frequency spectrum of each frame is divided into several frequency band regions along the frequency axis, and the average amplitude of the frequency spectrum in each frequency band region is calculated to form the frequency spectrum feature vector of the frame;

[0110] All the frequency spectrum feature vectors of the frames are arranged in time sequence to generate a pseudo-sound atlas frame sequence.

[0111] In the present embodiment, the improved RealFormer model specifically includes:

[0112] In each Transformer layer of the original RealFormer model, a standard feedforward network is replaced by a perturbation alignment dual-path residual module, which includes a main path convolution submodule and a pseudo sound perturbation enhancement submodule;

[0113] The main path convolution submodule includes two sequentially connected 1x1 convolution layers and a group convolution layer with a GELU activation function, which is used to extract the local spectral features of each frame in the pseudo sound spectrogram frame sequence and output a frame-level convolution feature vector;

[0114] The application uses a group convolution layer with a GELU activation function to extract the local spectral features of the pseudo sound spectrogram frame sequence in the main path. The convolution layer divides the input channels into multiple sub-channel groups and performs convolution operations respectively, effectively reducing the parameter quantity and computational complexity while improving the local sensitivity of feature extraction. The introduction of the GELU activation function can enhance the model's ability to model nonlinear features. Compared with traditional activation functions such as ReLU, its output is continuous and has high-order smoothness, which helps to maintain the stability of the feature distribution, improve the training convergence speed and recognition accuracy, and thus enhance the expression ability of complex pseudo sound interference patterns;

[0115] The pseudo sound perturbation enhancement submodule includes an inter-frame perturbation aggregation channel and a frequency band mutation detection channel;

[0116] The inter-frame perturbation aggregation channel forms a symmetric three-frame window with the current frame and the adjacent frames, respectively calculates the spectral amplitude vector difference between the current frame and the adjacent frames, and takes the average of the two spectral amplitude vector differences as the perturbation amplitude vector of the current frame;

[0117] The frequency band mutation detection channel performs a first-order difference operation on the spectral feature vector of each frame along the frequency dimension to obtain a frequency band gradient sequence corresponding to the frame. In the frequency band gradient sequence, a number of frequency band indexes with the largest gradient amplitudes and corresponding gradient values are extracted in turn, and the extracted frequency band gradient values are spliced in the order of frequency band indexes to generate a frequency band mutation vector;

[0118] The perturbation amplitude vector and the frequency band mutation vector are spliced in the feature dimension to generate a perturbation enhancement feature vector corresponding to the frame;

[0119] The frame-level convolution feature vector and the perturbation enhancement feature vector are spliced along the channel dimension to generate a fusion representation vector;

[0120] The fusion representation vector is input into a frame-level classification network, which includes a fully connected layer, a normalization layer and a Softmax classification layer, and outputs a frame-level pseudo sound state label vector consistent with the number of frames in the pseudo sound spectrogram frame sequence;

[0121] In order to improve the accuracy of pseudo-sound detection in bone conduction microphone voice signals, the application proposes an improved RealFormer model with optimized structure, the core improvement of which is that the standard feedforward network of the Transformer layer in the original RealFormer model is replaced by a perturbation alignment double-channel residual module, which is composed of a main channel convolution submodule and a pseudo-sound perturbation enhancement submodule. The main channel convolution submodule adopts a two-layer sequentially connected 1x1 convolution and a layer of grouped convolution structure with a GELU activation function, which can effectively extract the local spectral features of each frame in the pseudo-sound spectrum frame sequence and output a frame-level convolution feature vector.

[0122] On this basis, the pseudo-sound perturbation enhancement submodule is provided with an inter-frame perturbation aggregation channel and a frequency band mutation detection channel in parallel, which are used to mine the dynamic perturbation characteristics of the interference signal. Among them, the inter-frame perturbation aggregation channel forms a symmetrical three-frame window with the current frame and its adjacent frames before and after it, calculates the spectral amplitude vector difference of the current frame and the adjacent frames respectively, takes the average of the differences, generates an inter-frame perturbation amplitude vector, and thus captures the inter-frame perturbation trend; the frequency band mutation detection channel performs a first-order difference operation on the spectral feature vector of each frame along the frequency dimension, forms a frequency band gradient sequence, and extracts the frequency band index with the maximum gradient amplitude and the corresponding gradient value from the sequence, and finally splices them into a frequency band mutation vector to represent the local area with significant mutation in the spectral structure.

[0123] The output features of the two channels are spliced in the feature dimension to generate a perturbation enhancement feature vector corresponding to the frame, and the perturbation enhancement feature vector and the frame-level convolution feature vector output by the main channel convolution submodule are spliced along the channel dimension to form a fusion representation vector, which is input into the frame-level classification network as the basis for pseudo-sound classification. The frame-level classification network includes a fully connected layer, a normalization layer and a Softmax classification layer, and finally outputs a frame-level pseudo-sound state label vector consistent with the number of input frame sequences. This model structure can introduce explicit modeling of inter-frame perturbation and frequency local mutation while maintaining the spectral structure modeling capability, thereby improving the robustness and recognition accuracy of pseudo-sound interference.

[0124] In this embodiment, the Wav2Vec 2.0 model includes a time convolution encoder, a linear embedding layer, a Transformer context network and a speech recognition decoding head;

[0125] The residual voice signal is input into the time convolution encoder to extract the acoustic feature vector of each frame and generate a frame-level acoustic feature vector sequence;

[0126] The frame-level pseudo-sound state label vector is input into the linear embedding layer to be mapped into a frame-level pseudo-sound state embedding vector sequence;

[0127] The frame-level acoustic feature vector sequence and the frame-level artifact state embedding vector sequence are spliced frame by frame in feature dimensions to construct a joint feature sequence;

[0128] The joint feature sequence is input into the Transformer context network to extract a frame-level context representation vector sequence with long-distance context dependency;

[0129] The frame-level context representation vector sequence is input into a speech recognition decoding head, which includes a fully connected layer and a CTC decoding layer, to output a purified speech recognition result;

[0130] In the present application, the Wav2Vec 2.0 model includes a time convolutional encoder, a linear embedding layer, a Transformer context network, and a speech recognition decoding head, aiming to realize deep modeling and recognition of effective speech information in residual speech signals. The time convolutional encoder is composed of multiple one-dimensional convolutions, which is used for feature extraction in the time domain of the original residual speech signal, and outputs a frame-level acoustic feature vector sequence, which can effectively preserve the short-time acoustic mode information in the speech signal and enhance the local time structure perception ability. The linear embedding layer adopts a single fully connected structure, receives a frame-level artifact state label vector as input, and maps the discrete label to a continuous feature space with the same dimension as the acoustic feature vector through linear transformation, so that the artifact state information can be expressed together with the acoustic feature information.

[0131] After the frame-level acoustic feature vector sequence and the artifact state embedding vector sequence are spliced frame by frame in feature dimensions, a joint feature sequence is constructed. The joint feature sequence is input into the Transformer context network, which captures the long-distance context dependency relationship between speech frames through multiple layers of stacked self-attention mechanisms, enhances the semantic modeling capability, and enables the Wav2Vec 2.0 model to accurately locate effective speech information even in the presence of artifact interference. The number of attention heads, the dimension of the feedforward layer, and the position encoding structure in the Transformer context network are flexibly set according to actual precision and efficiency requirements.

[0132] In the present embodiment, the CTC decoding layer receives the frame-level character class prediction distribution output by the fully connected layer, performs path compression and blank symbol removal operations on each frame prediction result, generates a continuous character sequence, and outputs it as a purified speech recognition result;

[0133] In the present application, the speech recognition decoding head is composed of a fully connected layer and a CTC (Connectionist Temporal Classification) decoding layer, which is used to map the frame-level context representation vector sequence output by the Transformer context network into the final purified speech recognition result. The fully connected layer receives the context representation vector of each frame as input, performs linear transformation on the context representation vector through a set of trainable weight matrices, generates a probability distribution with a dimension equal to the size of the pre-defined character class set (such as Chinese characters, letters, numbers and special symbols), represents the prediction probability of each frame on all character classes, and constitutes a frame-level character class prediction distribution sequence.

[0134] The CTC decoding layer receives the above-mentioned frame-level character class prediction distribution sequence as input, and completes the prediction at the character level based on the alignment-free training and decoding mechanism. Specifically, the CTC decoding layer first performs path compression on the continuous repeated characters in the frame-level prediction path, that is, it combines the continuous same characters into a single character, which is used to eliminate redundant repeated labels; at the same time, for the frame positions predicted as blank labels, they are removed in the output sequence, thereby forming the final continuous character sequence. This decoding method does not require frame-level label and character alignment information, and can achieve efficient and accurate recognition under the condition that the input speech frame and the output text length are inconsistent. The final output continuous character sequence is the purified speech recognition result, representing the text information corresponding to the speech signal.

[0135] Reference Figure 2 A speech processing system of bone conduction microphone, comprising the following modules:

[0136] A speech signal acquisition and processing module for acquiring the original speech signal of the bone conduction microphone, performing frame slicing and frequency band decomposition, and generating a spectrum power spectrum;

[0137] A pseudo-sound candidate segment identification module for calculating the spectrum slope time sequence and the spectrum slope covariance matrix, identifying the time interval of multi-band slope synchronous mutation, and outputting it as a pseudo-sound candidate segment;

[0138] A pseudo-sound confirmation module for acquiring the electromyographic signal and the inertial measurement signal corresponding to the time of the pseudo-sound candidate segment, determining whether there is an electromyographic pulse mutation and an acceleration mutation, and if both exist, marking the pseudo-sound candidate segment as a confirmed pseudo-sound segment;

[0139] A frame state construction module for marking the silent frame, the speech frame and the confirmed pseudo-sound frame, and constructing a frame state sequence;

[0140] A residual speech extraction module for generating a residual speech signal and a pseudo-sound frame sequence;

[0141] The pseudo-sound modeling module is configured to perform a spectrum transformation on the pseudo-sound frame sequence, generate a pseudo-sound spectrogram frame sequence, and input the pseudo-sound spectrogram frame sequence into the improved RealFormer model to generate a frame-level pseudo-sound state label vector;

[0142] The speech recognition module is configured to input the residual speech signal and the frame-level pseudo-sound state label vector into a Wav2Vec 2.0 model to output a purified speech recognition result.

[0143] Embodiment 1:

[0144] To verify the feasibility of the application in implementation, the application is applied to a certain intelligent voice interaction terminal system to solve the problems of speech signal distortion, high recognition error rate, and difficulty in distinguishing pseudo-sound of the bone conduction microphone in a multi-source interference environment. The system is deployed in an intelligent head-mounted device, integrates a bone conduction microphone, an electromyography sensor, and an inertial measurement unit, can simultaneously collect speech signals and physiological motion data, and realizes multi-modal speech enhancement and recognition.

[0145] In the experimental scenario, the test object wears the device to perform voice interaction tests in three typical interference environments, i.e., indoor treadmill training, subway car riding, and outdoor cycling. The number of collected samples in each scenario is more than 200, each sample has a duration of about 10 to 15 seconds, the speech content is natural language instructions and digital short sentences, and the sampling rate is set to 16 kHz. In the experiment, the environmental noise intensity (A-weighted sound pressure level) fluctuates between 60-85 dB to ensure that the scenario is complex and representative.

[0146] First, the voice signal acquisition and processing module acquires the bone conduction original speech signal, frames the original speech signal in an 8-millisecond frame length and a 4-millisecond frame shift, decomposes the frequency band after a band-pass filter bank (50Hz-1500Hz), calculates the spectral energy density, and generates a spectral energy spectrogram. The spectrogram reveals the energy distribution characteristics of the bone conduction signal in different frequency bands. Subsequently, the pseudo-sound candidate segment recognition module performs a first-order difference operation on the spectral energy spectrogram to obtain a spectral slope time sequence, and detects a multi-band synchronous mutation interval based on a covariance matrix. The interval is marked as a pseudo-sound candidate segment.

[0147] In the case of obvious motion pseudo-sound, such as head shaking during cycling or jaw shaking during running, the mean value and change rate of the spectral slope covariance matrix will simultaneously increase. The application uses this feature to accurately identify the pseudo-sound occurrence period. Then, the pseudo-sound confirmation module introduces dual-channel verification of the electromyography signal and the inertial measurement signal. When an electromyography pulse mutation is detected and accompanied by an acceleration mutation, the pseudo-sound segment is confirmed and marked. The dual-signal joint detection mechanism effectively distinguishes between speech and non-speech motion states, avoiding the frequent misjudgments in traditional algorithms.

[0148] After confirming the pseudo-sound segment, the frame state construction module marks the frame-level signal as a mute frame, a speech frame or a confirmed pseudo-sound frame, and generates a frame state sequence. Then, the residual speech extraction module deletes the pseudo-sound frame according to the frame state sequence, retains only the continuous speech frames to form a residual speech signal, and synchronously extracts a pseudo-sound frame sequence. The short-time Fourier transform is performed on the pseudo-sound frame sequence to generate a pseudo-sound spectrogram frame sequence, which is input into the improved RealFormer model to model the pseudo-sound features through the perturbation alignment double-channel residual module, and output a frame-level pseudo-sound state label vector.

[0149] In the experiment, it is found that the improved RealFormer model improves the pseudo-sound classification accuracy on the training set by 7.8% compared with the standard Transformer model, and the response to the multi-band mutation feature is more stable. Finally, the residual speech signal and the pseudo-sound label vector are jointly input into the Wav2Vec 2.0 model, and after time convolution coding, Transformer context network analysis and CTC decoding, the final purified speech recognition result is output. The model parameters are set to 12 layers of Transformer layer, 768-dimensional hidden unit and 8 heads of self-attention mechanism, the training batch is 32, and the Adam optimizer is used for convergence.

[0150] In the test stage, in order to comprehensively verify the effect, the recognition performance under three conditions of original speech input, only band pass filtering processing and the method of the application is compared.

[0151] Table 1 Comparison of effects of different processing strategies in different experimental scenarios

[0152]

[0153] As can be seen from the data in the above table 1, in the three experimental scenarios, the method of the application performs superiorly in various indicators.

[0154] In the indoor treadmill training experimental scenario, the recognition accuracy of the original bone conduction speech signal is only 69.4%, the pseudo-sound detection rate is 55.3%, the false deletion rate is high, reaching 12.8%, the output signal-to-noise ratio is 9.1dB, and the word error rate (WER) is as high as 24.6%. Even after band pass filtering processing, although the recognition accuracy rises to 76.2% and the WER drops to 18.2%, there is still obvious pseudo-sound interference. In contrast, after processing by the method of the application, the recognition accuracy is significantly improved to 87.5%, the pseudo-sound detection rate is as high as 90.6%, the false deletion rate is reduced to 4.2%, the output signal-to-noise ratio is improved to 13.8dB, and the WER is also reduced to 11.3%, effectively improving the quality of speech recognition.

[0155] In the subway car, the recognition accuracy of the original signal is 73.5%, the pseudo-sound detection rate is only 61.2%, the false deletion rate is 11.5%, the output signal-to-noise ratio is 8.7dB, and the WER is 21.8%. After band-pass filtering processing, each index is slightly improved, but the improvement is limited. The recognition accuracy of the method of the application in this scene is 88.2%, the pseudo-sound detection rate is 92.4%, the false deletion rate is reduced to 4.5%, the output signal-to-noise ratio is improved to 13.2dB, and the WER is reduced to 11.7%, indicating that the application can effectively deal with the pseudo-sound problem in complex traffic environment.

[0156] In the outdoor riding environment, due to the influence of wind noise and head movement, the recognition accuracy of the original speech signal is only 71.1%, the pseudo-sound detection rate is 57.5%, the false deletion rate is 13.2%, the output signal-to-noise ratio is 8.9dB, and the WER is 23.1%. Although the band-pass filtering improves some indicators, it is not enough to completely suppress the pseudo-sound interference. The recognition accuracy of the method of the application in this scene is 86.3%, the pseudo-sound detection rate is improved to 89.7%, the false deletion rate is reduced to 4.8%, the output signal-to-noise ratio is improved to 13.5dB, and the WER is only 12.0%, showing excellent performance in strong interference motion scene.

[0157] The experimental results of the embodiment fully show that the application can effectively suppress non-speech disturbances in bone conduction signals by introducing spectral slope covariance detection, pseudo-sound detection and improved RealFormer model, realize high-fidelity speech purification, not only improve the speech recognition accuracy in complex noise environment, but also significantly improve the speech intelligibility and naturalness, and has high practical value and popularization potential.

[0158] The above describes only the preferred specific embodiments of the application, but the protection scope of the application is not limited thereto, and any skilled person in the art can make equivalent replacement or change according to the technical solution and inventive concept of the application within the technical range disclosed by the application, which should be covered within the protection scope of the application.

Claims

1. A voice processing method of a bone conduction microphone, characterized by, The method comprises the following steps: Step 1: Collecting the original speech signal of the bone conduction microphone and performing frame slicing, using a band-pass filter set to perform frequency band decomposition on each frame, calculating the spectral energy density value of each frequency band, and generating a spectral energy map; Step 2: Perform first-order difference operation on the spectral energy map to obtain a spectral slope time sequence, calculate the spectral slope covariance matrix between frequency bands, identify the time interval of multi-band slope synchronous mutation as the pseudo-sound candidate segment; Step 3: Collect the electromyographic signal and inertial measurement signal corresponding to the time of the pseudo-sound candidate segment, and judge whether there is an electromyographic pulse mutation and an acceleration mutation, if both exist, mark the pseudo-sound candidate segment as a confirmed pseudo-sound segment; Step 4: Based on the confirmed pseudo-sound segment, a frame state sequence is constructed, and the frames in the original speech signal are marked as silent frames, speech frames or confirmed pseudo-sound frames; Step 5: Delete the confirmed pseudo-sound frames in the frame state sequence to obtain a continuous speech segment composed of the remaining frames, which is defined as a residual speech signal; all deleted confirmed pseudo-sound frames constitute a pseudo-sound frame sequence; Step 6: Perform spectral transformation on the pseudo-sound frame sequence to generate a pseudo-sound map frame sequence, and input it into an improved RealFormer model to generate a frame-level pseudo-sound state label vector; Step 7: The residual speech signal and the frame-level pseudo-sound state label vector are jointly input into a Wav2Vec 2.0 model, and finally the purified speech recognition result is output.

2. The voice processing method of bone conduction microphone according to claim 1, wherein, The step one is specifically: Collecting the original speech signal of the bone conduction microphone, and performing frame slicing on the collected original speech signal of the bone conduction microphone in a frame length of 8 milliseconds and a frame shift of 4 milliseconds to generate a frame sequence; Input each frame in the frame sequence into a band-pass filter set with a set center frequency and bandwidth, the band-pass filter set includes no less than 16 filter channels, the center frequency covers the bone conduction effective speech frequency band from 50Hz to 1500Hz, and a frequency band signal set is output; Calculate the instantaneous energy value of each frequency band signal, which is the average value of the square of the amplitude of the frequency band signal sampling points in the current frame; Calculate the absolute mean of the first-order amplitude derivative in each frame as the mutation coefficient of the frequency band signal; Multiply the instantaneous energy value of each frequency band signal with the mutation coefficient to obtain the spectral energy density value of the corresponding frequency band; Splice all spectral energy density values of each frame into a spectral energy vector in frequency band order, and splice the spectral energy vectors of all frames in time order to obtain a spectral energy map.

3. The voice processing method of bone conduction microphone according to claim 1, wherein, The step two is specifically: Perform first-order difference operation on the spectral energy vectors of adjacent frames in the spectral energy map to obtain a spectral slope time sequence; In the spectral slope time sequence, extract the spectral slope vector corresponding to each frame, which is composed of the spectral energy density change values of each frequency band in the frame; Based on the spectral slope vector of each frame, calculate the covariance between the spectral energy density change values of any two frequency bands in the frame to construct a spectral slope covariance matrix; Using a time sequence sliding window method, the spectral slope covariance matrices corresponding to a plurality of continuous frames are statistically analyzed, and the average absolute value and the covariance change rate of each time sequence sliding window are calculated. The method has the advantages that the method can effectively identify the pseudo-sound segment in the original speech signal of the bone conduction microphone, and can effectively separate the pseudo-sound segment from the original speech signal of the bone conduction microphone, so as to obtain the residual speech signal and the pseudo-sound frame sequence, and then the purified speech recognition result can be output by using the Wav2Vec 2.0 model and the improved RealFormer model. When the average absolute value of the covariance exceeds a set average threshold, the covariance rate of change exceeds a set change rate threshold, and the number of frequency band pairs with the same change direction exceeds a preset proportion of the total number of frequency band pairs, it is determined that the time interval corresponding to the time sequence sliding window is a multi-band slope synchronous mutation segment; The time interval corresponding to the multi-band slope synchronous mutation segment is output as a pseudo-sound candidate segment.

4. The voice processing method of bone conduction microphone according to claim 1, wherein, The step three specifically includes: Extracting an electromyographic signal and an inertial measurement signal in the time interval corresponding to the pseudo-sound candidate segment, the electromyographic signal being an electrophysiological voltage sequence collected by a facial electromyographic sensor, and the inertial measurement signal being a three-axis acceleration time sequence collected by an acceleration sensor; Calculating an electromyographic mean square energy value in each sliding window for the electromyographic signal to obtain an electromyographic mean square energy value sequence; When the electromyographic mean square energy value in a sliding window exceeds a preset energy mutation threshold in the rising amplitude compared with a previous sliding window, and the amplitude difference value signs of a plurality of consecutive sampling points in the sliding window remain consistent, it is determined that there is an electromyographic pulse mutation in the sliding window; Performing sliding window maximum acceleration difference value operation on each axis of the inertial measurement signal respectively, and when the maximum acceleration difference value of any axis is greater than a set acceleration mutation threshold, it is determined that there is an acceleration mutation in the sliding window; When electromyographic pulse mutation and acceleration mutation are simultaneously detected in any sliding window corresponding to the pseudo-sound candidate segment, the pseudo-sound candidate segment is marked as a confirmed pseudo-sound segment.

5. The voice processing method of bone conduction microphone according to claim 1, wherein, The step four specifically includes: According to the time interval of the confirmed pseudo-sound segment, marking corresponding frames in the bone conduction microphone original speech signal as confirmed pseudo-sound frames; Among the frames that are not marked as confirmed pseudo-sound frames, marking continuous frames with an energy value lower than a set mute threshold and a frame number not less than a preset mute frame number as mute frames, and marking the remaining unmarked frames as speech frames; Constructing a frame state sequence arranged in time sequence, which is composed of all mute frames, speech frames and confirmed pseudo-sound frames.

6. The voice processing method of bone conduction microphone according to claim 1, wherein, The step of performing spectrum transformation on the pseudo-sound frame sequence to generate a pseudo-sound atlas frame sequence specifically includes: Performing short-time Fourier transformation on the audio signal corresponding to each frame in the pseudo-sound frame sequence to obtain a frequency spectrum of each frame; Dividing the frequency spectrum of each frame into a plurality of frequency band regions along the frequency axis, respectively calculating the frequency spectrum amplitude mean value in each frequency band region to form a frequency spectrum feature vector of the frame; Arranging the frequency spectrum feature vectors of all frames in time sequence to generate a pseudo-sound atlas frame sequence.

7. The method of claim 1, wherein, The improved RealFormer model specifically includes: In each Transformer layer of the original RealFormer model, a standard feedforward network is replaced by a perturbation alignment double-channel residual module, the perturbation alignment double-channel residual module including a main channel convolution submodule and a pseudo-sound perturbation enhancement submodule; The main channel convolution submodule includes two sequentially connected 1×1 convolution layers and a layer of grouped convolution layer with GELU activation function, which is used to extract the local frequency spectrum features of each frame in the pseudo-sound atlas frame sequence and output a frame-level convolution feature vector; The pseudo-sound perturbation enhancement submodule includes an inter-frame perturbation aggregation channel and a frequency band mutation detection channel; The inter-frame disturbance aggregation channel forms a symmetric three-frame window with the current frame and the adjacent frames before and after the current frame, respectively calculates the spectral amplitude vector difference between the current frame and the adjacent frames, and takes the average of the two spectral amplitude vector differences as the disturbance amplitude vector of the current frame; The frequency band mutation detection channel performs a first-order difference operation on the spectral feature vector of each frame along the frequency dimension to obtain a frequency band gradient sequence corresponding to the frame; in the frequency band gradient sequence, a plurality of frequency band indexes with the largest gradient amplitudes and corresponding gradient values are sequentially extracted, and the extracted frequency band gradient values are spliced in the order of the frequency band indexes to generate a frequency band mutation vector; The disturbance amplitude vector and the frequency band mutation vector are spliced in the feature dimension to generate a disturbance enhanced feature vector corresponding to the frame; The frame-level convolution feature vector and the disturbance enhanced feature vector are spliced along the channel dimension to generate a fusion representation vector; The fusion representation vector is input into a frame-level classification network, and the frame-level classification network includes a full connection layer, a normalization layer and a Softmax classification layer, and outputs a frame-level artifact state label vector consistent with the number of frames of the artifact map frame sequence.

8. The voice processing method of bone conduction microphone according to claim 1, wherein, The Wav2Vec 2.0 model includes a time convolution encoder, a linear embedding layer, a Transformer context network and a speech recognition decoding head; The residual speech signal is input into the time convolution encoder to extract an acoustic feature vector of each frame and generate a frame-level acoustic feature vector sequence; The frame-level artifact state label vector is input into the linear embedding layer to be mapped into a frame-level artifact state embedding vector sequence; The frame-level acoustic feature vector sequence and the frame-level artifact state embedding vector sequence are spliced in the feature dimension frame by frame to construct a joint feature sequence; The joint feature sequence is input into the Transformer context network to extract a frame-level context representation vector sequence with long-distance context dependence; The frame-level context representation vector sequence is input into the speech recognition decoding head, and the speech recognition decoding head includes a full connection layer and a CTC decoding layer, and outputs a purified speech recognition result.

9. The voice processing method of bone conduction microphone according to claim 8, wherein, The CTC decoding layer receives the frame-level character class prediction distribution output by the full connection layer, performs path compression and blank symbol removal operations on each frame prediction result, generates a continuous character sequence as a purified speech recognition result and outputs.

10. A speech processing system of a bone conduction microphone, performing the method of any one of claims 1 to 9, characterized in that, It comprises the following modules: A speech signal acquisition and processing module is used to acquire a bone conduction microphone original speech signal, perform frame slicing and frequency band decomposition, and generate a spectral quality energy map; An artifact candidate segment identification module is used to calculate a spectral slope time sequence and a spectral slope covariance matrix, identify a time interval of multi-frequency band slope synchronous mutation, and output as an artifact candidate segment; An artifact confirmation module is used to acquire an electromyographic signal and an inertial measurement signal corresponding to the time of the artifact candidate segment, determine whether there is an electromyographic pulse mutation and an acceleration mutation, and if both exist, mark the artifact candidate segment as a confirmed artifact segment; A frame state construction module is used to mark silent frames, speech frames and confirmed artifact frames, and construct a frame state sequence; A residual speech extraction module is used to generate a residual speech signal and an artifact frame sequence. The pseudo-tone modeling module is configured to perform a spectrum transformation on the pseudo-tone frame sequence, generate a pseudo-tone spectrogram frame sequence, and input the pseudo-tone spectrogram frame sequence into the improved RealFormer model to generate a frame-level pseudo-tone state label vector; The speech recognition module is configured to input the residual speech signal and the frame-level pseudo-tone state label vector into a Wav2Vec 2.0 model to output a purified speech recognition result.

Citation Information

Patent Citations

  • End-to-end bone and air conduction voice combined recognition method

    CN114495909A

  • Automated assessment of cognitive and speech motor impairment

    US20230172526A1