A noise-robust speech wake-up method, apparatus, device and medium

By jointly updating parameters, the feature enhancement network and the wake-up score acquisition network are optimized, which solves the reliability problem of voice wake-up in noisy environments and achieves efficient voice wake-up in noisy environments.

CN122511241APending Publication Date: 2026-08-04MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MALANSHAN AUDIO & VIDEO LABORATORY
Filing Date
2026-06-25
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, the optimization objectives of the enhancement module and the wake-up module are inconsistent in noisy environments, leading to a decrease in wake-up reliability.

Method used

By jointly updating the parameters of the initial feature enhancement network and the initial wake-up score acquisition network, the target feature enhancement network and the target wake-up score acquisition network are obtained. The noisy speech features are then input into the target feature enhancement network for feature enhancement processing to obtain the target enhanced speech features, which are then used as the sole input to the target wake-up score acquisition network.

Benefits of technology

This solves the problem of inconsistency between the optimization target and the wake-up discrimination target caused by independent training, avoids information loss and computational redundancy, and improves the reliability of voice wake-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511241A_ABST
    Figure CN122511241A_ABST
Patent Text Reader

Abstract

This application discloses a noise-resistant voice wake-up method, apparatus, device, and medium, relating to the field of speech signal processing. The method includes: jointly updating the parameters of an initial feature enhancement network and an initial wake-up score acquisition network to obtain a target feature enhancement network and a target wake-up score acquisition network; acquiring a noisy speech segment to be processed using a target data acquisition method, and extracting features from the noisy speech segment to obtain noisy speech features; inputting the noisy speech features into the target feature enhancement network to obtain target enhanced speech features; using the target enhanced speech features as the sole input to the target wake-up score acquisition network to obtain a target wake-up score, and waking up the device to be woken up based on the target wake-up score. By processing the target enhanced speech features through the target feature enhancement network and then using them as the sole input to the target wake-up score acquisition network, the problem of inconsistency between the enhancement optimization target and the wake-up discrimination target is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal processing, and in particular to a noise-resistant voice wake-up method, apparatus, device, and medium. Background Technology

[0002] Voice wake-up is a technology that activates devices or applications via voice commands, widely used in smart hardware, in-vehicle systems, and web applications. It works by listening to the user's voice input, recognizing specific wake-up words, and triggering corresponding actions. Voice wake-up needs to remain usable despite noise and other interference, and false wake-ups must be controlled.

[0003] To improve wake-up performance under noisy conditions, the industry often combines voice enhancement with wake-up recognition. In such combined solutions, the enhancement module and the wake-up module are often optimized according to their respective task objectives. Their requirements for feature representation may be inconsistent, and the enhancement process may change or lose time-frequency information that is beneficial to wake-up determination, thereby affecting wake-up reliability. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a noise-resistant voice wake-up method, apparatus, device, and medium, which can obtain target enhanced voice features through target feature enhancement network processing, and then use these features as the sole input to the target wake-up score acquisition network, thus avoiding the problem of inconsistency between the enhancement optimization target and the wake-up discrimination target. The specific solution is as follows: In a first aspect, this application provides a noise-resistant voice wake-up method, including: The initial feature enhancement network and the initial wake-up score acquisition network are jointly updated to obtain the corresponding target feature enhancement network and target wake-up score acquisition network. The noisy speech segment to be processed is obtained using the target data acquisition method, and the noisy speech segment is subjected to feature extraction to obtain the corresponding noisy speech features; The noisy speech features are input into the target feature enhancement network to perform feature enhancement processing on the noisy speech features, thereby obtaining the corresponding target enhanced speech features; The target enhanced speech features are used as the sole input to the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and the corresponding wake-up device is woken up by voice according to the target wake-up score.

[0005] Optionally, the joint parameter update of the initial feature enhancement network and the initial wake-up score acquisition network includes: Obtain a training sample set; each training sample in the training sample set includes a historical noisy speech segment, a historical noiseless speech segment corresponding to the historical noisy speech segment, and a wake-up tag used to characterize whether the historical noisy speech segment contains a preset wake-up word; Feature extraction is performed on the historical noisy speech segments and the historical noiseless speech segments respectively to obtain the corresponding noisy training features and noiseless training features; The noisy training features are input into the initial feature enhancement network to obtain the corresponding enhanced training features, and the enhanced training features are used as the sole input to the initial wake-up score acquisition network to obtain the training wake-up score corresponding to the noisy speech segment. A first loss is determined based on the difference between the enhanced training features and the noiseless training features, and a second loss is determined based on the training wake-up score and the wake-up label; The first loss and the second loss are weighted and fused to obtain the target total loss, and the initial feature enhancement network and the initial wake-up score acquisition network are jointly updated for several rounds based on the target total loss.

[0006] Optionally, the noisy speech feature is any one of spectral features, filter bank features, and Mel frequency cepstral coefficients. The feature extraction configuration of the noisy speech feature includes sampling rate, frame length, and frame shift. The feature extraction configuration is the configuration parameters used when extracting the noisy speech feature.

[0007] Optionally, the process by which the target feature enhancement network enhances the noisy speech features includes: Output the time-frequency mask corresponding to the noisy speech feature, and multiply the time-frequency mask by the noisy speech feature point by point to obtain the target enhanced speech feature; wherein the time-frequency mask and the noisy speech feature have the same number of time frames and feature dimension.

[0008] Optionally, the process of obtaining the target wake-up score from the network includes: Determine the frame-level wake-up score corresponding to each frame in the target enhanced speech features, and obtain the target wake-up score corresponding to the noisy speech segment based on each frame-level wake-up score.

[0009] Optionally, obtaining the target wake-up score corresponding to the noisy speech segment based on each of the frame-level wake-up scores includes: Determine the target weight value corresponding to each frame-level wake-up score, and perform a weighted summation of each frame-level wake-up score based on the target weight value to obtain the target wake-up score corresponding to the noisy speech segment.

[0010] Optionally, the step of performing voice wake-up on the corresponding device to be woken up based on the target wake-up score includes: Determine whether the target wake-up score is greater than a preset score threshold. If the corresponding determination result indicates that the target wake-up score is greater than the preset score threshold, then wake up the device to be woken up. If the corresponding judgment result indicates that the target wake-up score is not greater than the preset score threshold, then the device to be woken up is controlled to remain in standby mode.

[0011] Secondly, this application provides a noise-resistant voice wake-up device, comprising: The parameter update module is used to jointly update the parameters of the initial feature enhancement network and the initial wake-up score acquisition network to obtain the corresponding target feature enhancement network and target wake-up score acquisition network. The feature extraction module is used to acquire the noisy speech segment to be processed using the target data acquisition method, and to extract features from the noisy speech segment to obtain the corresponding noisy speech features. The feature enhancement module is used to input the noisy speech features into the target feature enhancement network to perform feature enhancement processing on the noisy speech features and obtain the corresponding target enhanced speech features; The device wake-up module is used to take the target enhanced speech features as the sole input to the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and to perform voice wake-up on the corresponding device to be woken up based on the target wake-up score.

[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned noise-resistant voice wake-up method.

[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the aforementioned noise-resistant voice wake-up method.

[0014] This application first performs joint parameter updates on the initial feature enhancement network and the initial wake-up score acquisition network to obtain the corresponding target feature enhancement network and target wake-up score acquisition network. Then, it uses a target data acquisition method to acquire the noisy speech segment to be processed and performs feature extraction on the noisy speech segment to obtain the corresponding noisy speech features. After that, the noisy speech features are input into the target feature enhancement network to perform feature enhancement processing on the noisy speech features to obtain the corresponding target enhanced speech features. Finally, the target enhanced speech features are used as the sole input of the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and the corresponding wake-up device is woken up by voice according to the target wake-up score. Therefore, this application solves the problem of inconsistent optimization targets between the initial feature enhancement network and the initial wake-up score acquisition network by jointly updating the parameters of the initial feature enhancement network and the initial wake-up score acquisition network, so that the loss function corresponding to the wake-up score can constrain the parameter optimization direction of the feature enhancement network. By processing the noisy speech features sequentially through the target feature enhancement network to obtain the target enhanced speech features, and then using them as the sole input of the target wake-up score acquisition network, this application solves the problem of information loss and computational redundancy introduced by the traditional scheme due to the cumbersome reconstruction and secondary extraction link of "feature-waveform-feature". Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of a noise-resistant voice wake-up method disclosed in this application; Figure 2 This application discloses a flowchart of a joint training process for a model. Figure 3 This application discloses a flowchart for obtaining a voice wake-up score; Figure 4 This is a schematic diagram of the structure of a noise-resistant voice wake-up device disclosed in this application; Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] See Figure 1 As shown, this embodiment of the invention discloses a noise-resistant voice wake-up method, including: Step S11: Perform joint parameter updates on the initial feature enhancement network and the initial wake-up score acquisition network to obtain the corresponding target feature enhancement network and target wake-up score acquisition network.

[0019] The core technical process for noise-resistant voice wake-up in this embodiment is as follows: The input speech feature X (i.e., noisy speech feature) first enters the feature enhancement model to obtain X′ (i.e., the target enhanced speech feature), and X′ is then used as the sole input to the wake-up recognition model to obtain the wake-up score S. Joint training: L_enh supervises the enhancement model with X_clean, and L_kws supervises the wake-up model with Y on X′. A single backpropagation updates both models simultaneously; L_kws gradient backpropagation enhances the model. Inference only takes X as input.

[0020] In this embodiment, the initial feature enhancement network and the initial wake-up score acquisition network are jointly updated in terms of parameters, including: acquiring a training sample set; each training sample in the training sample set includes a historical noisy speech segment, a historical noiseless speech segment corresponding to the historical noisy speech segment, and a wake-up tag used to characterize whether the historical noisy speech segment contains a preset wake-up word; extracting features from the historical noisy speech segment and the historical noiseless speech segment respectively to obtain corresponding noisy training features and noiseless training features; inputting the noisy training features into the initial feature enhancement network to obtain corresponding enhanced training features, and using the enhanced training features as the sole input of the initial wake-up score acquisition network to obtain the training wake-up score corresponding to the noisy speech segment; determining a first loss based on the difference between the enhanced training features and the noiseless training features, and determining a second loss based on the training wake-up score and the wake-up tag; performing a weighted fusion of the first loss and the second loss to obtain a target total loss, and performing several rounds of joint parameter updates on the initial feature enhancement network and the initial wake-up score acquisition network based on the target total loss.

[0021] Specifically, the order of forward computation, loss calculation, and backpropagation operations in a single batch of joint training is as follows: Step 1 (Sample Assembly): Take a batch of speech segments from the training set. Each sample contains noisy speech (i.e., historical noisy speech segments), clean reference speech (i.e. historical noiseless speech segments) that is time-aligned with it (usually obtained by adding noise to the same clean speech at a set signal-to-noise ratio), wake-up tags (positive samples contain wake-up words, and negative samples contain non-wake-up speech), and optional wake-up word end time annotations.

[0022] Step 2 (Feature Extraction): Extract features from noisy speech and clean reference speech using the exact same system configuration to obtain noisy feature X (i.e., noisy training feature) and clean feature X_clean (i.e., noise-free training feature); the configuration includes sampling rate, frame length, frame shift, feature type (such as FBank, MFCC or spectrum) and normalization method; both have the same number of frames and are time-aligned frame by frame.

[0023] That is, the noisy speech feature is any one of spectral features, filter bank features, and Mel frequency cepstral coefficients. The feature extraction configuration of the noisy speech feature includes sampling rate, frame length, and frame shift. The feature extraction configuration is the configuration parameters used when extracting the noisy speech feature.

[0024] Step 3 (Intra-batch Alignment): Since the duration of each sample in a batch is different, X and X_clean are aligned to the longest frame in the batch (e.g., padded with zeros), and the number of valid frames for each sample is recorded so that invalid frames can be masked during subsequent loss calculation.

[0025] Step 4 (Feature Enhancement Forward): Input the noisy feature X into the feature enhancement sub-network (i.e., the initial feature enhancement network), and output an enhanced feature X′ (i.e., the enhanced training feature) that is compatible with the shape of X; the preferred implementation is to output a time-frequency mask and multiply it point by point with X, so that X′ and X have the same number of time frames and feature dimension.

[0026] That is, in this embodiment, the process of the target feature enhancement network performing feature enhancement on the noisy speech features includes: outputting the time-frequency mask corresponding to the noisy speech features, and multiplying the time-frequency mask with the noisy speech features point by point to obtain the target enhanced speech features; wherein, the time-frequency mask and the noisy speech features have the same number of time frames and feature dimension.

[0027] Step 5 (Wake-up Recognition Forward): X′ is used as the sole input to the wake-up recognition sub-network. The frame-level wake-up response is obtained through temporal modeling and classification layers. Then, the segment-level wake-up score (i.e., training wake-up score) is obtained through the agreed temporal aggregation method (such as taking the maximum value, weighted summation, etc.), which is used for loss calculation or triggering decision.

[0028] Step 6 (Enhancement Loss L_enh): Within the effective frame range, calculate the reconstruction class loss L_enh (i.e., the first loss) between X′ and X_clean (such as mean square error, L1 loss, etc.). Only the effective frames are averaged, and the alignment padding region is not included in the calculation.

[0029] Step 7 (Wake-up Loss L_kws): Calculate the wake-up discrimination loss L_kws (i.e., the second loss) on X′, based on the wake-up label Y, constraining positive samples to have sufficiently high responses and negative samples to have sufficiently low responses within the target time period; the specific form can be cross-entropy, contrastive loss, or frame-level / segment-level aggregation loss, etc.

[0030] Step 8 (Total Loss): By Weighted total loss; The coefficients can be fixed or determined by optimization based on the validation set (see the loss balance explanation below).

[0031] Step 9 (Backpropagation): Perform backpropagation on the total loss, and simultaneously backpropagate the gradient to the wake-up recognition subnetwork and the feature enhancement subnetwork; among them, the gradient of L_kws with respect to X′ continues to be backpropagated to the enhancement subnetwork, so that the enhancement behavior is directly constrained by the wake-up target.

[0032] Step 10 (Parameter Update): Perform gradient pruning (optional) and optimizer update on the trainable parameters of the two sub-networks. After completing the parameter iteration of this batch, proceed to the next batch.

[0033] Step 11 (Verification / Inference Difference): During the verification and online inference stages, no clean reference feature X_clean is provided. Only noisy feature extraction → feature enhancement → wake-up determination → threshold comparison are performed. L_enh is not calculated, and backpropagation and parameter updates are not performed.

[0034] In addition, the core process of offline model training in this embodiment is as follows: Figure 2 As shown, the core lies in utilizing The total loss is determined, and the joint parameters of the feature enhancement model and the wakefulness recognition model are updated based on the total loss. In this embodiment, the specific process for offline model training is as follows: Prepare training corpus—collect positive wake-up words and negative non-wake-up words; synthesize noisy speech from clean speech with various noise types (such as environmental noise, music, human voice, etc.) at random signal-to-noise ratio, and retain time-aligned clean reference speech for enhanced supervision; Feature extraction—Features are extracted from noisy speech and clean reference speech respectively to obtain X and X_clean; training and validation use a unified feature configuration and ensure frame-by-frame temporal alignment; Batch training—combines X, X_clean, wake-up label and valid frame information into training batches, aligns them according to the longest frame length in the batch, and then feeds them into the model; Forward computation—X passes through the feature enhancement sub-network to obtain X′, X′ then passes through the wake-up recognition sub-network to obtain the wake-up response and calculate L_kws; at the same time, L_enh is calculated based on X_clean; Reverse update — press The total loss is synthesized, backpropagated, and the parameters of the two sub-networks are updated; gradient pruning can be optionally performed to improve training stability. Iterate until convergence—monitor the noisy wake-up rate and false wake-up rate on the validation set, and select the optimal model for deployment.

[0035] As can be seen from the above process, this embodiment performs joint parameter updates on the initial feature enhancement network and the initial wake-up score acquisition network, rather than training and updating the two networks separately.

[0036] Additionally, it should be noted that in step 3 above, the strict temporal alignment of the noisy feature X and the clean feature X_clean is ensured through the following method: (a) Homologous speech alignment: X and X_clean are time-aligned speech pairs from the same training sample. Noisy speech is usually obtained by mixing clean reference speech and noise according to the target signal-to-noise ratio, with the two having the same sampling rate, duration, and start and end times.

[0037] (b) Feature extraction with the same configuration: The same feature extraction parameters (frame length, frame shift, filter bank, pre-emphasis, normalization, etc.) are used for noisy speech and clean reference speech. Since the input speech is of the same length and has the same starting point, the number of output feature frames is the same.

[0038] (c) One-to-one correspondence frame by frame: the feature X[t] of frame t corresponds to the same original time segment as X_clean[t]; when the frame length is aligned within the batch, X and X_clean share the same effective frame range, ensuring that L_enh is calculated only between frame pairs at the same time.

[0039] (d) Enhance policy consistency: If data augmentation such as speed perturbation and volume change is performed on noisy speech during training, the clean reference speech and wake-up label will undergo the same transformation synchronously to avoid feature and label misalignment.

[0040] Furthermore, the balancing strategy between L_enh and L_kws during model training is as follows: (a) Fixed weighting (commonly used): Preset It is a constant, such as taking 1 for both, i.e., L = L_enh + L_kws; applicable to cases where the magnitudes of the two losses are close.

[0041] (b) Validation set parameter tuning (optional): Parameter tuning on the validation set... Perform grid search or Bayesian optimization to select the optimal combination with the wake-up rate and false wake-up rate under noisy conditions, and keep it unchanged during training and deployment after selection.

[0042] (c) Dynamic balancing (optional extension): 1. Multi-task uncertainty weighting: learn trainable weights for two tasks and automatically adjust the loss contribution; 2. Gradient normalization: scale according to the norm of the two loss gradients and then merge and update; 3. Staged training: focus on L_enh to stabilize and enhance the representation in the early stage, and increase the weight of L_kws to strengthen wake-up discrimination in the later stage.

[0043] (d) Scale matching suggestion: If the magnitudes of L_enh and L_kws differ significantly, loss normalization, logarithmic transformation, or adjustment can be used to match the magnitudes. To balance the two contributions and prevent one loss from dominating training and weakening the joint optimization effect.

[0044] Step S12: Obtain the noisy speech segment to be processed using the target data acquisition method, and extract features from the noisy speech segment to obtain the corresponding noisy speech features.

[0045] As can be seen from the foregoing, the noisy speech features in this embodiment are any one of spectral features, filter bank features, and Mel frequency cepstral coefficients. The feature extraction configuration of the noisy speech features includes sampling rate, frame length, and frame shift.

[0046] In this embodiment, the method for obtaining the noisy speech segment to be processed can be determined by speech activity detection, fixed analysis window, or business logic. The specific method for obtaining the speech segment is not specifically limited here.

[0047] The above process of feature extraction for noisy speech segments uses the same feature extraction configuration as the training phase, and will not be described again here.

[0048] Step S13: Input the noisy speech features into the target feature enhancement network to perform feature enhancement processing on the noisy speech features and obtain the corresponding target enhanced speech features.

[0049] In this embodiment, the process of the target feature enhancement network enhancing the noisy speech features includes: outputting the time-frequency mask corresponding to the noisy speech features, and multiplying the time-frequency mask with the noisy speech features point by point to obtain the target enhanced speech features; wherein, the time-frequency mask and the noisy speech features have the same number of time frames and feature dimension.

[0050] The feature enhancement process described above is consistent with the training phase process. That is, the noisy feature X is input into the feature enhancement subnetwork, and the enhanced feature X′ with the shape of X is output; the preferred implementation is to output a time-frequency mask and multiply it point by point with X, so that X′ and X have the same number of time frames and feature dimension.

[0051] Step S14: Use the target enhanced speech features as the sole input to the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and perform voice wake-up on the corresponding device to be woken up according to the target wake-up score.

[0052] In this embodiment, the process of obtaining the target wake-up score from the network output includes: determining the frame-level wake-up score corresponding to each frame in the target enhanced speech features, and obtaining the target wake-up score corresponding to the noisy speech segment based on each frame-level wake-up score.

[0053] The step of obtaining the target wake-up score corresponding to the noisy speech segment based on each frame-level wake-up score includes: determining the target weight value corresponding to each frame-level wake-up score, and performing a weighted summation of each frame-level wake-up score according to the target weight value to obtain the target wake-up score corresponding to the noisy speech segment.

[0054] In addition, the voice wake-up of the corresponding device to be woken up based on the target wake-up score includes: determining whether the target wake-up score is greater than a preset score threshold; if the determination result indicates that the target wake-up score is greater than the preset score threshold, then the device to be woken up is woken up; if the determination result indicates that the target wake-up score is not greater than the preset score threshold, then the device to be woken up is controlled to remain in standby mode.

[0055] The core process of online inference in this embodiment is as follows: Figure 3 As shown, this includes inputting speech features X into a feature enhancement model to obtain enhanced features X′, and then inputting enhanced features X′ into a wake-up recognition model to obtain a wake-up score S. Specifically, the detailed process of online model inference is as follows: Acquiring speech segments—the speech segments to be analyzed are determined by speech activity detection, fixed analysis windows, or business logic; Extracting noisy features X—using the same feature extraction configuration as the training phase; Feature enhancement—X is enhanced by outputting enhanced feature X′ through a feature enhancement subnetwork; Wake-up determination—X′ obtains a frame-level response (i.e., frame-level wake-up score) through the wake-up recognition sub-network, and then obtains a segment-level wake-up score S (i.e., target wake-up score) through temporal aggregation; Triggering decision—If S exceeds the preset threshold, it determines to wake up and trigger subsequent services; otherwise, it remains in standby mode. The entire process is completed in the feature domain, without the need for waveform reconstruction, secondary feature extraction, or independent general noise reduction modules.

[0056] It should be noted that the noisy speech feature X and the enhanced speech feature X′ in this embodiment are shape compatible. The shape compatibility mentioned above specifically refers to: (a) Basic definition: X′ and X have the same number of time frames and feature dimension, that is, the feature vector dimension remains unchanged at each time step before and after enhancement; the wake-up model can directly receive X′ without additional dimension transformation, interpolation or adaptation layer.

[0057] (b) Preferred implementation: The feature enhancement subnetwork outputs a time-frequency mask of the same size as X (the value is usually constrained to 0 to 1). The enhanced features are obtained by multiplying the mask with X point by point. Therefore, X′ has the same number of frames and dimension as X.

[0058] (c) Permitted extended forms: In addition to mask multiplication, residual mapping (X′= X + ...) is also permitted. Nonlinear mappings (X′= G(X)) or gated residuals, etc., are all considered shape compatible as long as the output is still a feature sequence with the same number of frames and dimensions as X.

[0059] (d) Different sizes: If the enhancement network changes the temporal resolution (such as reducing the number of frames due to undersampling) or changes the feature dimension, the output X′ cannot be trained by the model to obtain the label X_clean (i.e. the clean feature after denoising), which is not compatible with the present invention.

[0060] (e) Intra-batch alignment: When aligning frame lengths within a batch, the padding frames of X, X′, and X_clean are located at the same time position, the effective frame range is consistent, and the padding area does not participate in the calculation of L_enh and L_kws.

[0061] Compared to the previous process of feature extraction → reconstructing denoised speech → extracting features again → wake-up, this invention directly obtains X′ from the enhancement model in the feature domain and feeds it into the wake-up model, avoiding waveform reconstruction and secondary feature extraction. X′ is the only input to the wake-up model; it can adapt to various speech features such as spectrum, FBank, and MFCC.

[0062] In addition, in this embodiment, the enhancement model and the wake-up model are optimized using the same computation graph, and the L_kws gradient is backpropagated to enhance the model, thereby enhancing automatic adaptive wake-up.

[0063] Therefore, this application solves the problem of inconsistent optimization targets between the initial feature enhancement network and the initial wake-up score acquisition network by jointly updating the parameters of the initial feature enhancement network and the initial wake-up score acquisition network, so that the loss function corresponding to the wake-up score can constrain the parameter optimization direction of the feature enhancement network. By processing the noisy speech features sequentially through the target feature enhancement network to obtain the target enhanced speech features, and then using them as the sole input of the target wake-up score acquisition network, this application solves the problem of information loss and computational redundancy introduced by the traditional scheme due to the cumbersome reconstruction and secondary extraction link of "feature-waveform-feature".

[0064] See Figure 4 As shown, an embodiment of the present invention discloses a noise-resistant voice wake-up device, comprising: Parameter update module 11 is used to jointly update the parameters of the initial feature enhancement network and the initial wake-up score acquisition network to obtain the corresponding target feature enhancement network and target wake-up score acquisition network; The feature extraction module 12 is used to acquire the noisy speech segment to be processed using the target data acquisition method, and to extract features from the noisy speech segment to obtain the corresponding noisy speech features. The feature enhancement module 13 is used to input the noisy speech features into the target feature enhancement network to perform feature enhancement processing on the noisy speech features and obtain the corresponding target enhanced speech features; The device wake-up module 14 is used to take the target enhanced speech features as the only input to the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and to perform voice wake-up on the corresponding device to be woken up according to the target wake-up score.

[0065] In some specific embodiments, the parameter update module 11 may specifically include: The sample set acquisition unit is used to acquire a training sample set; each training sample in the training sample set includes a historical noisy speech segment, a historical noiseless speech segment corresponding to the historical noisy speech segment, and a wake-up tag used to characterize whether the historical noisy speech segment contains a preset wake-up word; The feature extraction unit is used to extract features from the historical noisy speech segments and the historical noiseless speech segments respectively, so as to obtain the corresponding noisy training features and noiseless training features. The scoring unit is used to input the noisy training features into the initial feature enhancement network to obtain the corresponding enhanced training features, and to use the enhanced training features as the only input of the initial wake-up score acquisition network to obtain the training wake-up score corresponding to the noisy speech segment. The loss determination unit is used to determine a first loss based on the difference between the enhanced training features and the noiseless training features, and to determine a second loss based on the training wake-up score and the wake-up label; The parameter update unit is used to perform weighted fusion of the first loss and the second loss to obtain the target total loss, and to perform several rounds of joint parameter updates on the initial feature enhancement network and the initial wake-up score acquisition network based on the target total loss.

[0066] In some specific embodiments, the target feature enhancement network may specifically include: The mask output unit is used to output the time-frequency mask corresponding to the noisy speech feature, and multiply the time-frequency mask by the noisy speech feature point by point to obtain the target enhanced speech feature; wherein the time-frequency mask and the noisy speech feature have the same number of time frames and feature dimension.

[0067] In some specific embodiments, the target wake-up score acquisition network may specifically include: The scoring determination submodule is used to determine the frame-level wake-up score corresponding to each frame in the target enhanced speech features, and to obtain the target wake-up score corresponding to the noisy speech segment based on each frame-level wake-up score.

[0068] In some specific embodiments, the score determination submodule may specifically include: The weight value determination unit is used to determine the target weight value corresponding to each of the frame-level wake-up scores, and to perform a weighted summation of each of the frame-level wake-up scores according to the target weight value to obtain the target wake-up score corresponding to the noisy speech segment.

[0069] In some specific embodiments, the device wake-up module 14 may specifically include: The device wake-up unit is used to determine whether the target wake-up score is greater than a preset score threshold. If the corresponding determination result indicates that the target wake-up score is greater than the preset score threshold, then the device to be woken up is woken up. The device standby unit is used to control the device to be woken up to remain in standby mode if the corresponding judgment result indicates that the target wake-up score is not greater than the preset score threshold.

[0070] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0071] Figure 5This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the noise-resistant voice wake-up method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0072] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0073] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0074] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the noise-resistant voice wake-up method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0075] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned noise-resistant voice wake-up method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0076] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0077] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0078] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0079] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0080] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A noise-resistant voice wake-up method, characterized in that, include: The initial feature enhancement network and the initial wake-up score acquisition network are jointly updated to obtain the corresponding target feature enhancement network and target wake-up score acquisition network. The noisy speech segment to be processed is obtained using the target data acquisition method, and the noisy speech segment is subjected to feature extraction to obtain the corresponding noisy speech features; The noisy speech features are input into the target feature enhancement network to perform feature enhancement processing on the noisy speech features, thereby obtaining the corresponding target enhanced speech features; The target enhanced speech features are used as the sole input to the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and the corresponding wake-up device is woken up by voice according to the target wake-up score.

2. The noise-resistant voice wake-up method according to claim 1, characterized in that, The joint parameter update of the initial feature enhancement network and the initial wake-up score acquisition network includes: Obtain a training sample set; each training sample in the training sample set includes a historical noisy speech segment, a historical noiseless speech segment corresponding to the historical noisy speech segment, and a wake-up tag used to characterize whether the historical noisy speech segment contains a preset wake-up word; Feature extraction is performed on the historical noisy speech segments and the historical noiseless speech segments respectively to obtain the corresponding noisy training features and noiseless training features; The noisy training features are input into the initial feature enhancement network to obtain the corresponding enhanced training features, and the enhanced training features are used as the sole input to the initial wake-up score acquisition network to obtain the training wake-up score corresponding to the noisy speech segment. A first loss is determined based on the difference between the enhanced training features and the noiseless training features, and a second loss is determined based on the training wake-up score and the wake-up label; The first loss and the second loss are weighted and fused to obtain the target total loss, and the initial feature enhancement network and the initial wake-up score acquisition network are jointly updated for several rounds based on the target total loss.

3. The noise-resistant voice wake-up method according to claim 1, characterized in that, The noisy speech feature is any one of spectral features, filter bank features, and Mel frequency cepstral coefficients. The feature extraction configuration of the noisy speech feature includes sampling rate, frame length, and frame shift. The feature extraction configuration is the configuration parameters used when extracting the noisy speech feature.

4. The noise-resistant voice wake-up method according to claim 1, characterized in that, The process by which the target feature enhancement network enhances the noisy speech features includes: Output the time-frequency mask corresponding to the noisy speech feature, and multiply the time-frequency mask by the noisy speech feature point by point to obtain the target enhanced speech feature; wherein the time-frequency mask and the noisy speech feature have the same number of time frames and feature dimension.

5. The noise-resistant voice wake-up method according to claim 1, characterized in that, The process of obtaining the target wake-up score from the network output includes: Determine the frame-level wake-up score corresponding to each frame in the target enhanced speech features, and obtain the target wake-up score corresponding to the noisy speech segment based on each frame-level wake-up score.

6. The noise-resistant voice wake-up method according to claim 5, characterized in that, The step of obtaining the target wake-up score corresponding to the noisy speech segment based on each of the frame-level wake-up scores includes: Determine the target weight value corresponding to each frame-level wake-up score, and perform a weighted summation of each frame-level wake-up score based on the target weight value to obtain the target wake-up score corresponding to the noisy speech segment.

7. The noise-resistant voice wake-up method according to claim 1, characterized in that, The step of voice-waking the corresponding device to be woken up based on the target wake-up score includes: Determine whether the target wake-up score is greater than a preset score threshold. If the corresponding determination result indicates that the target wake-up score is greater than the preset score threshold, then wake up the device to be woken up. If the corresponding judgment result indicates that the target wake-up score is not greater than the preset score threshold, then the device to be woken up is controlled to remain in standby mode.

8. A noise-resistant voice wake-up device, characterized in that, include: The parameter update module is used to jointly update the parameters of the initial feature enhancement network and the initial wake-up score acquisition network to obtain the corresponding target feature enhancement network and target wake-up score acquisition network. The feature extraction module is used to acquire the noisy speech segment to be processed using the target data acquisition method, and to extract features from the noisy speech segment to obtain the corresponding noisy speech features. The feature enhancement module is used to input the noisy speech features into the target feature enhancement network to perform feature enhancement processing on the noisy speech features and obtain the corresponding target enhanced speech features; The device wake-up module is used to take the target enhanced speech features as the sole input to the target wake-up score acquisition network to obtain the target wake-up score corresponding to the noisy speech segment, and to perform voice wake-up on the corresponding device to be woken up based on the target wake-up score.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the noise-resistant voice wake-up method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the noise-resistant voice wake-up method as described in any one of claims 1 to 7.