Wakeup word recognition method and device based on double-word joint detection, equipment and medium
By employing a wake word recognition method based on dual-word joint detection, and utilizing time warping and dual-word joint detection algorithms, the problems of low accuracy and high latency in existing wake word recognition technologies are solved, thereby improving recognition accuracy and response efficiency. This method is suitable for resource-constrained embedded systems and mobile devices.
Patent Information
- Application Number
- CN202411224954.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-05-11
AI Technical Summary
Existing wake word recognition technologies suffer from low accuracy and high latency in diverse environments, making them difficult to apply effectively in resource-constrained embedded systems or mobile devices.
A wake-up word recognition method based on dual-word joint detection is adopted. By acquiring pre-recorded audio segments of the target object and audio segments to be detected in various care scenarios, key feature information is extracted, and the target wake-up word is identified by using time warping algorithm and dual-word joint detection algorithm.
It improves the accuracy and response efficiency of wake word recognition, reduces the false alarm rate, and is suitable for environments requiring high sensitivity and low false alarm rate, such as childcare and security monitoring.
Smart Images

Figure CN118898991B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on May 11, 2024, entitled "A method, apparatus, device and medium for recognizing wake words in conjunction with dynamic time regularization", application number 202410580558.5. Technical Field
[0002] This invention relates to the field of audio processing technology, and in particular to a wake word recognition method, apparatus, device and medium based on dual-word joint detection. Background Technology
[0003] With the rise of intelligent voice assistants, voice control systems, and smart home devices, the demand for natural language and voice interaction continues to increase. Under this trend, wake word recognition has become a key technology for realizing natural voice interaction, allowing users to wake up devices and input voice commands with simple voice commands. Furthermore, energy efficiency and resource management have become particularly critical in mobile devices and embedded systems. By adopting wake word technology, the system can listen to ambient sounds in standby or low-power mode, waking up the more complex voice processing system only when a specific wake word is detected, effectively reducing power consumption and resource usage. With the widespread adoption of virtual assistants such as Apple's Siri, Amazon's Alexa, and Google Assistant, wake word recognition has become especially crucial. This technology allows these virtual assistants to wait for user commands in the background, and users can trigger various voice functions, including voice search, reminders, and music playback, simply with a voice command. In the real world, complex noisy environments, such as traffic noise and human voices, pose challenges to voice processing systems. Wake word recognition needs to be robust, capable of accurately detecting wake words in noisy environments to ensure the reliability of user experience. With the continuous development of deep learning technology, wake word recognition has been significantly improved. Deep learning methods are widely used in wake word recognition, greatly improving the accuracy and robustness of the model. In general, the advancement of wake word recognition technology not only meets users' urgent needs for natural voice interaction, but also solves a series of technical challenges in speech processing, promoting the development of the entire speech recognition field.
[0004] With the development of deep learning technology, neural networks have become increasingly popular in wake word recognition. Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs), and variants such as Gated Recurrent Units (GRUs) are all used to model speech features and sequences. However, firstly, existing wake word recognition technologies perform poorly in diverse environments. Variations in environmental factors such as noise levels, sound reflections, and spatial size can affect the accuracy of wake words, and existing wake word recognition technologies lack robustness when adapting to complex noisy environments and different recording conditions. Secondly, existing wake word recognition technologies are slow when processing large-scale speech data, especially on resource-constrained embedded systems or mobile devices. Slow recognition speeds lead to response time delays, reducing user experience.
[0005] Chinese patent CN107886957A discloses a voice wake-up method and apparatus combining voiceprint recognition. The method includes: S1: receiving a voice to be verified and extracting features to obtain the MFCC features of the voice to be verified; S2: caching the MFCC features of the voice to be verified within a preset time period; S3: determining whether the content of the voice to be verified is a preset wake-up word based on the cached MFCC features, and if so, proceeding to step S4; S4: inputting the cached MFCC features of the voice to be verified into a preset deep neural network model to obtain the i-vector vector of the voice to be verified; S5: combining the i-vector vector of the voice to be verified with a preset i-ve The ctor vector is compared, and the permission value of the voice to be verified is obtained based on the matching score obtained from the comparison. It is determined whether the permission value of the voice to be verified is greater than or equal to the permission value corresponding to the preset wake word of the voice to be verified. If so, the operation corresponding to the preset wake word of the voice to be verified is executed. Although the above patent discloses the acquisition of audio data, extraction of key features from the audio data, and identification of target keywords based on the key features, it is easily affected by the external environment when facing the diverse environment in reality, resulting in low accuracy of wake word recognition. At the same time, since the above patent has not made much improvement in recognition speed, the conventional high-latency response speed is difficult to apply to resource-constrained embedded systems or mobile devices.
[0006] Therefore, how to provide a low-latency, high-accuracy wake word recognition method that can be applied to resource-constrained embedded systems or mobile devices is an urgent problem to be solved. Summary of the Invention
[0007] In view of this, the present invention provides a wake word recognition method, device, equipment and medium based on dual-word joint detection to solve the problems of low accuracy and high latency in dual wake word recognition in the prior art.
[0008] The technical solution adopted in this invention is:
[0009] In a first aspect, the present invention provides a wake-word recognition method based on dual-word joint detection, the method comprising:
[0010] The system acquires pre-recorded audio segments of the target object and key feature information extracted from audio segments to be detected in various care scenarios, wherein the target object includes at least children and unauthorized users;
[0011] Feature extraction is performed on the pre-recorded audio segment to obtain feature template information corresponding to the target object;
[0012] The key feature information and the feature template information are matched. If the match is successful, the key feature information is removed.
[0013] If the matching fails, the audio frame corresponding to the key feature information is divided into multiple audio segments based on the key feature information.
[0014] Based on the amplitude difference between each audio segment, the step size of the algorithm detection within each audio segment is adjusted, and the target step size corresponding to each audio segment is output.
[0015] Based on each target step size, the time warping algorithm is used to perform similarity matching on each target audio segment within each audio segment at the target step size interval, and output the first target audio segment and the second target audio segment with a similarity greater than the preset similarity threshold.
[0016] The first target audio segment and the second target audio segment are identified using a two-word joint detection algorithm to identify the target wake word. The two-word joint detection algorithm is a target detection algorithm that detects two related words simultaneously.
[0017] Preferably, before acquiring the pre-recorded audio segment of the target object and the key feature information extracted from the audio segment to be detected in various care scenarios, the following steps are included:
[0018] Acquire real-time audio data in various caregiving scenarios;
[0019] Using a multi-feature fusion algorithm, the real-time audio data is subjected to silent detection, the detected silent audio segments are removed, and the speech audio segments are output.
[0020] Based on the preset duration threshold corresponding to the wake word, the corresponding audio segments with a duration longer than the duration threshold are removed from the speech audio segments, and the audio segments to be detected are output.
[0021] Spectral features are extracted from the audio segment to be detected to obtain key feature information related to the preset target wake word.
[0022] Preferably, the step of using a multi-feature fusion algorithm to perform silence recognition on the real-time audio data, removing the identified silent audio segments, and outputting speech audio segments includes:
[0023] Obtain the original audio signal corresponding to each audio frame in the real-time audio data, extract signal features from the original audio signal, and obtain multiple audio signal features;
[0024] The audio signal features are normalized, and the standard audio signal features corresponding to each audio signal feature are output.
[0025] Based on preset weights, the features of each standard audio signal are weighted and fused to output a fused feature vector;
[0026] The fused feature vector is input into a pre-trained classifier, which outputs the probability that the fused feature vector is identified as a speech audio segment.
[0027] The speech audio segment is identified based on the probability of the fused feature vector being identified as a speech audio segment and a preset probability threshold.
[0028] Preferably, the step of identifying the speech audio segment based on the probability of recognizing it as a speech audio segment according to the fused feature vector and a preset probability threshold includes:
[0029] Obtain the probability of each fused feature vector being identified as a speech audio segment and the preset smoothing factor;
[0030] Based on the smoothing factor, the probability of each fused feature vector being identified as a speech audio segment is smoothed, and the target probability of each fused feature vector being identified as a speech segment after smoothing is output.
[0031] The target probabilities are compared with the probability thresholds. When the target probability is greater than the probability threshold, the audio frame corresponding to the target probability is identified as the speech audio segment.
[0032] Preferably, the step of using the dual-word joint detection algorithm to identify the first target audio segment and the second target audio segment, and identifying the target wake word, includes:
[0033] The first target audio segment and the second target audio segment are filtered by time interval, and the first qualified audio segment and the second qualified audio segment that pass the filter are output.
[0034] Perform wake word recognition on the first and second qualified audio segments respectively, and output the recognition results;
[0035] Based on the recognition results, if the target wake word is recognized in both the first qualified audio segment and the second qualified audio segment, then the target wake word recognition is considered complete.
[0036] Preferably, the step of filtering the first target audio segment and the second target audio segment by time interval, and outputting the first qualified audio segment and the second qualified audio segment that have passed the filtering, includes:
[0037] Obtain the time interval and preset time window between the first target audio segment and the second target audio segment;
[0038] The time interval is compared with the time window. If the time interval is determined to be within the time window, the first qualified audio segment and the second qualified audio segment are output.
[0039] If the time interval is not within the time window, the corresponding first target audio segment and second target audio segment are removed.
[0040] Preferably, the wake-word recognition is performed on the first qualified audio segment and the second qualified audio segment respectively, and the output recognition result includes:
[0041] Acquire first training audio data and second training audio data, wherein the first training audio data includes the target wake word, and the second training audio data does not include the target wake word;
[0042] The first and second training audio data are labeled, and the label information is output.
[0043] Feature extraction is performed on the first training audio data and the second training audio data to output training feature information;
[0044] The label information and the training feature information are input into a preset classification model for training to obtain a wake word recognition model;
[0045] The first qualified audio segment and the second qualified audio segment are respectively input into the wake word recognition model, and the recognition results are output.
[0046] Secondly, the present invention provides a wake-up word recognition device based on dual-word joint detection, the device comprising:
[0047] The key feature information acquisition module is used to acquire key feature information extracted from pre-recorded audio segments of the target object and audio segments to be detected in various care scenarios, wherein the target object includes at least children and unauthorized users;
[0048] The feature extraction module is used to extract features from the pre-recorded audio segment and obtain feature template information corresponding to the target object;
[0049] The feature matching module is used to perform feature matching between the key feature information and the feature template information. If the match is successful, the key feature information is removed.
[0050] The audio frame segmentation module is used to divide the audio frame corresponding to the key feature information into multiple audio segments based on the key feature information if the matching fails.
[0051] The step size adjustment module is used to adjust the step size detected by the algorithm within each audio segment based on the amplitude difference between each audio segment, and output the target step size corresponding to each audio segment.
[0052] The similarity matching module is used to perform similarity matching on each target audio segment within each audio segment according to each target step size and using a time warping algorithm, and output the first target audio segment and the second target audio segment with a similarity greater than a preset similarity threshold.
[0053] The wake-up word recognition module is used to identify the first target audio segment and the second target audio segment using a two-word joint detection algorithm, and to identify the target wake-up word. The two-word joint detection algorithm is a target detection algorithm that detects two related words simultaneously.
[0054] Thirdly, embodiments of the present invention also provide an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect described above.
[0055] Fourthly, embodiments of the present invention also provide a storage medium storing computer program instructions thereon, which, when executed by a processor, implement the method of the first aspect described above.
[0056] In summary, the beneficial effects of the present invention are as follows:
[0057] This invention provides a wake-word recognition method, apparatus, device, and medium based on dual-word joint detection. The method includes: acquiring pre-recorded audio segments of a target object and key feature information extracted from audio segments to be detected in various care scenarios, wherein the target object includes at least children and unauthorized users; extracting features from the pre-recorded audio segments to obtain feature template information corresponding to the target object; performing feature matching between the key feature information and the feature template information; if the matching is successful, removing the key feature information; if the matching fails, dividing the audio frame corresponding to the key feature information into multiple audio segments based on the key feature information. Based on the amplitude difference between each audio segment, the step size of the algorithm detection within each audio segment is adjusted, and the target step size corresponding to each audio segment is output. Based on each target step size, a time warping algorithm is used to perform similarity matching on each target audio segment within each audio segment at the target step size interval, and the first target audio segment and the second target audio segment with a similarity greater than a preset similarity threshold are output. A two-word joint detection algorithm is used to identify the first target audio segment and the second target audio segment, and to identify the target wake word. The two-word joint detection algorithm is a target detection algorithm that simultaneously detects two related words. This invention matches the feature templates of pre-recorded audio with key feature information in real-time audio, enabling the rapid elimination of known sounds and reducing interference. Then, by adjusting the detection step size of audio segments and utilizing a time warping algorithm, it ensures accurate identification of target audio segments across different scenarios and audio variations. Finally, a dual-word joint detection algorithm further enhances the detection accuracy of dual wake words, making it particularly suitable for environments requiring high sensitivity and low false alarm rates, such as childcare and security monitoring. Overall, this method improves the accuracy and reliability of dual wake words, reduces false alarms, and increases response efficiency. Attached Figure Description
[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.
[0059] Figure 1 This is a schematic diagram of the overall working process of the wake word recognition method based on dual-word joint detection in Embodiment 1 of the present invention;
[0060] Figure 2 This is a schematic diagram of the process for silent recognition of the real-time audio data in Embodiment 1 of the present invention;
[0061] Figure 3 This is a schematic diagram of the process for recognizing speech audio segments in Embodiment 1 of the present invention;
[0062] Figure 4 This is a flowchart illustrating the process of processing key feature information and identifying the target wake word in Embodiment 1 of the present invention.
[0063] Figure 5 This is a schematic diagram of the process for identifying the first target audio segment and the second target audio segment in Embodiment 1 of the present invention;
[0064] Figure 6 This is a schematic diagram of the process of filtering the first target audio segment and the second target audio segment by time interval in Embodiment 1 of the present invention;
[0065] Figure 7 This is a schematic diagram of the process of recognizing wake words for the first qualified audio segment and the second qualified audio segment in Embodiment 1 of the present invention;
[0066] Figure 8 This is a structural block diagram of the wake word recognition device based on dual-word joint detection in Embodiment 3 of the present invention;
[0067] Figure 9 This is a schematic diagram of the electronic device in Embodiment 4 of the present invention. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the terms "center," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, the element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Where there is no conflict, embodiments of the present invention and the various features thereof can be combined with each other, all of which are within the scope of protection of the present invention.
[0069] Example 1
[0070] Please see Figure 1 Embodiment 1 of the present invention provides a wake word recognition method based on dual-word joint detection, the method comprising:
[0071] S1: Acquire real-time audio data in various caregiving scenarios;
[0072] Specifically, real-time audio data is acquired in various care scenarios, including at least medical institutions, nursing homes, and home care. Sensors and devices used in different care scenarios, such as microphones and surveillance cameras, are used to capture audio information about the monitored individuals as the real-time audio data.
[0073] S2: Using a multi-feature fusion algorithm, perform silent recognition on the real-time audio data, remove the recognized silent audio segments, and output the speech audio segments;
[0074] Specifically, the multi-feature fusion algorithm refers to extracting various feature information from real-time audio data. These feature information includes at least: time-domain features (such as energy and zero-crossing rate), frequency-domain features (such as Mel-frequency cepstral coefficients (MFCC), and time-frequency-domain features. These features reflect different aspects of the audio signal, facilitating more comprehensive silence recognition. The obtained features are fused using fusion algorithms such as Principal Component Analysis (PCA) or Linear Discriminant Analysis (LDA). A silence recognition model is then trained using labeled training data, employing machine learning algorithms such as Support Vector Machine (SVM) or deep learning models (such as Recurrent Neural Network (RNN) or Long Short-Term Memory Network (LSTM)). Real-time audio data is input into the trained silence recognition model for real-time silence recognition. Based on the model's output, audio segments marked as silent are removed, while segments marked as speech are retained. The audio segments that have undergone silence recognition and filtering are output as the final speech audio segments. By introducing silence recognition technology, all audio is segmented, and key audio segments are effectively extracted. This helps reduce the processing of non-critical silent audio and improves system efficiency and resource utilization.
[0075] In one embodiment, please refer to Figure 2 S2 includes:
[0076] S21: Obtain the original audio signal corresponding to each audio frame in the real-time audio data, extract signal features from the original audio signal, and obtain multiple audio signal features;
[0077] Specifically, the real-time audio data is decomposed into multiple consecutive audio frames. For each audio frame, its corresponding original audio signal is obtained, and feature extraction is performed on each original audio signal to output four features corresponding to the original audio signal: short-time energy, spectral centroid, harmonic ratio, and zero crossover rate.
[0078] S22: Normalize each of the audio signal features and output the standard audio signal features corresponding to each of the audio signal features;
[0079] Specifically, for the four audio signal features mentioned above, their mean μ and standard deviation σ are calculated. For example, the calculation of the mean and standard deviation for feature X is as follows:
[0080]
[0081] Where N is the number of samples of audio signal features, X i This is the i-th sample; based on the mean and standard deviation of feature X, the audio signal feature X is normalized. The normalization formula is x. norm = (xu) / σ, where the feature before normalization is x, and the feature after normalization is x. normThe short-time energy, spectral centroid, harmonic ratio, and zero crossover rate mentioned above are all normalized using the method described above. This process ensures that all audio signal features are on the same scale, avoiding model training difficulties that may result from different feature dimensions. This has a positive impact on improving the convergence speed and performance of machine learning models.
[0082] S23: Based on preset weights, perform weighted fusion processing on the features of each standard audio signal and output a fused feature vector;
[0083] Specifically, a preset weighting factor is defined for each standard audio signal feature. These weighting factors are adjusted according to the importance of the feature or the task requirements. Assuming that the four features, short-time energy, spectral centroid, harmonic ratio, and zero crossover rate, are X1, X2, X3, and X4 respectively, the corresponding weighting factors are w1, w2, w3, and w4. The weights w1, w2, w3, and w4 are applied to each standard audio signal feature Xi for weighted processing. Using a linear combination method, the fused feature F is calculated. The weighted fusion processing allows the contribution of different features to the final feature vector to be balanced by adjusting the weighting factors, so that the model can better adapt to the requirements of the task or better capture the features of the data.
[0084] S24: Input the fused feature vector into a pre-trained classifier and output the probability that the fused feature vector is identified as a speech audio segment;
[0085] Specifically, a training set is prepared, which contains samples labeled as speech or silence. A support vector machine (SVM) classifier is trained using this training set. The fused feature vector F is provided as input to the SVM classifier. The SVM classifier processes the input fused feature vector and outputs the probability that the vector is identified as a speech activity, denoted as P (speech activity).
[0086] S25: Identify the speech audio segment based on the probability of the fused feature vector being identified as a speech audio segment and a preset probability threshold.
[0087] Specifically, first, a probability threshold T is set. This threshold will be used to determine whether the fused feature vector is classified as a speech audio segment, and this threshold can be adjusted according to task requirements and performance requirements. The probability P(speech|F) of the fused feature vector is compared with the preset probability threshold T. If P(speech|F) ≥ T, then the audio segment is classified as a speech segment. If P(speech|F) < T, then the audio segment is classified as a non-speech silent segment. Through this process, the audio segment is binary-classified based on the predicted probability value, and is divided into speech or non-speech. The selection of the probability threshold depends on the specific requirements of the task, and the optimal threshold can be determined through experiments and performance evaluations. In some applications, more attention is paid to accuracy, while in other applications, more attention is paid to recall rate.
[0088] In one embodiment, please refer to Figure 3 , where S25 includes:
[0089] S251: Obtain the probability that each fused feature vector is recognized as a speech audio segment and a preset smoothing factor;
[0090] Specifically, for each fused feature vector Fi, obtain the probability P(speech|Fi) that it is recognized as a speech audio segment, and obtain a preset smoothing factor α, which is used for smoothing the probability.
[0091] S252: According to the smoothing factor, smooth the probability that each fused feature vector is recognized as a speech audio segment, and output the target probability that each fused feature vector is recognized as a speech segment after smoothing;
[0092] Specifically, for the probability P(speech|Fi) of each fused feature vector Fi, use the smoothing factor α for smoothing, and use the moving average method to calculate the smoothed target probability y smooth :
[0093] y smooth = αy current +(1 - α)y previous
[0094] where y previous is the probability after smoothing at the previous moment, and y current is the probability that the current fused feature vector is recognized as a speech audio segment.
[0095] S253: Compare each of the target probabilities with the probability threshold. When the target probability is greater than the probability threshold, identify the audio frame corresponding to the target probability as the speech audio segment.
[0096] Specifically, for each smoothed target probability y smooth,i , compare it with the preset probability threshold T; if ysmooth,i ≥ T, the corresponding audio frame is recognized as a speech audio segment; if y smooth,i < T, the corresponding audio frame is recognized as a silent segment of non-speech audio. By introducing a smoothing factor α, the jump of the decision is reduced to achieve the smoothing of the recognition probability, which helps to reduce sudden decision changes, make the boundary of speech activity smoother, and ensure that the smoothing factor and threshold are adjusted to meet the specific requirements of the task.
[0097] S3: According to the preset duration threshold corresponding to the wake-up word, the audio segments in the speech audio segment with a duration higher than the duration threshold are removed, and the待检测音频片段 is output;
[0098] Specifically, according to the preset duration threshold corresponding to the wake-up word, the audio segments in the speech audio segment with a duration higher than the duration threshold are removed, and the待检测音频片段 is output. The preset time threshold is to delete redundant audio information. Considering the actual scenario, the speech segment when the user calls the target wake-up word "Haima Haima" is usually very short. For example, the time threshold obtained here is 2 seconds.
[0099] S4: Perform spectral feature extraction on the待检测音频片段 to obtain key feature information related to the preset target wake-up word;
[0100] Specifically, the待检测音频片段 is preprocessed, such as denoising, normalization, etc., to improve the accuracy of feature extraction; the MFCC algorithm is used to process the audio segment to extract its frequency characteristics. MFCC is a commonly used audio feature extraction method. It decomposes the audio signal into frequency bands and then extracts the Mel-frequency cepstral coefficients of each frequency band. These coefficients reflect the characteristics of the speech signal in the frequency domain and help to capture the speech information of the speaker; from the features extracted by MFCC, the key feature information related to the preset target wake-up word "Haima Haima" is selected, at least including the energy and spectral shape within a specific time window.
[0101] S5: Use the dynamic time warping algorithm and the bigram joint detection algorithm to process the key feature information to identify the target wake-up word.
[0102] Specifically, DTW is an algorithm for comparing the similarity between two sequences. It allows bending the two sequences along the timeline to find the best match. The two-word joint detection algorithm is a target detection algorithm that simultaneously detects two related words. This algorithm is used to simultaneously detect the wake word "hippocampus." For wake word detection tasks, the two-word joint detection algorithm can improve the accuracy of wake word recognition, especially when there is a contextual relationship. After extracting key feature information, the dynamic time warping algorithm is used in conjunction with the two-word joint detection algorithm to jointly detect the wake word and its context. This processing flow can improve the robustness of wake word recognition and better cope with changes in pronunciation, speech rate, etc.
[0103] In one embodiment, please refer to Figure 4 S5 includes:
[0104] S51: Based on the key feature information, the audio frame corresponding to the key feature information is divided into multiple audio segments;
[0105] Specifically, based on the key feature information, the audio frame corresponding to the key feature information is divided into several windows according to time, and then divided into multiple audio segments, so that each segment can better adapt to the characteristics of different parts and form multiple continuous audio segments.
[0106] S52: Based on the amplitude difference between each audio segment, adjust the step size of the algorithm detection within each audio segment, and output the target step size corresponding to each audio segment;
[0107] Specifically, amplitude calculation is performed on each audio segment, such as calculating the average amplitude within each audio segment and the amplitude difference between adjacent audio segments, resulting in a difference sequence representing the amplitude variation trend. An amplitude threshold is set to divide the difference sequence into several intervals. A larger amplitude difference indicates significant changes between adjacent audio segments, while a smaller difference may indicate relatively stable adjacent segments. For each audio segment, the algorithm detection step size within that segment is dynamically adjusted based on its interval in the difference sequence. A larger amplitude difference means that the segment contains more variations, thus requiring a smaller step size for more accurate matching; a smaller amplitude difference indicates that adjacent segments are relatively stable, requiring a larger step size to improve computational efficiency. The target step size corresponding to each audio segment is output, so that each segment has a step size adapted to its own characteristics. Here, the step size refers to the distance the window moves on the time axis when applying the feature extraction algorithm to the audio signal.
[0108] S53: Based on each target step length, using a time warping algorithm, perform similarity matching on each target audio segment within each audio segment at the target step length interval, and output the first target audio segment and the second target audio segment with a similarity greater than a preset similarity threshold.
[0109] Specifically, based on the target step size, within each audio segment, a time warping algorithm (e.g., DTW) is used for similarity matching to ensure that the matching time interval adapts to the characteristics of each segment, and outputs the first target audio segment and the second target audio segment with a similarity greater than a preset similarity threshold.
[0110] S54: Using the dual-word joint detection algorithm, identify the first target audio segment and the second target audio segment, and identify the target wake word.
[0111] Specifically, the two-word joint detection algorithm involves simultaneously detecting two related words. Here, the two words refer to the target wake word "seahorse seahorse". The two-word joint detection algorithm is applied to the first target audio segment, with the target wake word "seahorse seahorse" as the main focus word. Similarly, the two-word joint detection algorithm is applied to the second target audio segment, with the target wake word "seahorse seahorse" also as the main focus word. The algorithm outputs the recognition result of whether the target wake word "seahorse seahorse" exists in each target audio segment.
[0112] In one embodiment, please refer to Figure 5 S54 includes:
[0113] S541: Perform time interval filtering on the first target audio segment and the second target audio segment, and output the first qualified audio segment and the second qualified audio segment that have passed the filtering.
[0114] In one embodiment, please refer to Figure 6 S541 includes:
[0115] S5411: Obtain the time interval and preset time window between the first target audio segment and the second target audio segment;
[0116] Specifically, from the first target audio segment and the second target audio segment obtained by the above DTW algorithm, the time interval between the first target audio segment and the second target audio segment is obtained, and a preset time window is obtained, that is, the minimum and maximum time interval between two "hippocampus" words are specified. The setting of this time window can be based on the actual application scenario and requirements.
[0117] S5412: Compare the time interval with the time window. If the time interval is determined to be within the time window, output the first qualified audio segment and the second qualified audio segment.
[0118] Specifically, the acquired time interval is compared with a preset time window. It is determined whether the time interval is within the time window; if the time interval is within the preset time window, it means that the time between the two "hippocampus" words is reasonable, and the corresponding first target audio segment and second target audio segment are output as qualified audio segments.
[0119] S5413: If the time interval is not within the time window, then the corresponding first target audio segment and second target audio segment are removed.
[0120] Specifically, if the time interval is not within the preset time window, it means that the time between the two "hippocampus" words does not meet the set conditions. In this case, the corresponding first target audio segment and second target audio segment need to be removed and not included in the output results.
[0121] S542: Perform wake word recognition on the first qualified audio segment and the second qualified audio segment respectively, and output the recognition results;
[0122] In one embodiment, please refer to Figure 7 S542 includes:
[0123] S5421: Obtain first training audio data and second training audio data, wherein the first training audio data includes the target wake-up word, and the second training audio data does not include the target wake-up word;
[0124] Specifically, a speech dataset containing a large number of "hippocampus" words is collected as the first training audio data to ensure coverage of various speech characteristics, environments, and speakers; at the same time, negative example data, i.e. speech data that does not contain "hippocampus" words, is collected as the second training audio data to train the model to distinguish between positive and negative examples.
[0125] S5422: Label the first training audio data and the second training audio data, and output the label information;
[0126] Specifically, the first and second training audio data are labeled, and each audio sample is associated with its corresponding label to indicate whether it contains the target wake word. This can be a binary label, such as "contains" or "does not contain".
[0127] S5423: Extract features from the first training audio data and the second training audio data, and output training feature information;
[0128] Specifically, the MFCC algorithm is used to extract features from the first and second training audio data, converting them into a matrix containing audio features so that the computer learning model can process them.
[0129] S5424: Input the label information and the training feature information into a preset classification model for training to obtain a wake word recognition model;
[0130] Specifically, the label information and MFCC features are integrated to create a training dataset, ensuring the balance of the dataset, i.e., the number of positive and negative examples are similar. The ResNet18 model is selected, which is a classic convolutional neural network (CNN) suitable for image and audio processing tasks. The ResNet18 model is trained using the label information and the training feature information. By iteratively optimizing the model parameters, it can accurately distinguish between audio samples containing and not containing the target wake word, thus deriving the wake word recognition model.
[0131] S5425: Input the first qualified audio segment and the second qualified audio segment into the wake word recognition model respectively, and output the recognition result.
[0132] Specifically, the first and second qualified audio segments, filtered by time intervals, are input into the trained wake word recognition model. The model outputs a recognition result corresponding to each input audio segment. The recognition result is a probability value indicating whether the target wake word is contained. A threshold is set to determine whether the target wake word is considered to exist.
[0133] S543: Based on the recognition results, if the target wake-up word is recognized in both the first qualified audio segment and the second qualified audio segment, then the target wake-up word recognition is considered complete.
[0134] Specifically, based on the recognition results, if both the first and second qualified audio segments are analyzed by the wake word recognition model and identified as containing the target wake word "hippocampus", then the target wake word recognition task is considered to have been successfully completed. This means that the wake word expected by the user has been successfully detected in both audio segments that have been filtered over a specific time interval. This result can be used as a trigger condition to initiate subsequent voice commands or other corresponding processing, thereby enabling a wider range of voice interaction applications.
[0135] Example 2
[0136] Example 1 provides a wake-word recognition method based on dual-word joint detection. However, in real-life scenarios, challenges related to false wake-ups, resource management, and privacy security need to be considered, especially for battery-powered mobile and embedded devices. This undoubtedly increases the pressure on energy management, thereby affecting device performance and user experience. Therefore, based on Example 1, further avoiding false wake-ups and unauthorized operations is not only necessary to improve user experience but also an important measure to maintain home security and protect personal privacy. Therefore, after S4, the method further includes:
[0137] The key feature information extracted from the pre-recorded audio segments and the audio segments to be detected of the target object is obtained, wherein the target object includes at least children and unauthorized users;
[0138] Specifically, feature extraction of pre-recorded audio segments of the target object is to obtain key feature information in the audio of children and unauthorized users, so as to effectively identify wake words and prevent false wake-ups. First, pre-recorded audio segments of the target object are collected and stored. Then, digital signal processing preprocessing, such as noise reduction and normalization, is performed on these audio segments to reduce noise interference and maintain consistent audio quality.
[0139] Feature extraction is performed on the pre-recorded audio segment to obtain feature template information corresponding to the target object;
[0140] Specifically, the MFCC (Mel-Frequency Cepstral Coefficients) feature extraction method is used to convert the audio signal into a spectral feature representation. The MFCC method involves performing short-time Fourier transform, filter bank analysis, and Mel-frequency cepstral transform on the audio signal, ultimately obtaining a set of coefficients reflecting the audio features. The MFCC coefficients contain the frequency and energy distribution information of the audio signal, have good discriminative and expressive power, and can accurately capture the audio features of the target object. The MFCC coefficients are used as the feature template information corresponding to the target object. The feature template information is used to build the audio feature model of the target object so that the audio segment to be detected can be matched and verified in the future. In this step, machine learning algorithms are involved, such as support vector machine (SVM) and Gaussian mixture model (GMM). These algorithms are used to train and model the MFCC features to learn the audio feature pattern of the target object and output the trained audio feature model.
[0141] The key feature information and the feature template information are matched. If the match is successful, the key feature information is removed.
[0142] Specifically, feature matching and elimination are used to verify whether the audio segment to be detected contains the voice features of the target object and to prevent false wake-ups or unauthorized operations. In this step, the extracted MFCC features of the audio segment to be detected are input into the trained audio feature model and matched with the feature template of the target object. The matching process usually uses similarity calculation methods, such as Euclidean distance and cosine similarity. If the MFCC features of the audio segment to be detected successfully match the feature template of the target object, it means that the audio segment contains the audio features of the target object, and the audio segment is immediately eliminated to avoid false wake-ups or unauthorized operations. The elimination method can be to directly discard the segment or mark it as untrusted for further follow-up measures, such as reminding the user or recording abnormal events for subsequent analysis. Through feature matching and elimination, the false wake-up behavior of children and unauthorized users can be effectively identified and prevented, improving the security of the device and the user experience.
[0143] Example 3
[0144] Please see Figure 8 Embodiment 3 of the present invention also provides a wake word recognition device based on dual-word joint detection, the device comprising:
[0145] The real-time audio acquisition module is used to acquire real-time audio data in various care scenarios;
[0146] The speech segment extraction module is used to perform silent recognition on the real-time audio data using a multi-feature fusion algorithm, remove the recognized silent audio segments, and output the speech audio segments.
[0147] The audio segment extraction module is used to remove audio segments with a duration longer than the preset wake word duration threshold and output the audio segment to be detected.
[0148] The key feature extraction module is used to extract key features from the audio segment to be detected and obtain key feature information related to the preset target wake word;
[0149] The wake word recognition module is used to process the key feature information using a dynamic time warping algorithm and a two-word joint detection algorithm to identify the target wake word.
[0150] Specifically, the wake-up word recognition device based on dual-word joint detection provided in this embodiment of the invention includes: a real-time audio acquisition module for acquiring real-time audio data in various care scenarios; a speech segment extraction module for performing silent recognition on the real-time audio data using a multi-feature fusion algorithm, removing the identified silent audio segments, and outputting speech audio segments; a target audio segment extraction module for removing corresponding audio segments with a duration higher than a preset wake-up word duration threshold, and outputting target audio segments; a key feature extraction module for extracting key features from the target audio segments to obtain key feature information related to a preset target wake-up word; and a wake-up word recognition module for processing the key feature information using a dynamic time warping algorithm and a dual-word joint detection algorithm to recognize the target wake-up word. This device acquires real-time audio data from various caregiving scenarios, fully considering the sound characteristics of different environments, thus better adapting to complex and ever-changing real-world usage scenarios. Secondly, in the silence recognition stage, it improves the accuracy of distinguishing between audible and silent parts, more effectively removing background noise and non-speech audio, reducing invalid audio segments for subsequent processing, improving work efficiency, and minimizing response latency. Furthermore, by setting a preset wake-up word duration threshold, it filters out excessively long speech segments that clearly do not conform to the wake-up word, avoiding wasting work resources on long audio segments. Finally, in the wake-up word recognition stage, it employs a dynamic time warping algorithm and a two-word joint detection algorithm. The combination of these two algorithms allows for more flexible handling of changes in speech signals, and through joint detection, it not only helps adapt to different speech speeds and change patterns but also reduces response latency and improves real-time performance. Overall, by comprehensively considering multiple factors, from data diversity to processing accuracy and real-time performance, this device provides strong technical support for voice wake-up of resource-constrained embedded systems or mobile devices in various scenarios.
[0151] Example 4
[0152] In addition, combined Figure 1 The wake word recognition method based on dual-word joint detection described in Embodiment 1 of the present invention can be implemented by an electronic device. Figure 9 A schematic diagram of the hardware structure of the electronic device provided in Embodiment 4 of the present invention is shown.
[0153] Electronic devices may include processors and memory storing computer program instructions.
[0154] Specifically, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.
[0155] The memory may include a large-capacity storage device for data or instructions. For example, and not limitingly, the memory may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to a data processing device. In a particular embodiment, the memory is a non-volatile solid-state memory. In a particular embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0156] The processor reads and executes computer program instructions stored in memory to implement any of the wake word recognition methods based on dual-word joint detection in the above embodiments.
[0157] In one example, the electronic device may also include a communication interface and a bus. For example, Figure 9 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.
[0158] The communication interface is mainly used to enable communication between various modules, devices, units and / or equipment in the embodiments of the present invention.
[0159] A bus, including hardware, software, or both, couples components of the device together. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, a bus may include one or more buses. While specific buses are described and illustrated in embodiments of the invention, the invention contemplates any suitable bus or interconnect.
[0160] Example 5
[0161] Furthermore, in conjunction with the wake word recognition method based on dual-word joint detection in Embodiment 1 above, Embodiment 5 of the present invention can also provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the wake word recognition methods based on dual-word joint detection in the above embodiments.
[0162] In summary, the embodiments of the present invention provide a wake word recognition method, apparatus, device, and medium based on dual-word joint detection.
[0163] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0164] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0165] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0166] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A wake-word recognition method based on dual-word joint detection, characterized in that, The method includes: The system acquires pre-recorded audio segments of the target object and key feature information extracted from audio segments to be detected in various care scenarios, wherein the target object includes at least children and unauthorized users; Feature extraction is performed on the pre-recorded audio segment to obtain feature template information corresponding to the target object; The key feature information and the feature template information are matched. If the match is successful, the key feature information is removed. If the matching fails, the audio frame corresponding to the key feature information is divided into multiple audio segments based on the key feature information. Based on the amplitude difference between each audio segment, the step size of the algorithm detection within each audio segment is adjusted, and the target step size corresponding to each audio segment is output. Based on each target step size, the time warping algorithm is used to perform similarity matching on each target audio segment within each audio segment at the target step size interval, and output the first target audio segment and the second target audio segment with a similarity greater than the preset similarity threshold. The two-word joint detection algorithm is used to identify the first target audio segment and the second target audio segment, and to identify the target wake word. The two-word joint detection algorithm is a target detection algorithm that detects two related words at the same time.
2. The wake word recognition method based on dual-word joint detection according to claim 1, characterized in that, Before acquiring the pre-recorded audio segments of the target object and the key feature information extracted from the audio segments to be detected in various care scenarios, the following is included: Acquire real-time audio data in various caregiving scenarios; Using a multi-feature fusion algorithm, the real-time audio data is subjected to silent detection, the detected silent audio segments are removed, and the speech audio segments are output. Based on the preset duration threshold corresponding to the wake word, the corresponding audio segments with a duration longer than the duration threshold are removed from the speech audio segments, and the audio segments to be detected are output. Spectral features are extracted from the audio segment to be detected to obtain key feature information related to the preset target wake word.
3. The wake word recognition method based on dual-word joint detection according to claim 2, characterized in that, The method of using a multi-feature fusion algorithm to perform silent audio recognition on the real-time audio data, removing the identified silent audio segments, and outputting speech audio segments includes: Obtain the original audio signal corresponding to each audio frame in the real-time audio data, extract signal features from the original audio signal, and obtain multiple audio signal features; The audio signal features are normalized, and the standard audio signal features corresponding to each audio signal feature are output. Based on preset weights, the features of each standard audio signal are weighted and fused to output a fused feature vector; The fused feature vector is input into a pre-trained classifier, which outputs the probability that the fused feature vector is identified as a speech audio segment. The speech audio segment is identified based on the probability of the fused feature vector being identified as a speech audio segment and a preset probability threshold.
4. The wake word recognition method based on dual-word joint detection according to claim 3, characterized in that, The process of identifying the speech audio segment based on the probability of recognizing it as a speech audio segment according to the fused feature vector and a preset probability threshold includes: Obtain the probability of each fused feature vector being identified as a speech audio segment and the preset smoothing factor; Based on the smoothing factor, the probability of each fused feature vector being identified as a speech audio segment is smoothed, and the target probability of each fused feature vector being identified as a speech segment after smoothing is output. The target probabilities are compared with the probability thresholds. When the target probability is greater than the probability threshold, the audio frame corresponding to the target probability is identified as the speech audio segment.
5. The wake word recognition method based on dual-word joint detection according to claim 1, characterized in that, The method of using the dual-word joint detection algorithm to identify the first target audio segment and the second target audio segment, and identifying the target wake word includes: The first target audio segment and the second target audio segment are filtered by time interval, and the first qualified audio segment and the second qualified audio segment that pass the filter are output. Perform wake word recognition on the first and second qualified audio segments respectively, and output the recognition results; Based on the recognition results, if the target wake word is recognized in both the first qualified audio segment and the second qualified audio segment, then the target wake word recognition is considered complete.
6. The wake word recognition method based on dual-word joint detection according to claim 5, characterized in that, The step of filtering the first target audio segment and the second target audio segment by time interval, and outputting the first qualified audio segment and the second qualified audio segment that have passed the filtering, includes: Obtain the time interval and preset time window between the first target audio segment and the second target audio segment; The time interval is compared with the time window. If the time interval is determined to be within the time window, the first qualified audio segment and the second qualified audio segment are output. If the time interval is not within the time window, the corresponding first target audio segment and second target audio segment are removed.
7. The wake word recognition method based on dual-word joint detection according to claim 5, characterized in that, The wake word recognition is performed on the first and second qualified audio segments respectively, and the recognition results are output as follows: Acquire first training audio data and second training audio data, wherein the first training audio data includes the target wake word, and the second training audio data does not include the target wake word; The first and second training audio data are labeled, and the label information is output. Feature extraction is performed on the first training audio data and the second training audio data to output training feature information; The label information and the training feature information are input into a preset classification model for training to obtain a wake word recognition model; The first qualified audio segment and the second qualified audio segment are respectively input into the wake word recognition model, and the recognition results are output.
8. A wake-up word recognition device based on dual-word joint detection, characterized in that, The device includes: The key feature information acquisition module is used to acquire key feature information extracted from pre-recorded audio segments of the target object and audio segments to be detected in various care scenarios, wherein the target object includes at least children and unauthorized users; The feature extraction module is used to extract features from the pre-recorded audio segment and obtain feature template information corresponding to the target object; The feature matching module is used to perform feature matching between the key feature information and the feature template information. If the match is successful, the key feature information is removed. The audio frame segmentation module is used to divide the audio frame corresponding to the key feature information into multiple audio segments based on the key feature information if the matching fails. The step size adjustment module is used to adjust the step size detected by the algorithm within each audio segment based on the amplitude difference between each audio segment, and output the target step size corresponding to each audio segment. The similarity matching module is used to perform similarity matching on each target audio segment within each audio segment according to each target step size and using a time warping algorithm, and output the first target audio segment and the second target audio segment with a similarity greater than a preset similarity threshold. The wake-up word recognition module is used to identify the first target audio segment and the second target audio segment using a two-word joint detection algorithm, and to identify the target wake-up word. The two-word joint detection algorithm is a target detection algorithm that detects two related words at the same time.
9. An electronic device, characterized in that, include: At least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method as described in any one of claims 1-7.
10. A storage medium storing computer program instructions thereon, characterized in that, The method as described in any one of claims 1-7 is implemented when the computer program instructions are executed by the processor.
Citation Information
Patent Citations
Voice wakeup method and device based on voiceprint recognition
CN107886957A
Speech recognition method and wake-up word detection method and device
CN109192210A
Electronic device for processing user utterance, and control method therefor
CN112639962A