An Offline Speech Keyword Recognition Method and Its System Implementation for Mixed Chinese-English in a Specific Scenario

By adopting a mixed offline recognition method and an online waste model in the voice keyword recognition technology, the problems of low recognition accuracy, slow response time and privacy leakage in the prior art are solved, and the recognition accuracy rate does not decrease and response speed is improved when replacing the keyword list.

CN114530141BActive Publication Date: 2025-06-24BEIHANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202011323748.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-23
Publication Date
2025-06-24
Estimated Expiration
2040-11-23

AI Technical Summary

Technical Problem

The existing voice keyword recognition technology has problems such as low recognition accuracy, slow response time, privacy leakage when data is uploaded to the cloud, and the recognition accuracy drops sharply when the number of keywords to be identified increases.

Method used

The offline speech keyword recognition method mixed with Chinese and English is adopted, and the basic modeling unit of the hidden Markov model acoustic model is used to match the non-keyword parts in continuous speech by using the online waste model. Strategies such as adaptive keyword matching windows, simplified speech activity detection and optimized path decoding are proposed.

Benefits of technology

It realizes that when changing keyword lists in specific scenarios, the recognition accuracy rate will not be reduced, the response speed will be improved, and the risk of privacy leakage in data upload to the cloud is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114530141B_ABST
    Figure CN114530141B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method and system for offline speech keyword recognition with a mixture of Chinese and English. A specific implementation of the method includes: obtaining a speech digital signal, performing voice activity detection on it to obtain a speech segment to be recognized; defining an adaptive keyword matching window and segmenting the speech segment to be recognized; extracting features from the speech segment to obtain Mel-frequency cepstral coefficient embedding feature vectors; analyzing a custom keyword list and combining it with a pre-trained phoneme padding model to obtain a Chinese decoding network space and an English decoding network space; sequentially inputting the Mel-frequency cepstral coefficient embedding feature vectors into the decoding network spaces to obtain a recognition result; and post-processing the recognition result to generate a target recognition result. This implementation has a low computational load, can perform offline recognition, has a high recognition accuracy, a fast response speed, supports mixed Chinese and English recognition, and can flexibly replace the keyword list to adapt to applications in different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of speech recognition, and particularly to an offline speech keyword recognition method and system for mixed Chinese and English. Background Art

[0002] Speech keyword recognition technology is a research branch in the field of automatic speech recognition. Automatic speech recognition technology requires complete decoding and conversion of the collected continuous speech stream, which has higher requirements for hardware computing power. It often adopts an online recognition method of uploading data to the cloud for computing. When the network is offline, the recognition effect drops sharply, and there is also a risk of privacy leakage when uploading data to the cloud. Speech keyword recognition only focuses on the keyword part, has lower hardware dependence, and can complete decoding and calculation locally, thus realizing offline recognition, and has broad application prospects in specific scenarios such as the military field, air traffic control field, and speech monitoring field.

[0003] Speech keyword recognition began in the 1970s. After years of technological development and accumulation, speech keyword recognition algorithms can be roughly divided into two categories. One is speech keyword recognition based on a phoneme padding model. This method regards the speech to be recognized as consisting of two parts: keywords and non-keywords. The other is speech keyword recognition based on template matching. This method directly compares the speech to be recognized with the keyword template speech and calculates the distance difference.

[0004] For the speech keyword recognition method based on the phoneme padding model, there are two implementation schemes based on the hidden Markov model and the neural network. The phoneme padding model based on the hidden Markov model establishes HMMs (Hidden Markov Models) for both keywords and non-keywords. HMM can better describe the statistical distribution probability of the speech signal feature states, but this implementation method has disadvantages such as low recognition accuracy and slow response time and needs to be further optimized. The phoneme padding model based on the neural network regards keyword recognition as a classification problem of keywords and non-keywords. This scheme requires a large amount of corpus for neural network training. When changing keywords, it is necessary to re-collect training data and re-train network parameters. Therefore, this scheme is relatively limited in practical applications.

[0005] For the speech keyword recognition method based on template matching, there are also two implementation schemes based on DTW (Dynamic Time Warping) and embedding learning. The keyword recognition based on DTW is a method used in the early stage of speech keyword recognition. The core idea is to perform sequence alignment through dynamic programming and then calculate the distance between sequences. This method is relatively simple to implement, but it is mainly used for isolated word recognition. The speech keyword recognition based on embedding learning is to train a neural network feature extractor (e.g., LSTM feature extractor), convert the speech to be recognized and the keyword template speech into feature vectors of the same length through the feature extractor, and then calculate the vector distance. This implementation method has a high recognition accuracy when recognizing a single keyword, so it has a wide application in the field of intelligent device wake-up. However, as the number of keywords to be recognized increases, the recognition accuracy will drop sharply. Although only a small amount of keyword template corpus needs to be collected to replace the keywords to be recognized, this method brings the problem of poor recognition effect for non-specific people. Summary of the Invention

[0006] This section of the disclosure is used to introduce concepts in a brief form, which will be described in detail in the following detailed implementation section. This section of the disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] In view of the problems existing in the above speech keyword recognition, some embodiments of the present disclosure propose an offline speech keyword recognition method and its system implementation that combines Chinese and English, and propose using context-related phonemes as the basic modeling unit of the hidden Markov model acoustic model, and using an online waste model to match the non-keyword part in continuous speech. Through these methods, keyword recognition can easily replace the keyword list according to a specific scenario without having too much impact on the recognition accuracy, and it is not necessary to retrain the waste model when replacing the keyword list. Some embodiments of the present disclosure also propose strategies such as simplifying speech activity detection and optimizing path decoding to improve the system recognition accuracy and response speed.

[0008] Some embodiments of the present disclosure provide a method for offline speech keyword recognition in a Chinese-English mixture. The method includes: obtaining a speech digital signal, performing speech activity detection on it to obtain a speech segment to be recognized; defining an adaptive keyword matching window and segmenting the speech segment to be recognized; extracting features from the speech segment within the window to obtain a Mel Frequency Cepstral Coefficient (MFCC) embedding feature vector; analyzing a custom keyword list in a specific scenario and combining a pre-trained phoneme padding model to obtain a Chinese decoding network space and an English decoding network space for the custom keywords; sequentially inputting the MFCC embedding feature vector into the decoding network spaces to obtain a recognition result; and post-processing the recognition result to generate a target recognition result as the output.

[0009] Optionally, the speech activity detection includes: defining parameter information for speech acquisition, and calling an audio processing interface to perform quantization processing on the original speech with the following parameters: the sampling frequency is 16,000 Hz, the number of channels is 1, and the number of speech frames included in each speech block is 1,024, to obtain the speech frame coding information x(n) = (x1(n), x2(n),..., x m (n)) at the nth moment. The combination of k speech frame coding information gives the original speech block information f = (x(t1), x(t2),..., x(t k ) within the time period from t1 to t k ); calculating the average sound intensity for the quantized speech frame coding information as follows: where n represents the nth moment, x represents the speech frame coding information, x(n) represents the speech frame coding information collected at the nth moment, x1(n) represents the first bit in the speech frame coding information collected at the nth moment, x2(n) represents the second bit in the speech frame coding information collected at the nth moment, x m (n) represents the mth bit in the speech frame coding information collected at the nth moment, f represents the speech block, t represents the moment, t1 represents the first moment, t2 represents the second moment, t k represents the kth moment, x(t1) represents the speech frame coding information collected at the t1 moment, x(t2) represents the speech frame coding information collected at the t2 moment, x(t k ) represents the speech frame coding information collected at the t k moment, threshold represents the average sound intensity and also serves as the silence threshold in the current environment, γ represents the influence factor, and its specific value is γ = 2.5, k represents the serial number, i represents the serial number, m represents the number of speech frame coding information, X i represents the ith bit in the speech frame coding information, x i(n) represents the i-th bit in the speech frame coding information collected at the n-th moment; analyze the change in sound intensity, and dynamically update the mute threshold when keyword recognition is completed or when the sound intensity does not exceed the threshold for a long time.

[0010] Optionally, the defining of the adaptive keyword matching window includes: calculating the average keyword length with reference to the keyword list: where l represents the average keyword length, n represents the number of keywords, i represents the serial number, and l i represents the length of the i-th keyword; define the length wnd of the matching window and the moving distance rwnd of the window based on the average keyword length, where wnd satisfies 1.5l ≤ wnd ≤ 2l, and when a keyword is recognized, rwnd = 0.8l, and if no keyword is recognized, rwnd = 0.4l.

[0011] Optionally, the extracting features from the speech segment within the window to obtain the Mel-frequency cepstral coefficient embedding feature vector includes: pre-emphasizing the speech signal within the keyword matching window to compensate for the loss of high-frequency signals during sound propagation; overlapping and framing the speech signal with a fixed frame length and frame shift to obtain the framed speech signal; windowing the framed speech signal to obtain a speech signal with enhanced central part and zeroed-out rest; performing Fourier transform on the windowed speech signal to obtain the linear spectrum of each frame of speech signal; inputting the linear spectrum into the Mel-frequency filter bank to obtain the Mel-frequency cepstral coefficient embedding feature vector.

[0012] Optionally, the analyzing the custom keyword list in a specific scenario and combining with a pre-trained phoneme interpolation model to obtain the Chinese decoding network space and English decoding network space of the custom keyword includes: training a hidden Markov model acoustic model with context-dependent phonemes as the basic modeling unit, and constructing a phoneme interpolation model with an online waste model, where the phoneme is the smallest basic unit constituting speech, and the online waste model directly calculates the local waste probability score of each speech frame in the phoneme model without the need to train a separate waste model; customizing the keyword list according to the application requirements of different scenarios, and generating dictionary information on the correspondence between keywords and phonemes in the partitioning method of the Carnegie Mellon University dictionary; using the keyword text as the language model corpus and performing language modeling based on the statistical language model. For a given keyword sequence S = (s1, s2,..., s n ), the 3-gram language model probability is expressed as follows: where S represents the keyword sequence, s1 represents the first character in the keyword sequence, s2 represents the second character in the keyword sequence, s n represents the n-th character in the keyword sequence, n represents the length of the keyword sequence, i represents the serial number, P represents the probability, and P(s1, s2,..., sn ) represents the probability of a keyword sequence occurring in the order of (s1, s2,..., s n ), P(s i |s i-1 , s i-2 ) represents the probability of s i given s i-1 and S i-2 , where S i-1 represents the (i - 1)-th word in the keyword sequence, and S i-2 represents the (i - 2)-th word in the keyword sequence. represents the product calculation of the probabilities from the 1st to the nth; the pre-trained phoneme padding model, dictionary information, and 3-gram language model probabilities together constitute the Chinese decoding network space and English decoding network space of the custom keyword list. Among them, when the keyword list changes, the phonemes constituting the speech do not need to be retrained, and only the dictionary information and 3-gram language model probabilities of the keyword list to be recognized need to be regenerated.

[0013] Optionally, inputting the Mel-frequency cepstral coefficient embedded feature vectors into the decoding network space in sequence to obtain the recognition result includes: obtaining the Mel-frequency cepstral coefficient embedded feature vectors within the adaptive keyword matching window as the speech observation sequence: O = (o1, o2,..., o M ), where O represents the speech observation sequence, o1 represents the 1st frame in the speech observation sequence, o2 represents the 2nd frame in the speech observation sequence, and O M represents the Mth frame in the speech observation sequence; in the multi-language decoder where the Chinese decoding network space λ1 and the English decoding network space λ2 are in parallel, for the same speech observation sequence O = (o1, o2,..., o M ), the Viterbi algorithm is used to perform parallel decoding in the two decoding network spaces respectively to obtain the best state sequences P c and P e containing keyword phonemes and non-keyword phonemes of the given speech observation sequence, and calculating the confirmation score as follows: Among them, S1 represents the confirmation score of the speech observation sequence O in the Chinese decoding network space, P represents probability, O represents the speech observation sequence, P c represents the Chinese best state sequence, P(P c |O) represents the conditional probability of P c occurring when the speech observation sequence is O, P(P c ) represents the probability of P c occurring in the language model, P(O) represents the probability of the speech observation sequence O, and P(O|P c ) represents the probability of O given P cThe conditional probability that appears when, S2 represents the confirmation score of the speech observation sequence O in the English decoding network space, P e represents the English best state sequence, P(P e |O) represents P e The conditional probability that appears when the speech observation sequence is O, P(P e ) represents P e The probability that appears in the language model, P(O|P e ) represents the conditional probability that O appears when P e ; where P(P c ) and P(P e ) are obtained from the language model, P(O|P c ) and P(O|P e ) are obtained from the hidden Markov model acoustic model. The denominators of the two formulas are the same. Specifically, it is to compare the numerators, that is, to compare the probability of generating the current speech observation sequence in which language decoding network space is the largest. If S1 > S2, it is considered that Chinese is recognized, otherwise it is considered that English is recognized.

[0014] Optionally, the post-processing of the recognition result to generate a target recognition result as output includes: under the guidance of dictionary information and a language model, combining the best state path to obtain a recognition result including keyword and non-keyword information; for the recognition result including non-keyword information, using the keyword list as the evaluation criterion for the phoneme output probability, and obtaining a target recognition result as output, which satisfies: Among them, P represents probability, W represents keyword, W i represents the i-th keyword in the keyword list, C represents the recognition result obtained by combining the best state path, represents the most likely recognized keyword when the recognition result is C.

[0015] Some embodiments of the present disclosure also provide an offline speech keyword recognition system for mixed Chinese and English, including: a voice real-time monitoring module for the microphone to real-time monitor the voice signal in the current environment; a voice activity detection module for detecting the speech segment to be recognized in the voice signal; a keyword recognition module for determining whether there is a keyword in the voice signal; a data record storage area retrieval module for recording the information related to the keyword in the database and providing a data query function.

[0016] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: First, some embodiments of the present disclosure propose a simplified voice activity detection algorithm, which is simple to implement, has a low computational complexity, and can complete the calculation locally, contributing to the offline implementation of keyword recognition. It can dynamically update the silence threshold, quickly sense the background noise changes in the current environment, and has better endpoint discrimination ability in the case of a relatively complex background environment. Second, some embodiments of the present disclosure propose an adaptive sliding window for voice keyword matching. The sliding window length is defined according to the length of the keyword to be recognized, and the long speech segment is further segmented into shorter speech segments. This method combines the advantages of the template matching algorithm and the hidden Markov model acoustic model. The former makes the algorithm have a low computational complexity in the decoding calculation process, can achieve real-time response, and also makes it possible to recognize mixed Chinese and English. The latter makes the matching network have the characteristics of the hidden Markov model acoustic model, recognizes by phonemes during decoding, has a higher recognition accuracy for keywords with the same prefix, and improves the response speed of the system. Then, some embodiments of the present disclosure propose post-processing operations for keyword recognition. For the recognition results containing waste information obtained from the path decoding part, using the keyword list as the criterion for judging the posterior probability score of phonemes, removing the waste state to obtain the final recognition result, which improves the system's rejection rate. Finally, some embodiments of the present disclosure propose a phoneme padding model with phonemes as the basic modeling unit and an online waste model, which can conveniently replace the keyword list without reducing the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0018] Figure 1 is a design diagram of the implementation module of an exemplary system to which some embodiments of the present disclosure can be applied;

[0019] Figure 2 is a flowchart of some embodiments of the method for offline speech keyword recognition of mixed Chinese and English according to the present disclosure;

[0020] Figure 3 is a flowchart of some other embodiments of the method for offline speech keyword recognition of mixed Chinese and English according to the present disclosure;

[0021] Figure 4 is a flowchart of some other embodiments of the method for offline speech keyword recognition of mixed Chinese and English according to the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not used to limit the protection scope of the present disclosure.

[0023] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0024] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0025] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0027] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0028] Figure 1 It is a design diagram of an implementation module of an exemplary system to which some embodiments of the present disclosure can be applied. It includes a voice listening module, a voice recognition module, and a data query module. The voice listening module and the voice recognition module together constitute the offline voice keyword recognition in Chinese and English mixture of the present disclosure. The data query module mainly provides data recording and retrieval functions.

[0029] The voice monitoring module mainly includes a microphone recording module, a voice activity detection module, and a voice feature extraction module. Among them, the microphone recording module is mainly responsible for calling relevant audio processing interfaces (e.g., APIs) to obtain and quantize and store voice signals. The voice activity detection module mainly detects the voice collected by the microphone to determine whether there is a sound coming. Some embodiments of the present disclosure implement voice activity detection using a simplified double-threshold method. The voice preprocessing module mainly performs adaptive matching window segmentation and feature extraction on the speech segment to be recognized after voice activity detection processing. Among them, since the collected voice digital signal contains some background noise, silence, and the voice signal part of the actual speaker. The voice signal part of the actual speaker is what is needed for keyword recognition. The speech segment to be recognized extracts the real speaking voice part in the voice signal, that is, the voice signal part. As an example, Mel Frequency Cepstral Coefficients can be used for feature extraction.

[0030] Specifically, the microphone recording module uses an interface (e.g., the PyAudio interface). Define the parameter information of the voice signal. Define the parameter information of voice collection. Call the audio processing interface to perform quantization processing on the original voice with the following parameters: the sampling frequency is 16,000 Hz. The number of channels is 1. The number of voice frames included in each voice block is 1,024. Obtain the voice frame coding information x(n) = (x1(n), x2(n),..., x m (n)) of the quantization processing at the nth moment. The k voice frame coding information is combined to obtain the original voice block information f = (x(t1), x(t2),..., x(t k )) within the time period from t1 to t k . Among them, n represents the nth moment. x represents the voice frame coding information. x(n) represents the voice frame coding information collected at the nth moment. x1(n) represents the 1st bit in the voice frame coding information collected at the nth moment. x2(n) represents the 2nd bit in the voice frame coding information collected at the nth moment. x m (n) represents the mth bit in the voice frame coding information collected at the nth moment. f represents the original voice block information. t represents the moment. t1 represents the 1st moment. t2 represents the 2nd moment. t k represents the kth moment. x(t1) represents the voice frame coding information collected at the t1 moment. x(t2) represents the voice frame coding information collected at the t2 moment. x(t k ) represents the voice frame coding information collected at the t k moment.

[0031] Continuing to refer to Figure 2 , a flowchart showing some embodiments of the offline voice keyword recognition method in Chinese and English mixture according to the present disclosure is shown.

[0032] Analyze the collected speech signal, and collect the speech frame coding information x(n) in the current environment for 1 s. Calculate the average sound intensity of the sound signal during this time period:

[0033]

[0034] Among them, n represents the nth moment. x represents the speech frame coding information. x(n) represents the speech frame coding information collected at the nth moment. x1(n) represents the 1st bit in the speech frame coding information collected at the nth moment. x2(n) represents the 2nd bit in the speech frame coding information collected at the nth moment. x m (n) represents the mth bit in the speech frame coding information collected at the nth moment. f represents the original speech block information. t represents the moment. t1 represents the 1st moment. t2 represents the 2nd moment. t k represents the kth moment. x(t1) represents the speech frame coding information collected at the t1 moment. x(t2) represents the speech frame coding information collected at the t2 moment. x(t k ) represents the speech frame coding information collected at the t k moment. threshold represents the average sound intensity and also serves as the silence threshold in the current environment. γ represents the influence factor. The specific value is γ = 2.5. k represents the serial number. i represents the serial number. m represents the number of speech frame coding information. X i represents the ith bit in the speech frame coding information. x i (n) represents the ith bit in the speech frame coding information collected at the nth moment.

[0035] Take the threshold as the silence threshold in the current environment. Judge the start and end of the sound according to this silence threshold. If there is still no sound intensity exceeding the silence threshold within 10 s, recalculate and update the silence threshold.

[0036] The speech feature extraction module mainly converts the speech signal into an acoustic feature vector. First, pre-emphasize the speech signal to make up for the high-frequency energy loss during sound propagation. Then, overlap and frame the speech signal with a frame length of 25 ms and a frame shift of 10 ms. Next, perform a windowing operation on each frame of the signal y(n) = x(n) × w(n). Where w(n) represents the window function. Finally, perform a Fourier transform on each windowed frame of the signal, and then perform Mel-frequency cepstral coefficient extraction to obtain the Mel-frequency cepstral coefficient embedding feature vector.

[0037] Further refer to Figure 3 , which shows a flowchart of some other embodiments of the Chinese-English mixed offline speech keyword recognition method according to the present disclosure.

[0038] The speech recognition module mainly includes two parts: acoustic model training and keyword recognition. The acoustic model training part includes a custom keyword module and a decoding network training module corresponding to the custom keyword. The keyword recognition module mainly includes three parts: adaptive keyword matching window definition, path decoding, and post-processing of recognition results. Among them, the adaptive keyword matching window is a method in the present disclosure to solve the problems of low keyword recognition rate and slow recognition response speed. This method uses a sliding window to further divide the real speech signal part obtained by voice activity detection into small speech segments one by one. The size of the window is determined by the length of the keyword to be recognized. The moving distance of the window is determined by whether a keyword is recognized. Since the size and moving distance of the window are not fixed but vary with the keyword, it is "adaptive". The predetermined number of speech segments segmented by the sliding window is the "keyword matching window" mentioned here. As an example, the predetermined number can be 5.

[0039] The following refers to Figure 4 , which shows a flowchart of some other embodiments of the Chinese-English mixed offline speech keyword recognition method according to the present disclosure.

[0040] Obtain a keyword list to be recognized according to the application requirements in a specific scenario, define the dictionary information corresponding to keywords and phonemes according to the division method of the Carnegie Mellon University dictionary, and train to generate a language model. Phonemes are the smallest basic units that make up speech. In this implementation, the division method of the Carnegie Mellon University dictionary is adopted. Use the context-related phonemes as the basic modeling units to train the hidden Markov model acoustic model. An online waste model is used to match the non-keyword part in the speech. The construction of the hidden Markov model acoustic model λ mainly determines the state transition probability matrix A and the state alignment probability matrix B. According to the above Mel-frequency cepstral coefficient observation sequence O=(o1, o2,..., o M ) continuously iterate the model parameters. Among them, O represents the speech observation sequence. o1 represents the first frame in the speech observation sequence. o2 represents the second frame in the speech observation sequence. o M represents the Mth frame in the speech observation sequence. Make the probability P(O|λ) the largest. In this implementation, the parameter training part is completed using an algorithm (for example: Baum-Welch algorithm). Use the keyword text as the language model corpus and perform language modeling based on the statistical language model. For a given keyword sequence S=(s1, s2,..., s n ), the 3-gram language model probability is expressed as follows:

[0041]

[0042] Among them, S represents the keyword sequence. s1 represents the first character in the keyword sequence. s2 represents the second character in the keyword sequence. sn Represents the nth word in the keyword sequence. n represents the length of the keyword sequence. i represents the serial number. P represents the probability. P(s1, s2,..., s n ) represents the probability of the keyword sequence that appears in the order of (s1, s2,..., s n ). P(s i |s i-1 , s i-2 ) represents the probability of si given s i-1 and s i-2 . s i-1 represents the (i - 1)th word in the keyword sequence. S i-2 represents the (i - 2)th word in the keyword sequence. Represents the calculation of the product of the probabilities from the 1st to the nth.

[0043] The pre-trained phoneme padding model, the dictionary information of the custom keywords, and the language model. The three together constitute the Chinese decoding network space and the English decoding network space of the custom keyword list. When the keyword list is changed, the phonemes forming the speech do not need to be retrained. Only the dictionary information and the language model of the keyword list to be recognized need to be regenerated. Among them, the phoneme padding model is a solution in the present disclosure for keyword recognition, and is also called the "junk model" or "scrap model". Continuous speech consists of keywords and non-keyword parts other than keywords. The phoneme padding model is mainly used to solve the misrecognition problem of the non-keyword parts in continuous speech. Keyword recognition based on the hidden Markov model usually adopts this method. This method trains a hidden Markov model acoustic model for each keyword to be recognized. A phoneme padding model / junk model is trained for the non-keyword parts. When decoding the keyword recognition path, the keyword speech part is decoded into the corresponding keyword. The non-keyword part is decoded as padding. Thus, only the keywords can be recognized without recognizing the non-keywords.

[0044] The adaptive keyword matching part first calculates the average length of the keywords where l represents the average keyword length. n represents the number of keywords. i represents the serial number. l i represents the length of the i-th keyword. Based on the average keyword length, the length wnd of the matching window and the moving distance rwnd of the window are defined. wnd satisfies 1.5l ≤ wnd ≤ 2l. When a keyword is recognized, rwnd = 0.8l. If no keyword is recognized, then rwnd = 0.4l. The adaptive window further divides the speech to be recognized into smaller speech segments. Thereby improving the response speed and recognition accuracy of the system.

[0045] The keyword recognition part mainly aims at the speech observation sequence O = (o1, o2,..., o M ). The Viterbi algorithm is used to calculate the decoding in parallel in the parallel Chinese and English decoding spaces, and the optimal state sequences P c and P e containing keyword phonemes and non-keyword phonemes of the given speech observation sequence are obtained. The confirmation scores are calculated as follows:

[0046]

[0047] Among them, S1 represents the confirmation score of the speech observation sequence O in the Chinese decoding network space. P represents probability. O represents the speech observation sequence. P c represents the Chinese optimal state sequence. P(P c |O) represents the conditional probability of P c appearing when the speech observation sequence is O. P(P c ) represents the probability of P c appearing in the language model. P(O) represents the probability of the speech observation sequence O. P(O|P c ) represents the conditional probability of O appearing when P c occurs. S2 represents the confirmation score of the speech observation sequence O in the English decoding network space. P e represents the English optimal state sequence. P(P e |O) represents the conditional probability of P e appearing when the speech observation sequence is O. P(P e ) represents the probability of P e appearing in the language model. P(O|P e ) represents the conditional probability of O appearing when P e occurs.

[0048] Among them, P(P c ) and P(P e ) are obtained from the language model. P(O|P c ) and P(O|P e ) are obtained from the hidden Markov model acoustic model. The denominators of the two formulas are the same. Specifically, it is to compare the numerators, that is, to compare which language decoding space has the highest probability of generating the current speech observation sequence. If S1 > S2, it is considered that Chinese is recognized, otherwise it is considered that English is recognized.

[0049] Under the guidance of the dictionary information and the language model, the recognition result post-processing part combines the optimal state path to obtain the recognition result C containing keyword and non-keyword information. Using the keyword list as the evaluation criterion for the phoneme output probability score, the final keyword recognition result W satisfies:

[0050]

[0051] Among them, P represents probability. W represents a keyword. W i represents the i-th keyword in the keyword list. C represents the recognition result obtained by combining the optimal state paths. represents the keyword most likely to be recognized when the recognition result is C.

[0052] The data query module is mainly involved in operations on data records. As an example, this part is implemented using database technology (e.g., sqlite database technology). A query model is defined and a query statement is set. Then, the query statement is concatenated with the user's query conditions to form a more complete and accurate query. In this embodiment, a paging query is also designed, and only 5 data records are displayed on each page. The paging query not only makes the interface more beautiful but also limits the number of data records queried from the database each time. When the amount of data in the database is large, the query result of the database can be obtained quickly.

[0053] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0054] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0055] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. An offline speech keyword recognition method for mixed Chinese and English, comprising: Obtaining a speech digital signal, performing speech activity detection on it to obtain a speech segment to be recognized; Defining an adaptive keyword matching window and segmenting the speech segment to be recognized; Performing feature extraction on the speech segment within the window to obtain a Mel-frequency cepstral coefficient embedding feature vector; Analyzing a custom keyword list in a specific scenario and combining a pre-trained phoneme padding model to obtain a Chinese decoding network space and an English decoding network space for the custom keyword; Sequentially inputting the Mel-frequency cepstral coefficient embedding feature vector into the decoding network space to obtain a recognition result; Post-processing the recognition result to generate a target recognition result as output; Wherein, the analyzing a custom keyword list in a specific scenario and combining a pre-trained phoneme padding model to obtain a Chinese decoding network space and an English decoding network space for the custom keyword includes: Training a hidden Markov model acoustic model with context-dependent phonemes as the basic modeling unit, and constructing a phoneme padding model with an online garbage model. Wherein, a phoneme is the smallest basic unit constituting speech, and the online garbage model directly calculates the local garbage probability score of each speech frame in the phoneme padding model without the need to separately train a garbage model; According to the application requirements of different scenarios, customizing a keyword list, and generating dictionary information on the corresponding relationship between keywords and phonemes in the partitioning method of the Carnegie Mellon University dictionary; Using the keyword text as the language model corpus, language modeling is performed based on the statistical language model. For a given keyword sequence , the trigram language model probability is expressed as follows: , Among them, represents a keyword sequence, represents the th character in the keyword sequence, represents the th character in the keyword sequence, represents the th character in the keyword sequence, represents the length of the keyword sequence, represents the serial number, represents the probability, represents the probability of the keyword sequence appearing in the order, represents under the condition of known and probability, represents the th character in the keyword sequence, represents the th character in the keyword sequence, represents the th to the th probability for consecutive multiplication calculation; The pre-trained phoneme padding model, the dictionary information, and the trigram language model probability together constitute the Chinese decoding network space and the English decoding network space for the custom keyword list. Wherein, when the keyword list is changed, the phonemes constituting the speech do not need to be retrained, and only the dictionary information and the trigram language model probability of the keyword list to be recognized need to be regenerated; 2. The method according to claim 1, wherein The speech activity detection includes: Define the parameter information for voice acquisition, and call the audio processing interface to perform quantization processing on the original voice with the following parameters: the sampling frequency is 16,000 Hz, the number of channels is 1, and the number of voice frames included in each voice block is 1,024, to obtain the voice frame coding information after quantization processing at the nth moment , The combination of voice frame coding information to The original voice block information within the time period ; Calculating the average sound intensity for the quantized speech frame coding information as follows: , Among them, represents the moment, represents the voice frame coding information, represents the voice frame coding information collected at the moment, represents the th bit in the voice frame coding information collected at the moment, represents the th bit in the voice frame coding information collected at the moment, represents the th bit in the voice frame coding information collected at the moment, represents the original voice block information, represents the moment, represents the th moment, represents the th moment, represents the th moment, represents the voice frame coding information collected at the moment, represents the voice frame coding information collected at the moment, represents the average sound intensity which is also used as the silence threshold in the current environment, represents the influence factor, and the specific value is , represents the serial number, represents the serial number, represents the quantity of the voice frame coding information, represents the th bit in the voice frame coding information, represents the th bit in the voice frame coding information collected at the nth moment; Analyzing the change in sound intensity, and dynamically updating the mute threshold when keyword recognition is completed or when there is no sound intensity exceeding the threshold for a long time; 3. The method according to claim 2, wherein, The defining an adaptive keyword matching window includes: Calculating the average keyword length by comparing with the keyword list as: , Among them, represents the average keyword length, represents the number of keywords, represents the serial number, represents the length of the nth keyword; Define the length of the matching window based on the average keyword length and the distance of window movement , Meet , when a keyword is recognized , if no keyword is recognized then .

4. The method according to claim 3, wherein, The performing feature extraction on the speech segment within the window to obtain a Mel-frequency cepstral coefficient embedding feature vector includes: Pre-emphasizing the speech signal within the keyword matching window to compensate for the loss of high-frequency signals during sound propagation; Overlapping and framing the speech signal with a fixed frame length and frame shift to obtain a framed speech signal; Windowing the framed speech signal to obtain a speech signal with enhanced central part and the rest tending to zero; Performing Fourier transform on the windowed speech signal to obtain the linear spectrum of each frame of the speech signal; Inputting the linear spectrum into a Mel-frequency filter bank to obtain a Mel-frequency cepstral coefficient embedding feature vector.

5. According to the method described in claim 4, the sequentially inputting the Mel-frequency cepstral coefficient embedding feature vector into the decoding network space to obtain a recognition result includes: Obtain the Mel-frequency cepstral coefficient embedding feature vectors within the adaptive keyword matching window as the speech observation sequence: , Among them, represents the speech observation sequence, represents the th frame in the speech observation sequence, represents the th frame in the speech observation sequence, represents the th frame in the speech observation sequence; In the Chinese decoding cyberspace and the English decoding cyberspace which constitute a parallel multi - language decoder, for the same speech observation sequence perform parallel decoding using the Viterbi algorithm in the two decoding cyberspaces respectively to obtain the optimal state sequences of the given speech observation sequence that contain keyword phonemes and non - keyword phonemes and , calculate the confirmation scores as follows: , Among them, represents the speech observation sequence the confirmation score in the Chinese decoding network space, represents probability, represents the speech observation sequence, represents the Chinese best state sequence, represents when the speech observation sequence is the conditional probability that occurs, represents the probability that occurs in the language model, represents the speech observation sequence the probability of, represents at the conditional probability that occurs, represents the speech observation sequence the confirmation score in the English decoding network space, represents the English best state sequence, represents when the speech observation sequence is the conditional probability that occurs, represents the probability that occurs in the language model, represents at the conditional probability that occurs; Among them and are obtained from the language model, and are obtained from the Hidden Markov Model acoustic model. The denominators of the two formulas are the same. Specifically, it is to compare the numerators, that is, to compare which language decoding network space has the highest probability of generating the current speech observation sequence. If it is considered that Chinese is recognized, otherwise it is considered that English is recognized.

6. The method according to claim 5, wherein, The post-processing of the recognition result to generate the target recognition result as the output includes: Under the guidance of the dictionary information and the language model, combine the optimal state path to obtain the recognition result containing keyword and non-keyword information; For the recognition result of the non-keyword information, using the keyword list as the evaluation criterion for the phoneme output probability, the target recognition result is obtained as the output, which satisfies: , where represents probability, represents keyword, represents the th keyword in the keyword list, represents the recognition result obtained by combining the optimal state paths, represents that in the case where the recognition result is the most likely recognized keyword.

7. An offline speech keyword recognition system for mixed Chinese and English in a specific scenario applied to the method described in claim 1, comprising: The real-time speech monitoring module is used to monitor the speech signal in the current environment in real time through the microphone; The voice activity detection module is used to detect the speech segment to be recognized in the speech signal; The keyword recognition module is used to judge whether there are keywords in the speech signal; The data record storage domain retrieval module is used to record the information related to the keyword in the database and provide the data query function.

Citation Information

Patent Citations

  • Teacher movement tracing method based on movement detection combining multi-channel fusion

    CN101394479A

  • Voice activation method and system

    CN105374352A

  • Real-time gesture recognition method and system

    CN110163142A

  • Deep neural network construction method for voice command word recognition and recognition method and device

    CN111210815A