Method and device for dynamic monitoring of sensitive words in Tibetan language speech recognition in complex environments

Through multi-dimensional feature fusion scoring and decoding confidence evaluation, the problems of dialect differences and environmental noise in Tibetan speech recognition are solved, and sensitive word recognition with high accuracy and robustness is achieved.

CN120375808BActive Publication Date: 2025-09-23BEIJING WANGZHI TIANYUAN BIG DATA TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510876494.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-23
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In Tibetan speech recognition, the traditional single phoneme matching method cannot adapt to different dialects, resulting in a high misjudgment rate. The special acoustic environment of the plateau and wind noise interference make the recognition of sensitive words more difficult, and existing technologies make it difficult to achieve multi-dimensional fusion scoring.

Method used

By obtaining the Tibetan speech input stream, performing frame processing, and extracting multi-dimensional features, including phoneme-level, semantic-level, scene-level, and emotion-level scores, combined with the Tibetan acoustic model and decoding confidence, a multi-dimensional perception fusion evaluation is performed, and the sensitive word discrimination threshold is dynamically adjusted to achieve graded warning.

Benefits of technology

It significantly improves the accuracy and adaptability of sensitive word recognition, reduces false alarm and missed detection rates, enhances robustness and dynamic decision-making capabilities, and optimizes resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375808B_ABST
    Figure CN120375808B_ABST
Patent Text Reader

Abstract

This application provides a method and device for dynamically monitoring sensitive words in Tibetan speech recognition in complex environments. The method segments the Tibetan speech input stream to be monitored in a complex environment to obtain a framed speech signal of the Tibetan speech in the complex environment. The method then performs multi-dimensional perceptual fusion on the phoneme-level, semantic-level, scene-level, and sentiment-level scores of each candidate sensitive word in the framed speech signal to obtain perceptual fusion features corresponding to the context of each candidate sensitive word. The method also determines the decoding confidence of the Tibetan syllable sequence based on the path stability of the decoding path in speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence. The method then determines the contextual sensitivity of each candidate sensitive word based on the decoding confidence and the perceptual fusion features. The method then uses the contextual sensitivity to provide a graded warning for the Tibetan speech input stream in the complex environment. This approach enables multi-dimensional fusion scoring of sensitive words in Tibetan speech recognition in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of Tibetan speech recognition, and more specifically, to a method and device for dynamically monitoring sensitive words in Tibetan speech recognition in complex environments. Background Art

[0002] As a language with rich dialects and complex phonetic structure, Tibetan's sensitive word recognition needs to comprehensively consider the diversity of speech and the complexity of semantics. In Tibetan speech recognition, the recognition of sensitive words is an important technical challenge. At present, Tibetan speech recognition technology has made certain progress, but sensitive word recognition is still a problem that needs to be solved urgently.

[0003] The pronunciation differences among the three major Tibetan dialects pose a severe challenge to the monitoring of sensitive words. Traditional single phoneme matching methods, due to their lack of adaptability to different dialects, can easily misjudge these legitimate dialect variants as sensitive words, resulting in a high false alarm rate. At the same time, the special acoustic environment of the plateau further exacerbates the difficulty of recognition: the reverberation effect of religious sites can cause blurred syllable boundaries, and legitimate dialects may be misidentified as sensitive words under echo interference, while continuous wind noise can mask the characteristics of voiceless consonants, making it difficult to distinguish minimal opposite pairs and causing missed detections. Therefore, how to achieve multi-dimensional fusion scoring of sensitive words in Tibetan speech recognition in complex environments has become a difficult problem facing the industry. Summary of the Invention

[0004] The present application provides a method and device for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition, which can realize multi-dimensional fusion scoring of sensitive words in Tibetan language complex environment speech recognition.

[0005] In a first aspect, the present application provides a method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition, comprising:

[0006] Obtain the Tibetan speech input stream to be monitored in a complex environment;

[0007] Segmenting the Tibetan speech input stream to obtain a framed speech signal of the Tibetan speech in a complex environment, and then extracting multi-dimensional features of the speech syllables in the context of the framed speech signal;

[0008] Preliminarily matching the framed speech signal with a sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream, and then performing multi-dimensional perceptual fusion on the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features to obtain a perceptual fusion feature of the context corresponding to each candidate sensitive word;

[0009] Performing speech recognition on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, thereby determining a path stability of a decoding path in the speech recognition and a posterior probability of each Tibetan syllable in the Tibetan syllable sequence, and determining a decoding confidence of the Tibetan syllable sequence based on the path stability and each posterior probability;

[0010] Based on the decoding confidence and the various perceptual fusion features, each candidate sensitive word is fused and evaluated to obtain the context sensitivity of each candidate sensitive word, and then the Tibetan speech input stream in a complex environment is graded and warned based on the context sensitivity.

[0011] In some embodiments, extracting the multi-dimensional features of the speech syllables in the context of the framed speech signal specifically includes:

[0012] Set the window radius and window time interval of the time context window;

[0013] For each speech syllable word in the framed speech signal, matching the speech syllable word with a standard sensitive word according to a Tibetan phonetic fuzzy matching algorithm to obtain a phoneme-level score of the speech syllable word;

[0014] Based on the Tibetan semantic model, the semantic rationality of the phonetic syllable words in the current context is evaluated to obtain the semantic level score of the phonetic syllable words;

[0015] The sensitivity threshold is adjusted by the acoustic scene classification label and the speaker's dialect characteristics to obtain the scene-level score of the speech syllable vocabulary;

[0016] The emotional polarity strength of candidate sensitive words is evaluated using the Tibetan emotional dictionary to obtain the emotional level score of the phonetic syllable vocabulary;

[0017] Then, the phoneme-level score, semantic-level score, scene-level score and emotion-level score of each speech syllable word in the framed speech signal are obtained, and the multi-dimensional features of the speech syllable in the context of the framed speech signal are determined through all the phoneme-level scores, semantic-level scores, scene-level scores and emotion-level scores.

[0018] In some embodiments, the framed speech signal is preliminarily matched with a sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream, specifically including:

[0019] Obtain the sensitive word library and candidate sensitive thresholds for Tibetan speech;

[0020] Performing edit distance matching on each speech syllable word in the framed speech signal and the sensitive word library to obtain the Tibetan pinyin edit distance of each speech syllable word;

[0021] A plurality of candidate sensitive words in the Tibetan speech input stream are screened out from the framed speech signal according to the Tibetan pinyin edit distances and the candidate sensitivity threshold.

[0022] In some embodiments, the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features are subjected to multi-dimensional perceptual fusion to obtain the perceptual fusion features of the context corresponding to each candidate sensitive word, specifically including:

[0023] For each candidate sensitive word, determining the fusion information of the candidate sensitive word in the multi-dimensional features;

[0024] Based on the fusion information, the phoneme-level score, semantic-level score, scene-level score and sentiment-level score of the candidate sensitive words are fused into the perceptual fusion features of the context corresponding to the candidate sensitive words, thereby obtaining the perceptual fusion features of the context corresponding to each candidate sensitive word.

[0025] In some embodiments, determining the path stability of the decoding path in speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence specifically includes:

[0026] Determining the path stability interval of the decoding path in speech recognition;

[0027] Counting the proportion of stable frames in the framed voice signal based on the path stability interval;

[0028] Determining the path stability of a decoding path in speech recognition by using the stable frame ratio;

[0029] The posterior probability of each Tibetan syllable in the Tibetan syllable sequence is obtained from the Tibetan acoustic model.

[0030] In some embodiments, determining the decoding confidence of the Tibetan syllable sequence by using the path stability and the respective posterior probabilities specifically includes:

[0031] Initialize a linear probability model based on dynamic weight fusion;

[0032] The path stability is considered as a temporal continuity feature in a linear probability model;

[0033] The individual posterior probabilities are used as instantaneous reliability features in the linear probability model;

[0034] The linear probability model after feature input is used to perform weighted scoring on the multi-frame decoding results of the Tibetan syllable sequence to obtain the decoding confidence of the Tibetan syllable sequence.

[0035] In some embodiments, performing a fusion evaluation on each candidate sensitive word based on the decoding confidence and each perceptual fusion feature to obtain the context sensitivity of each candidate sensitive word specifically includes:

[0036] For each candidate sensitive word, obtain the influence weight of the context on the candidate sensitive word;

[0037] Based on the influence weight, the perception fusion feature of the candidate sensitive word and the decoding confidence are fused to obtain the context sensitivity of the candidate sensitive word, and then the context sensitivity of each candidate sensitive word is obtained.

[0038] In a second aspect, the present application provides a device for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition, comprising:

[0039] An acquisition module is used to obtain the Tibetan speech input stream to be monitored in a complex environment;

[0040] a processing module configured to segment the Tibetan speech input stream to obtain a framed speech signal of the Tibetan speech in a complex environment, and then extract multi-dimensional features of the speech syllables in the context of the framed speech signal;

[0041] The processing module is further configured to perform a preliminary match between the framed speech signal and a sensitive word library of Tibetan speech to obtain a plurality of candidate sensitive words in the Tibetan speech input stream, and then perform multi-dimensional perceptual fusion of the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features to obtain a perceptual fusion feature corresponding to the context of each candidate sensitive word;

[0042] The processing module is further configured to perform speech recognition on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, thereby determining a path stability of a decoding path in the speech recognition and a posterior probability of each Tibetan syllable in the Tibetan syllable sequence, and determining a decoding confidence of the Tibetan syllable sequence based on the path stability and each posterior probability;

[0043] The execution module is used to perform a fusion evaluation on each candidate sensitive word based on the decoding confidence and each perceptual fusion feature to obtain the context sensitivity of each candidate sensitive word, and then use the context sensitivity to perform a graded warning on the Tibetan speech input stream in a complex environment.

[0044] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned method for dynamic monitoring of sensitive words in Tibetan complex environment speech recognition.

[0045] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions or codes. When the instructions or codes are run on a computer, the computer implements the above-mentioned method for dynamic monitoring of sensitive words in Tibetan complex environment speech recognition.

[0046] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0047] The present application provides a method and device for dynamically monitoring sensitive words in Tibetan speech recognition in complex environments, which comprises the following steps: obtaining a Tibetan speech input stream to be monitored in a complex environment; segmenting the Tibetan speech input stream to obtain a framed speech signal of Tibetan speech in the complex environment, and then extracting multi-dimensional features of speech syllables in the context of the framed speech signal; preliminarily matching the framed speech signal with a sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream, and then performing multi-dimensional sensing on the phoneme-level score, semantic-level score, scene-level score, and emotion-level score of each candidate sensitive word in the multi-dimensional features. The method comprises the following steps: performing perceptual fusion on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, and then determining the path stability of the decoding path in the speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence, and determining the decoding confidence of the Tibetan syllable sequence based on the path stability and the posterior probabilities; performing a fusion evaluation on each candidate sensitive word based on the decoding confidence and the perceptual fusion features to obtain the context sensitivity of each candidate sensitive word, and then performing a graded warning on the Tibetan speech input stream in a complex environment based on the context sensitivity.

[0048] It can be seen that in this application, each candidate sensitive word is fused and evaluated based on the decoding confidence and each perceptual fusion feature to obtain the contextual sensitivity of each candidate sensitive word, and then the Tibetan speech input stream in a complex environment is graded and warned by the contextual sensitivity; first, the perceptual fusion feature is determined to obtain a multi-dimensional context representation, which can significantly improve the accuracy and adaptability of sensitive word discrimination. The construction of perceptual fusion features provides a comprehensive contextual representation for Tibetan sensitive word recognition by integrating the four-dimensional scoring of phoneme level, semantic level, scene level and emotion level. The phoneme level scoring is based on the Tibetan pinyin fuzzy matching algorithm, which can capture the acoustic differences between dialect variants and standard pronunciation, and solve the misjudgment problem of single phoneme matching; the semantic level scoring uses the bidirectional long short-term memory network (Bi-LSTM) model to analyze the dependency of syllable sequences and effectively distinguish the pronunciation similarity between homophones and sensitive words; the scene level scoring is combined with the acoustic environment classification to dynamically adjust the threshold for sensitive word discrimination; the emotion level scoring quantifies the emotional polarity of the vocabulary through the Tibetan emotion dictionary. Multi-dimensional fusion can improve the F1 score of sensitive word recognition, significantly reducing the false alarm rate in mixed dialect scenarios. Decoding confidence is then determined to obtain a reliability measure for speech recognition results, thereby enhancing the robustness and dynamic decision-making capabilities of sensitive word monitoring. Decoding confidence is calculated by decoding path stability and syllable posterior probability, reflecting the credibility of speech recognition results. Decoding path stability measures the volatility of the acoustic model output path and can identify low-quality speech segments. Syllable posterior probability quantifies the certainty of each syllable in the joint decoding of the acoustic language model. Combining decoding path stability and syllable posterior probability can automatically filter out low-confidence recognition results caused by environmental noise or dialect differences, thereby reducing the missed detection rate in noisy environments. Dynamic weight allocation ensures the detection priority of key sensitive words. Decoding confidence can also trigger a hierarchical response mechanism: high-confidence sensitive words are warned in real time, while low-confidence results are transferred to a manual review process, achieving optimal resource allocation. In summary, the above scheme can achieve multi-dimensional fusion scoring of sensitive words in Tibetan speech recognition in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0050] Figure 1 This is an exemplary flow chart of a method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition according to some embodiments of the present application;

[0051] Figure 2 is a schematic diagram of the process of Tibetan speech recognition according to some embodiments of the present application;

[0052] Figure 3 is a schematic diagram of a process for determining path stability and posterior probability according to some embodiments of the present application;

[0053] Figure 4 1 is a schematic diagram of a structure of a device for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition according to some embodiments of the present application;

[0054] Figure 5 This is a structural diagram of a computer device for implementing a method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition according to some embodiments of the present application. DETAILED DESCRIPTION

[0055] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0056] refer to Figure 1 This figure is an exemplary flow chart of a method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition according to some embodiments of the present application. The method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition mainly includes the following steps:

[0057] In step 101, a Tibetan speech input stream to be monitored in a complex environment is obtained.

[0058] It should be noted that in this application, the Tibetan speech input stream refers to the continuous Tibetan speech signal collected in real time; the complex environment refers to the speech collection scenario containing background noise (such as wind, human voice, mechanical sound), multi-speaker overlap, dialect variants or emotional intonation changes, which directly affects the robustness of speech recognition; in specific implementation, a microphone array is used to collect continuous Tibetan speech signals in a complex environment as the Tibetan speech input stream to be monitored.

[0059] In some embodiments, reference Figure 2As described, the figure is a flow chart of Tibetan speech recognition according to some embodiments of the present application. The figure describes the overall processing flow of Tibetan speech recognition, which mainly includes the integration and optimization of the acoustic model, pronunciation dictionary and language model. After constructing the search space, Viterbi Beam search is performed based on the input voice file to output the optimal recognition result. Specifically, the system first integrates the acoustic model, pronunciation dictionary and language model, and sequentially performs the determination and minimization steps to optimize the search space structure. Then, the input voice file and the optimized search space are input into the Viterbi Beam search module, and the initialization, judgment score, path clipping and backtracking operations are completed in sequence to obtain the final optimal recognition result. The entire process realizes efficient modeling and accurate recognition of Tibetan speech signals.

[0060] In step 102, the Tibetan speech input stream is segmented to obtain a framed speech signal of the Tibetan speech in a complex environment, and then multi-dimensional features of the speech syllables in the context of the framed speech signal are extracted.

[0061] In some embodiments, segmenting the Tibetan speech input stream to obtain a framed speech signal of Tibetan speech in a complex environment can be achieved by using the following steps:

[0062] Normalizing the Tibetan language speech input stream to obtain a continuous speech stream to be monitored;

[0063] Using a fixed duration to divide the continuous speech stream into a plurality of short-time signal segments;

[0064] The framed speech signal of Tibetan speech in a complex environment is determined through all short-time signal segments.

[0065] It should be noted that, in this application, the term "framed voice signal" refers to a continuous voice stream and a short-term signal segment.

[0066] In the specific implementation, first, the Tibetan speech input stream is normalized to obtain the continuous speech stream to be monitored, which can be achieved in the following way, namely: the Tibetan speech input stream is sampled at a uniform rate and the amplitude is normalized, an anti-aliasing filter is used to ensure signal quality, speech from different sources is uniformly converted to a 16kHz sampling rate, the signal amplitude is adjusted by automatic gain control to avoid numerical overflow or quantization error in subsequent processing, thereby completing the normalization adjustment of the Tibetan speech input stream, and the normalized Tibetan speech input stream is used as the continuous speech stream to be monitored; then, a fixed time is used to The continuous speech stream is divided into multiple short-time signal segments, which can be achieved by adopting the following method: framing the continuous speech stream using an analysis window with a frame length of 25ms and a frame shift of 10ms; each analysis frame is weighted by a Hamming window function to effectively reduce spectral leakage; to maintain continuity, an overlapping area of ​​1ms is retained between adjacent frames, and multiple short-time signal segments are obtained; finally, determining the framed speech signal of Tibetan speech in a complex environment through all the short-time signal segments can be achieved by adopting the following method: taking the collection of all the short-time signal segments as the framed speech signal of Tibetan speech in a complex environment.

[0067] In some embodiments, extracting the multi-dimensional features of the speech syllables in the context of the framed speech signal may be achieved by using the following steps:

[0068] Set the window radius and window time interval of the time context window;

[0069] For each speech syllable word in the framed speech signal, matching the speech syllable word with a standard sensitive word according to a Tibetan phonetic fuzzy matching algorithm to obtain a phoneme-level score of the speech syllable word;

[0070] Based on the Tibetan semantic model, the semantic rationality of the phonetic syllable words in the current context is evaluated to obtain the semantic level score of the phonetic syllable words;

[0071] The sensitivity threshold is adjusted by the acoustic scene classification label and the speaker's dialect characteristics to obtain the scene-level score of the speech syllable vocabulary;

[0072] The emotional polarity strength of candidate sensitive words is evaluated using the Tibetan emotional dictionary to obtain the emotional level score of the phonetic syllable vocabulary;

[0073] Then, the phoneme-level score, semantic-level score, scene-level score and emotion-level score of each speech syllable word in the framed speech signal are obtained, and the multi-dimensional features of the speech syllable in the context of the framed speech signal are determined through all the phoneme-level scores, semantic-level scores, scene-level scores and emotion-level scores.

[0074] It should be noted that, in this application, multi-dimensional features; the time context window represents a fixed time range before and after the current syllable as the center, which is used to capture the semantic association between syllables; the phoneme-level score is a parameter that quantifies the degree of similarity in pronunciation between the current syllable and the standard sensitive word; the semantic-level score is an indicator that reflects the reasonableness of the appearance of the vocabulary in the current context; the scene-level score is the environmental adaptation score that affects the final sensitivity judgment; the emotion-level score is a parameter that quantifies the intensity of emotion carried by the vocabulary.

[0075] In the specific implementation, first, setting the window radius and window time interval of the time context window can be achieved in the following way, namely: setting the window radius of the time context window to the size of the analysis window in the continuous speech stream, obtaining the interval from the start time to the end time of each time context window as the window time interval; secondly, for each speech syllable word in the framed speech signal, matching the speech syllable word with the standard sensitive word according to the Tibetan pinyin fuzzy matching algorithm, and obtaining the phoneme-level score of the speech syllable word can be achieved in the following way, namely: for each speech syllable word in the framed speech signal, matching the speech syllable word with the standard sensitive word according to the Tibetan pinyin fuzzy matching algorithm The phonetic syllable vocabulary is matched with the standard sensitive words, and the matching result is used as the phoneme-level score of the phonetic syllable vocabulary. The Tibetan pinyin fuzzy matching algorithm is to use the number of pinyin modifications required to convert the phonetic syllable vocabulary into the standard sensitive words as the matching result; then, the semantic rationality of the phonetic syllable vocabulary in the current context is evaluated based on the Tibetan semantic model. The semantic-level score of the phonetic syllable vocabulary can be achieved in the following way, namely: based on the pre-trained Tibetan bidirectional long short-term memory network model, the co-occurrence probability of the candidate sensitive words in the front and back windows is calculated, and the semantic deviation is calculated in combination with Tibetan grammatical rules, and the rationality score in the range of 0-1 is output.

[0076] Furthermore, in the specific implementation, the sensitivity threshold is adjusted by the acoustic scene classification label and the speaker dialect feature, and the scene-level score of the speech syllable vocabulary can be obtained in the following way, namely: matching the preset sensitivity coefficient according to the acoustic scene classification label, superimposing the threshold offset corresponding to the speaker dialect feature, and dividing the product of the sensitivity coefficient and the threshold offset by the sensitivity threshold as the scene-level score of the speech syllable vocabulary; then, the emotional polarity intensity of the candidate sensitive words is evaluated through the Tibetan emotional dictionary, and the emotional level score of the speech syllable vocabulary can be obtained in the following way, namely: querying the Tibetan emotional dictionary to obtain the basic emotional polarity and intensity value of the candidate sensitive words, combining the emotional word density in the time context window for weighting, and finally outputting the normalized emotional polarity. Sentiment tendency score, wherein the number of occurrences of candidate sensitive words in the time context window can be counted as the sentiment word density; finally, the phoneme-level score, semantic-level score, scene-level score and sentiment-level score of each speech syllable word in the framed speech signal are obtained, and the multi-dimensional features of the speech syllable in the context in the framed speech signal are determined by all the phoneme-level scores, semantic-level scores, scene-level scores and sentiment-level scores. This can be achieved in the following way, namely: the phoneme-level score, semantic-level score, scene-level score and sentiment-level score of each speech syllable word in the framed speech signal can be obtained by the above method, so that the set of phoneme-level score, semantic-level score, scene-level score and sentiment-level score is used as the multi-dimensional feature of the speech syllable in the context in the framed speech signal.

[0077] It should be noted that in this application, the Tibetan bidirectional long short-term memory network model is a deep learning architecture designed specifically for Tibetan sequence modeling. It captures contextual semantic associations through bidirectional temporal analysis, and uses two forward and backward long short-term memory networks to process speech sequences respectively. The forward network analyzes the syllable stream in chronological order and learns the relationship between the current word and the previous text; the backward network processes in reverse order to capture the association between the current word and the following text. After the hidden layer outputs of the two networks are spliced, the co-occurrence probability of each candidate sensitive word in the specified window is calculated through the fully connected layer. The Tibetan grammatical rules are learned in the pre-training stage. During actual reasoning, the predicted probability is compared with the grammatical rule library, and the candidate words that violate the conventional grammatical structure are scored lower. Finally, a semantic rationality score of 0-1 is output through normalization. The higher the semantic rationality score, the greater the possibility that the word appears in the context and conforms to the Tibetan grammatical norms.

[0078] In step 103, the framed speech signal is preliminarily matched with the sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream, and then the phoneme-level score, semantic-level score, scene-level score and emotion-level score of each candidate sensitive word in the multi-dimensional features are subjected to multi-dimensional perceptual fusion to obtain the perceptual fusion features of the corresponding context of each candidate sensitive word.

[0079] In some embodiments, the framed speech signal is preliminarily matched with a sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream by using the following steps:

[0080] Obtain the sensitive word library and candidate sensitive thresholds for Tibetan speech;

[0081] Performing edit distance matching on each speech syllable word in the framed speech signal and the sensitive word library to obtain the Tibetan pinyin edit distance of each speech syllable word;

[0082] A plurality of candidate sensitive words in the Tibetan speech input stream are screened out from the framed speech signal according to the Tibetan pinyin edit distances and the candidate sensitivity threshold.

[0083] It should be noted that in this application, candidate sensitive words are syllable words that are initially matched successfully in the Tibetan voice input stream; the sensitive word library refers to a predefined set of Tibetan sensitive words, which contains standard spellings, common variants and pinyin transliterations; the candidate sensitive threshold is the similarity critical value for determining whether a syllable is a sensitive word, and if it is lower than the candidate sensitive threshold, it is considered a non-sensitive word; the Tibetan pinyin edit distance indicates the degree of difference between the syllable word and the pinyin of the sensitive word, which is calculated by the number of character addition, deletion and modification operations. The smaller the Tibetan pinyin edit distance, the more similar it is.

[0084] In the specific implementation, first, the sensitive word library and candidate sensitive threshold of Tibetan speech are obtained, which can be achieved in the following way, namely: the sensitive word library and candidate sensitive threshold of Tibetan speech are obtained from the console of the Tibetan recognition device; then, each speech syllable word in the framed speech signal is matched with the sensitive word library for editing distance, and the Tibetan pinyin editing distance of each speech syllable word is obtained by the following way, namely: for each speech syllable word in the framed speech signal, the speech syllable word is converted into pinyin, and the dynamic programming algorithm is used to calculate the pinyin editing distance of each word with the sensitive word library one by one. The calculation can be carried out according to the characteristics of Tibetan. Optimization, setting a higher weight for initial consonant difference, followed by final vowel difference, and the lowest weight for tone difference, recording the minimum value of the pinyin edit distance as the Tibetan pinyin edit distance of the speech syllable vocabulary, and the Tibetan pinyin edit distance of each speech syllable vocabulary can be obtained by the above method; finally, screening out multiple candidate sensitive words in the Tibetan speech input stream from the framed speech signal through each Tibetan pinyin edit distance and the candidate sensitive threshold can be achieved in the following way, namely: taking the speech syllable vocabulary whose Tibetan pinyin edit distance in the framed speech signal is less than the candidate sensitive threshold as the candidate sensitive word, and obtaining multiple candidate sensitive words in the Tibetan speech input stream.

[0085] In some embodiments, the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features are subjected to multi-dimensional perceptual fusion to obtain the perceptual fusion features of the context corresponding to each candidate sensitive word. This can be achieved by the following steps:

[0086] For each candidate sensitive word, determining the fusion information of the candidate sensitive word in the multi-dimensional features;

[0087] Based on the fusion information, the phoneme-level score, semantic-level score, scene-level score and sentiment-level score of the candidate sensitive words are fused into the perceptual fusion features of the context corresponding to the candidate sensitive words, thereby obtaining the perceptual fusion features of the context corresponding to each candidate sensitive word.

[0088] It should be noted that, in this application, the perceptual fusion feature is an indicator that reflects the overall sensitivity of the candidate sensitive word in the context; in specific implementation, first, for each candidate sensitive word, determining the fusion information of the candidate sensitive word in the multi-dimensional feature can be implemented in the following way, namely: obtaining the dimension weights of the candidate sensitive word in each dimension from the console of the Tibetan language recognition device, wherein the phoneme dimension weight defaults to 0.4, the semantic dimension weight defaults to 0.3, the scene dimension weight defaults to 0.2, and the emotion dimension weight defaults to 0.1, thereby taking the set of all dimension weights as the fusion information of the candidate sensitive word in the multi-dimensional feature, and the fusion information represents The degree of influence of each dimension on the overall sensitivity of the candidate sensitive word in the context; then, based on the fusion information, the phoneme-level score, semantic-level score, scene-level score and sentiment-level score of the candidate sensitive word are fused into the perceptual fusion feature of the candidate sensitive word corresponding context, and then the perceptual fusion feature of each candidate sensitive word corresponding context is obtained. This can be achieved in the following way, namely: using the weights of each dimension in the fusion information to calculate the weighted sum of the phoneme-level score, semantic-level score, scene-level score and sentiment-level score as the perceptual fusion feature of the candidate sensitive word corresponding context. The perceptual fusion feature of each candidate sensitive word corresponding context can be obtained by the above method.

[0089] In step 104, speech recognition is performed on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, and then the path stability of the decoding path in the speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence are determined. The decoding confidence of the Tibetan syllable sequence is determined based on the path stability and each posterior probability.

[0090] In some embodiments, a Tibetan acoustic model is used to perform speech recognition on the framed speech signal to generate a Tibetan syllable sequence; it should be noted that in this application, the Tibetan acoustic model is a statistical model specially trained for Tibetan speech characteristics, and its core function is to map the acoustic features of the speech signal into corresponding Tibetan phonemes or syllable sequences. This Tibetan acoustic model uses a hybrid modeling approach that combines a deep neural network with a hidden Markov model (HMM). The neural network is responsible for learning the common acoustic patterns of the three major Tibetan dialects (U-Tsang, Amdo, and Kham) and extracting discriminative feature representations through multi-layer nonlinear transformations. The HMM models the temporal dynamic characteristics of Tibetan phonemes. Each phoneme corresponds to an HMM state, and the temporal relationship between phonemes is described by the state transition probability. The maximum likelihood criterion is used during training. Using labeled Tibetan speech data, the network parameters are optimized through the backpropagation algorithm, enabling the Tibetan acoustic model to accurately output the posterior probability that each frame of speech belongs to each Tibetan phoneme. During the decoding stage, the Tibetan acoustic model works in conjunction with the Tibetan language model, combining the Viterbi algorithm to search for the optimal syllable path, achieving robust recognition in complex acoustic environments (such as plateau noise and reverberation).

[0091] In some embodiments, the path stability of the decoding path in speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence are determined, referring to Figure 3 As described above, this figure is a schematic diagram of the process of determining path stability and posterior probability in some embodiments of the present application. In this embodiment, determining path stability and posterior probability can be achieved by using the following steps:

[0092] In step 1041, a path stability interval of a decoding path in speech recognition is determined;

[0093] In step 1042, a proportion of stable frames in the framed speech signal is counted based on the path stability interval;

[0094] In step 1043, the path stability of the decoding path in speech recognition is determined according to the stable frame ratio;

[0095] In step 1044, the posterior probability of each Tibetan syllable in the Tibetan syllable sequence is obtained from the Tibetan acoustic model.

[0096] It should be noted that in this application, the posterior probability represents the probability value that the current speech frame corresponds to the specified Tibetan syllable, and the posterior probability reflects the confidence of single-frame recognition; path stability is an indicator that comprehensively measures the consistency of recognition results in the decoding path. The higher the path stability, the more reliable the recognition result; path stability interval; stable frame ratio refers to the ratio of the number of stable frames to the total number of frames. The stable frame ratio can quantify the recognition stability of the entire framed speech signal.

[0097] In specific implementation, first, determining the path stability interval of the decoding path in speech recognition can be achieved in the following manner, namely: setting the minimum stable interval length to 3 frames, and when the same syllable is recognized in 3 consecutive frames in speech recognition, it is marked as a stable interval, and the interval boundary can be dynamically adjusted in a sliding window manner to ensure that all stable segments can be captured. For each detected stable interval, its starting frame, ending frame and corresponding syllable label are recorded to obtain the path stability interval of the decoding path in speech recognition; secondly, counting the proportion of stable frames in the framed speech signal based on the path stability interval can be achieved in the following manner, namely: counting the number of frames of the framed speech signal in the path stability interval, and calculating the ratio of the number of frames to the total number of frames of the framed speech signal as the proportion of stable frames in the framed speech signal; then, through the stable frame The path stability of the decoding path in speech recognition can be determined by the following method, namely: using a normalization algorithm (for example: minimum-maximum normalization) to normalize the stable frame ratio and use it as the path stability of the decoding path in speech recognition; finally, obtaining the posterior probability of each Tibetan syllable in the Tibetan syllable sequence from the Tibetan acoustic model can be achieved by the following method, namely: for each Tibetan syllable in the Tibetan syllable sequence, obtain the posterior probability value of the Tibetan syllable from the Tibetan acoustic model; for the frames within the path stability interval, calculate the average of the posterior probability values ​​of all frames in the path stability interval where the Tibetan syllable is located as the posterior probability of the Tibetan syllable; for the frames not in the path stability interval, directly use the single-frame probability value as the posterior probability of the Tibetan syllable. The posterior probability of each Tibetan syllable in the Tibetan syllable sequence can be obtained by the above method.

[0098] In some embodiments, determining the decoding confidence of the Tibetan syllable sequence by using the path stability and each posterior probability can be achieved by using the following steps:

[0099] Initialize a linear probability model based on dynamic weight fusion;

[0100] The path stability is considered as a temporal continuity feature in a linear probability model;

[0101] The individual posterior probabilities are used as instantaneous reliability features in the linear probability model;

[0102] The linear probability model after feature input is used to perform weighted scoring on the multi-frame decoding results of the Tibetan syllable sequence to obtain the decoding confidence of the Tibetan syllable sequence.

[0103] It should be noted that the decoding confidence in this application is a measure of the reliability of Tibetan syllable sequence recognition. Specifically, when decoding the framed speech signal through the Tibetan acoustic model, the reliability measure value is calculated by the dynamic weight fusion model based on the stability of the decoding path and the posterior probability of the Tibetan syllable sequence. In this application, the linear probability model is a mathematical model based on weighted summation, which is used to fuse multi-dimensional features and output a probabilistic score. In the weighted score, the linear probability model takes the time continuity feature) and the instantaneous reliability feature as two key input dimensions, and performs linear weighting through preset weight coefficients. The two features are first normalized separately to eliminate the dimensional difference. Then, according to the characteristics of Tibetan speech, a fixed weight of 0.4 is assigned to temporal continuity and a weight of 0.6 is assigned to instantaneous reliability to reflect the dominant role of single-frame recognition confidence. Finally, the decoding confidence is calculated through the weighted summation formula, and the result is mapped to the interval [0,1] using normalization to form an output with probabilistic interpretation. The linear characteristics of the linear probability model ensure computational efficiency and are suitable for real-time processing scenarios. At the same time, the weight coefficient can be dynamically adjusted according to different dialect environments, taking into account both flexibility and accuracy.

[0104] In step 105, a fusion evaluation is performed on each candidate sensitive word based on the decoding confidence and each perceptual fusion feature to obtain the context sensitivity of each candidate sensitive word, and then a graded warning is performed on the Tibetan speech input stream in a complex environment based on the context sensitivity.

[0105] In some embodiments, performing a fusion evaluation on each candidate sensitive word based on the decoding confidence and each perceptual fusion feature to obtain the context sensitivity of each candidate sensitive word can be achieved by using the following steps:

[0106] For each candidate sensitive word, obtain the influence weight of the context on the candidate sensitive word;

[0107] Based on the influence weight, the perception fusion feature of the candidate sensitive word and the decoding confidence are fused to obtain the context sensitivity of the candidate sensitive word, and then the context sensitivity of each candidate sensitive word is obtained.

[0108] In a specific implementation, first, for each candidate sensitive word, obtaining the influence weight of the context on the candidate sensitive word can be achieved in the following manner, namely: for each candidate sensitive word, obtaining the influence weight of the context on the candidate sensitive word from the console of the Tibetan language recognition device; then, based on the influence weight, the perceptual fusion feature of the candidate sensitive word and the decoding confidence are fused to obtain the context sensitivity of the candidate sensitive word, and then obtaining the context sensitivity of each candidate sensitive word can be achieved in the following manner, namely: using the fusion formula for feature fusion, namely: context sensitivity = perceptual fusion feature * influence weight + decoding confidence * (1-influence weight), to obtain the context sensitivity of the candidate sensitive word. The context sensitivity of each candidate sensitive word can be obtained in the above manner.

[0109] In some embodiments, the graded warning of the Tibetan voice input stream in a complex environment through the context sensitivity can be achieved in the following manner, namely: obtaining a warning level mapping table from the console of the Tibetan language recognition device, obtaining the warning level corresponding to the context sensitivity from the warning level mapping table, and then using the warning level as the warning level of the Tibetan voice input stream in the complex environment.

[0110] In addition, in another aspect of the present application, in some embodiments, the present application provides a dynamic monitoring device for sensitive words in Tibetan complex environment speech recognition, referring to Figure 4 This figure is a schematic diagram of the structure of a device for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition according to some embodiments of the present application. The device for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described as follows:

[0111] Acquisition module 201, in this application, acquisition module 201 is mainly used to acquire the Tibetan speech input stream to be monitored in a complex environment;

[0112] Processing module 202, in this application, is used to segment the Tibetan speech input stream to obtain a framed speech signal of Tibetan speech in a complex environment, and then extract multi-dimensional features of the speech syllables in the context of the framed speech signal;

[0113] It should be noted that the processing module 202 is further configured to perform a preliminary match between the framed speech signal and a sensitive word library of Tibetan speech to obtain a plurality of candidate sensitive words in the Tibetan speech input stream, and then perform multi-dimensional perceptual fusion on the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features to obtain a perceptual fusion feature corresponding to the context of each candidate sensitive word;

[0114] In addition, the processing module 202 is further configured to perform speech recognition on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, thereby determining the path stability of a decoding path in the speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence, and determining the decoding confidence of the Tibetan syllable sequence based on the path stability and each posterior probability;

[0115] Execution module 203. In this application, execution module 203 is mainly used to perform a fusion evaluation on each candidate sensitive word based on the decoding confidence and each perceptual fusion feature, obtain the context sensitivity of each candidate sensitive word, and then use the context sensitivity to perform graded warning on the Tibetan speech input stream in a complex environment.

[0116] The above describes in detail the examples of the method and device for dynamic monitoring of sensitive words in Tibetan complex environment speech recognition provided by the embodiments of the present application. It can be understood that, in order to realize the above functions, the corresponding device includes a hardware structure and / or software module corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to realize the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0117] In some embodiments, the present application also provides a computer device, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned method for dynamic monitoring of sensitive words in Tibetan complex environment speech recognition.

[0118] In some embodiments, reference Figure 5 The dotted line in the figure indicates that the unit or module is optional. The figure is a schematic diagram of the structure of a computer device for implementing a method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition according to an embodiment of the present application. The method for dynamically monitoring sensitive words in Tibetan language complex environment speech recognition described in the above embodiment can be Figure 5 The computer device shown in the figure is implemented, and the computer device includes at least one processor 301, a memory 302 and at least one communication unit 305. The computer device can be a terminal device, a server or a chip.

[0119] The processor 301 may be a general-purpose processor or a dedicated processor. For example, the processor 301 may be a central processing unit (CPU), which may be used to control the computer device, execute software programs, and process data from the software programs. The computer device may also include a communication unit 305 for inputting (receiving) and outputting (transmitting) signals.

[0120] For example, the computer device may be a chip, the communication unit 305 may be an input and / or output circuit of the chip, or the communication unit 305 may be a communication interface of the chip, and the chip may be a component of a terminal device, a network device, or other device.

[0121] For another example, the computer device may be a terminal device or a server, and the communication unit 305 may be a transceiver of the terminal device or the server, or the communication unit 305 may be a transceiver circuit of the terminal device or the server.

[0122] The computer device may include one or more memories 302, on which a program 304 is stored. The program 304 can be executed by the processor 301 to generate instructions 303, so that the processor 301 executes the method described in the above method embodiment according to the instructions 303. Optionally, data (such as a target audit model) can also be stored in the memory 302. Optionally, the processor 301 can also read data stored in the memory 302. The data can be stored at the same storage address as the program 304, or at a different storage address from the program 304.

[0123] The processor 301 and the memory 302 may be provided separately or integrated together, for example, integrated on a system on chip (SOC) of a terminal device.

[0124] It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or software-based instructions in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.

[0125] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] For example, in some embodiments, the present application also provides a computer-readable storage medium, which stores instructions or codes. When the instructions or codes are run on a computer, the computer implements the above-mentioned method for dynamic monitoring of sensitive words in Tibetan complex environment speech recognition.

[0127] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0128] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for dynamic monitoring of sensitive words in Tibetan language complex environment speech recognition, characterized in that: The steps include: Obtain the Tibetan speech input stream to be monitored in a complex environment; Segmenting the Tibetan speech input stream to obtain a framed speech signal of the Tibetan speech in a complex environment, and then extracting multi-dimensional features of the speech syllables in the context of the framed speech signal; Preliminarily matching the framed speech signal with a sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream, and then performing multi-dimensional perceptual fusion on the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features to obtain a perceptual fusion feature of the context corresponding to each candidate sensitive word; Performing speech recognition on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, thereby determining a path stability of a decoding path in the speech recognition and a posterior probability of each Tibetan syllable in the Tibetan syllable sequence, and determining a decoding confidence of the Tibetan syllable sequence based on the path stability and each posterior probability; Based on the decoding confidence and the various perceptual fusion features, each candidate sensitive word is fused and evaluated to obtain the context sensitivity of each candidate sensitive word, and then the Tibetan speech input stream in a complex environment is graded and warned based on the context sensitivity.

2. The method according to claim 1, wherein Extracting the multi-dimensional features of the speech syllables in the context of the framed speech signal specifically includes: Set the window radius and window time interval of the time context window; For each speech syllable word in the framed speech signal, matching the speech syllable word with a standard sensitive word according to a Tibetan pinyin fuzzy matching algorithm to obtain a phoneme-level score of the speech syllable word, wherein when matching the speech syllable word with the standard sensitive word according to the Tibetan pinyin fuzzy matching algorithm, the Tibetan pinyin fuzzy matching algorithm is to calculate the number of pinyin modifications required to convert the speech syllable word into a standard sensitive word as a matching result, and use the matching result as the phoneme-level score of the speech syllable word; Based on the Tibetan semantic model, the semantic rationality of the phonetic syllable words in the current context is evaluated to obtain the semantic level score of the phonetic syllable words; The sensitivity threshold is adjusted by the acoustic scene classification label and the speaker's dialect characteristics to obtain the scene-level score of the speech syllable vocabulary; The emotional polarity strength of candidate sensitive words is evaluated using the Tibetan emotional dictionary to obtain the emotional level score of the phonetic syllable vocabulary; Then, the phoneme-level score, semantic-level score, scene-level score and emotion-level score of each speech syllable word in the framed speech signal are obtained, and the multi-dimensional features of the speech syllable in the context of the framed speech signal are determined through all the phoneme-level scores, semantic-level scores, scene-level scores and emotion-level scores.

3. The method according to claim 1, wherein The framed speech signal is preliminarily matched with a sensitive word library of Tibetan speech to obtain multiple candidate sensitive words in the Tibetan speech input stream, specifically including: Obtain the sensitive word library and candidate sensitive thresholds for Tibetan speech; Performing edit distance matching on each speech syllable word in the framed speech signal and the sensitive word library to obtain the Tibetan pinyin edit distance of each speech syllable word; A plurality of candidate sensitive words in the Tibetan speech input stream are screened out from the framed speech signal according to the Tibetan pinyin edit distances and the candidate sensitivity threshold.

4. The method according to claim 1, wherein The phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features are subjected to multi-dimensional perceptual fusion to obtain the perceptual fusion features of the context corresponding to each candidate sensitive word. Specifically, the perceptual fusion features include: For each candidate sensitive word, determining the fusion information of the candidate sensitive word in the multi-dimensional features; Based on the fusion information, the phoneme-level score, semantic-level score, scene-level score and sentiment-level score of the candidate sensitive words are fused into the perceptual fusion features of the context corresponding to the candidate sensitive words, thereby obtaining the perceptual fusion features of the context corresponding to each candidate sensitive word.

5. The method according to claim 1, wherein Determining the path stability of the decoding path in speech recognition and the posterior probability of each Tibetan syllable in the Tibetan syllable sequence specifically includes: Determining the path stability interval of the decoding path in speech recognition; Counting the proportion of stable frames in the framed voice signal based on the path stability interval; Determining the path stability of a decoding path in speech recognition by using the stable frame ratio; The posterior probability of each Tibetan syllable in the Tibetan syllable sequence is obtained from the Tibetan acoustic model.

6. The method according to claim 1, wherein Determining the decoding confidence of the Tibetan syllable sequence by using the path stability and the respective posterior probabilities specifically includes: Initialize a linear probability model based on dynamic weight fusion; The path stability is considered as a temporal continuity feature in a linear probability model; The individual posterior probabilities are used as instantaneous reliability features in the linear probability model; The linear probability model after feature input is used to perform weighted scoring on the multi-frame decoding results of the Tibetan syllable sequence to obtain the decoding confidence of the Tibetan syllable sequence.

7. The method according to claim 1, wherein Based on the decoding confidence and each perceptual fusion feature, each candidate sensitive word is subjected to a fusion evaluation to obtain the context sensitivity of each candidate sensitive word, specifically including: For each candidate sensitive word, obtain the influence weight of the context on the candidate sensitive word; Based on the influence weight, the perception fusion feature of the candidate sensitive word and the decoding confidence are fused to obtain the context sensitivity of the candidate sensitive word, and then the context sensitivity of each candidate sensitive word is obtained.

8. A dynamic monitoring device for sensitive words in Tibetan language complex environment speech recognition, characterized by: include: An acquisition module is used to obtain the Tibetan speech input stream to be monitored in a complex environment; a processing module configured to segment the Tibetan speech input stream to obtain a framed speech signal of the Tibetan speech in a complex environment, and then extract multi-dimensional features of the speech syllables in the context of the framed speech signal; The processing module is further configured to perform a preliminary match between the framed speech signal and a sensitive word library of Tibetan speech to obtain a plurality of candidate sensitive words in the Tibetan speech input stream, and then perform multi-dimensional perceptual fusion of the phoneme-level score, semantic-level score, scene-level score, and sentiment-level score of each candidate sensitive word in the multi-dimensional features to obtain a perceptual fusion feature corresponding to the context of each candidate sensitive word; The processing module is further configured to perform speech recognition on the framed speech signal using a Tibetan acoustic model to generate a Tibetan syllable sequence, thereby determining a path stability of a decoding path in the speech recognition and a posterior probability of each Tibetan syllable in the Tibetan syllable sequence, and determining a decoding confidence of the Tibetan syllable sequence based on the path stability and each posterior probability; The execution module is used to perform a fusion evaluation on each candidate sensitive word based on the decoding confidence and each perceptual fusion feature to obtain the context sensitivity of each candidate sensitive word, and then use the context sensitivity to perform a graded warning on the Tibetan speech input stream in a complex environment.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the method for dynamic monitoring of sensitive words in Tibetan complex environment speech recognition according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions or codes, which, when executed on a computer, enable the computer to implement the method for dynamically monitoring sensitive words in Tibetan complex environment speech recognition according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Uygur language sensitive word filtration system

    CN104504091A

  • Automated speech recognition using dynamically adjustable listening timeout

    CN110491414A