Language information dynamic detection method based on deep neural network

By adopting a dynamic detection method of language information based on deep neural networks in speech processing technology, combining attention mechanism and dynamic window selection, the recognition accuracy problem of traditional methods when dealing with frequent language switching is solved, and higher language switching detection accuracy and recognition accuracy are achieved.

CN120148478AActive Publication Date: 2025-06-13GLOBAL TONE COMM TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510303486.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Traditional language classification methods cannot effectively handle frequent and short-term language switching, resulting in insufficient accuracy of language recognition in bilingual mixed pronunciation streams.

Method used

The dynamic detection method of language information based on deep neural network is adopted, combining attention mechanism, multi-frame voice information, dynamic window selection, long-term information statistics and recurrent neural network, and dynamic adjustment of feature windows is made to improve the detection accuracy of language switching points.

Benefits of technology

It improves the detection accuracy of language switching points and the accuracy of language recognition in bilingual mixed speech streams, and is suitable for bilingual mixed speech streams, and has high flexibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148478A_ABST
    Figure CN120148478A_ABST
Patent Text Reader

Abstract

The invention provides a language information dynamic detection method based on a deep neural network, and relates to the technical field of voice processing, and the method comprises the steps: obtaining a to-be-detected mixed language voice stream, and extracting original acoustic features; generating an attention vector of the current time step based on an attention mechanism, and dynamically selecting a voice feature frame sequence window of the current time step in combination with unidirectional time selection and limitation of a specific time span; multiplying the dynamic window signal by the original acoustic feature signal to generate a local acoustic feature of the current time step detection language information; normalizing the local acoustic features of any length into fixed dimension features, inputting the fixed dimension features into a deep neural network classifier, and outputting a language probability value corresponding to the voice features of the current time step; based on the dynamic window information and the language probability value, the starting and ending time of each voice segment in the mixed voice stream and the corresponding language label are output, the time point of language switching is determined, and the method is particularly suitable for language switching recognition in the bilingual mixed voice stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly to a method for dynamically detecting language information based on a deep neural network. Background Art

[0002] For a bilingual mixed speech stream, it is very difficult to identify the language because the start and end positions of the languages are unknown. For the problem of language identification in a bilingual mixed speech stream with word-level switching, second-language words are usually intermittently inserted to replace first-language words during the communication in the first language. In this case of word-level switching, on the one hand, the position of the language-switching time point is not known in advance. On the other hand, since the language information switches frequently and most of the second-language speech stream information has a short duration, the traditional language classification method based on the entire test speech stream is also not applicable. For a bilingual mixed speech stream, it is not feasible to directly construct a classifier to discriminate the language category for each input frame of speech features. Because there will actually be a large error when the frame-level language classifier makes a discrimination. For a single-language stream, a post-processing smoothing operation can be performed on the discrimination result sequence to obtain a language identification result for this single-language stream. However, when two languages are mixed in the speech stream and the time point position of the language cut-off point is unknown, the output result sequence of the frame-level classifier with a large error will show a continuous jump of language labels. In this case, the post-processing smoothing cannot solve the problem, and an ideal language classification result cannot be obtained.

[0003] To solve the problem that it is impossible to effectively classify and identify the language when other languages are mixed in one speech. Bilingual mixed speech streams can be divided into sentence-level switching bilingual mixed speech streams and word-level switching bilingual mixed speech streams according to the existence mode of information. For the problem of language identification in a sentence-level switching bilingual mixed speech stream, it generally appears in the application scenario of bilingual identification. Usually, a conversation contains several sentences, and each sentence can freely choose a certain language. In this case, since the position of the language switch is not known in advance, and the start and end time point positions of the language are also not known in advance, and the traditional language classification is carried out based on the assumption that the entire test speech stream is of the same language, the traditional language classification method is not applicable to the case where language switching exists in the same speech stream. Summary of the Invention

[0004] In view of this, the object of the present invention is to propose a dynamic language information detection method based on a deep neural network, which is particularly suitable for language switch recognition in a bilingual mixed speech stream. By combining the attention mechanism, multi-frame speech information, dynamic window selection, long-term information statistics, and recurrent neural network, the detection accuracy of language switch points and the accuracy of language recognition in the bilingual mixed speech stream are improved, and the problem that traditional language classification methods cannot effectively handle frequent and short language switches is solved.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] Based on the above object, in the first aspect, the present invention provides a dynamic language information detection method based on a deep neural network, including the following steps:

[0007] Obtain the mixed-language speech stream to be detected and extract the original acoustic features;

[0008] Generate the attention vector at the current time step based on the attention mechanism, and dynamically select the speech feature frame sequence window at the current time step by combining unidirectional time selection and the limitation of a specific time span;

[0009] Multiply the dynamic window signal by the original acoustic feature signal to generate the local acoustic features for detecting the language information at the current time step;

[0010] Regularize the local acoustic features of any length into fixed-dimensional features, and use the fixed-dimensional features as the input for the neural network language classification;

[0011] Input the fixed-dimensional features into the deep neural network classifier to output the language probability value corresponding to the speech features at the current time step;

[0012] Based on the dynamic window information and the language probability value, output the start and end times and the corresponding language labels of each speech segment in the mixed speech stream, and determine the time points of language switches.

[0013] As a further solution of the present invention, the attention vector at the current time step generated based on the attention mechanism is:

[0014] a i =[a i,1 ,...,a i,j ,...,a i,T

[0015] Generate the speech feature frame sequence window w i at the current time step according to the attention vector a i , where:

[0016] w i =f(a i ) ​

[0017] Introduce a unidirectional time selection limit, where the time point positions of the voice feature frame sequence window conform to the monotonically increasing property. When calculating the position information of the current window w i the position information of the previous time step window w i-1 is added so that w i is later than w i-1 to appear:

[0018] w i = f(a i , w i-1 )

[0019] And add a limit on the finite time span. When calculating the position information of the current window w i the window length limit parameter L w is added so that the window shape does not appear scattered:

[0020] w i = f(a i , w i-1 , L w ).

[0021] As a further solution of the present invention, multiply the dynamic window signal by the original acoustic feature signal to generate the local acoustic feature for detecting the language information at the current time step, where the local acoustic feature is:

[0022]

[0023] In the formula, X is the original acoustic feature input, and X' is the acoustic feature sequence selected in real time, serving as the feature input for language recognition at the current time step.

[0024] As a further solution of the present invention, regularize the local acoustic features of any length into fixed-dimension features. The regularization methods include two ways: long-term information statistics and recurrent neural network; among them, the regularization method of long-term information statistics includes:

[0025] Calculate the mean and variance statistics of the frame sequence selected in real time;

[0026] Combine the mean, variance statistics with the frame sequence selected in real time as the fixed-dimension feature.

[0027] As a further solution of the present invention, the calculation formula of the language feature vector formed by long-term information statistics is as follows:

[0028] v t = f(Append(μ t , σ t ))

[0029]

[0030] Among them, X' i represents the i-th frame acoustic feature in the frame sequence that is selected in real time. The length of the frame sequence selected in real time is T, μ t and σ t respectively represent the mean and variance statistics. Finally, X” t =[μ t ,σ t is combined as the input feature of the language classifier at the t-th time step; v t represents the language feature vector of the long-term information statistical component; T represents the length of the frame sequence selected in real time; t + T represents the cut-off time step position of T frames counted backward starting from the t-th time step.

[0031] As a further solution of the present invention, the regularization method of the recurrent neural network includes:

[0032] Generating fixed-dimension features based on the recurrent neural network;

[0033] Taking the output of the recurrent neural network as the fixed-dimension features;

[0034] Among them, the calculation formula of the fixed-dimension features is as follows

[0035] s t =f(UX′ t +Ws t-1 )

[0036] X” t =softmax(Vs t )

[0037] Among them, U, W, and V are the input mapping matrix, the hidden layer mapping matrix, and the output mapping matrix respectively. Finally, X” t is used as the input feature of the language classifier at the t-th time step; s t represents the hidden layer state of the recurrent neural network at the t-th time step; X' t represents the input feature at the t-th time step; s t-1 represents the hidden layer state of the recurrent neural network at the (t - 1)-th time step, which is used to participate in the calculation of the hidden layer state at the current time step; Vs t represents the result of the linear transformation of the hidden layer state s t after passing through the output mapping matrix V, and it is the intermediate representation before output.

[0038] As a further solution of the present invention, when inputting fixed-dimensional features into the deep neural network classifier and outputting the language probability value corresponding to the speech feature at the current time step, the probability value that the speech feature at the current time step is judged to be the corresponding language is determined according to the fixed-dimensional language feature; wherein, a deep neural network is used as the language classifier, and the output of the neural network language classifier adopted has two output endpoints, and the output results respectively represent the probability value that the speech feature at the current time step is judged to be the corresponding language.

[0039] As a further solution of the present invention, the deep neural network classifier is a multi-level structure, including:

[0040] An input layer for receiving fixed-dimensional features;

[0041] A hidden layer that extracts high-level language features using a feedforward neural network or a bidirectional LSTM;

[0042] An output layer that outputs the probability value of the language corresponding to the speech feature at the current time step through Softmax.

[0043] As a further solution of the present invention, the language classification process of the neural network language classifier is as follows:

[0044] h' t =W i X” t

[0045] h t =f(h′ t )

[0046]

[0047] Among them, X” t is the fixed-dimensional language feature generated by the long-term information statistical method or the recurrent neural network method, and is used as the input feature of the language classifier at the t-th time step. W i is the input layer mapping matrix, W o is the output layer mapping matrix, p(l t |X” t ) is the probability of being judged as the corresponding language. p(l t |X” t ) is normalized by spft max to ensure that the sum of the two language probabilities is 1; among them, represents the probability of being judged as the corresponding language ; represents the probability of being judged as the corresponding language ; f' t represents the intermediate variable after the input layer transformation; h tis the output vector of the last hidden layer, which is the highest-level feature of the language information and can be transmitted to the decoder as a language information vector to assist in the recognition of bilingual mixed speech content; f() is the hidden layer function, which varies according to the type of neural network in the hidden layer. A feedforward neural network or a recurrent neural network can be adopted.

[0048] As a further solution of the present invention, the method for dynamically detecting language information further includes an adjustment step of a window length limit parameter:

[0049] Obtain a mixed data set and corresponding labeled text and language labels;

[0050] Generate a dynamic window according to the window length limit parameter and calculate the recognition accuracy;

[0051] Determine the optimization interval of the window length limit parameter according to the recognition accuracy.

[0052] As a further solution of the present invention, the adjustment step of the window length limit parameter includes:

[0053] In real-time time domain selection, a dynamic window length limit parameter is adopted;

[0054] Dynamically adjust the window length limit parameter according to real-time confidence feedback.

[0055] Compared with the prior art, a method for dynamically detecting language information based on a deep neural network proposed by the present invention has the following beneficial effects:

[0056] 1. Achieve accurate language switch detection: By introducing an attention mechanism and dynamic window selection, the present invention can accurately locate the language switch point in the bilingual mixed speech stream, jointly use multiple frames of speech as the language classification input at the current time step, output the language discrimination at the current time step, and can better predict the language of the current speech stream according to historical speech information. The present invention can dynamically adjust the feature window, focus on the key time steps of language switching, and effectively improve the detection accuracy of language switching.

[0057] 2. In order to make the window have time concentration, a limit of a finite time span is added, that is, when calculating the position information of the current window w i the window length limit parameter L w is added to ensure that the window shape does not show a scattered state; through training, the window length limit parameter L w is calibrated.

[0058] 3. The output of the adopted neural network language classifier correspondingly has two output endpoints, and the output results respectively represent the probability values that the speech features at the current time step are judged to be the corresponding languages. The present invention is not only applicable to bilingual mixed speech streams, but also can be extended and applied in multilingual speech streams. By adjusting the output layer of the classifier, this method can support the recognition of more languages, adapt to the switching situations between different languages, and has high flexibility and scalability. It can be widely applied in fields such as speech recognition, speech translation, and voice assistants.

[0059] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the related art, the following will briefly introduce the accompanying drawings required for the description of the exemplary embodiments or the related art. The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings:

[0061] Figure 1 is a flowchart of a method for dynamically detecting language information based on a deep neural network according to an embodiment of the present invention.

[0062] Figure 2 is a schematic diagram of the working principle of long-term information statistics in a method for dynamically detecting language information based on a deep neural network according to an embodiment of the present invention.

[0063] Figure 3 is a schematic diagram of the working principle of generating fixed-dimensional features based on a recurrent neural network in a method for dynamically detecting language information based on a deep neural network according to an embodiment of the present invention.

[0064] Figure 4 is a schematic diagram of a neural network language classifier in a method for dynamically detecting language information based on a deep neural network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] Next, in combination with the accompanying drawings and the specific embodiments, the present application will be further described. It should be noted that, on the premise of no conflict, the following described embodiments or technical features can be arbitrarily combined to form new embodiments.

[0066] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the following further describes the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0067] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are used to distinguish two non-identical entities or non-identical parameters with the same name. It can be seen that "first" and "second" are only for the convenience of expression and should not be construed as a limitation on the embodiments of the present invention. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units inherently includes other steps or units.

[0068] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0069] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0070] Next, some embodiments of the present application will be described in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0071] Aiming at the problem that traditional language classification methods cannot effectively handle frequent and short language switches. The present invention proposes a dynamic detection method for language information based on a deep neural network, which is particularly suitable for language switch recognition in bilingual mixed speech streams. By combining an attention mechanism, multi-frame speech information, dynamic window selection, long-term information statistics, and a recurrent neural network, the detection accuracy of language switch points and the accuracy of language recognition in bilingual mixed speech streams are improved, and the problem that traditional language classification methods cannot effectively handle frequent and short language switches is solved.

[0072] See Figure 1 As shown, the embodiments of the present invention provide a dynamic detection method for language information based on a deep neural network. The method includes the following steps:

[0073] Step S10: Obtain the mixed-language speech stream to be detected and extract the original acoustic features.

[0074] Step S20: Generate an attention vector for the current time step based on the attention mechanism, and dynamically select a sequence window of speech feature frames for the current time step by combining unidirectional time selection and the limitation of a specific time span.

[0075] In this step, the attention vector for the current time step generated based on the attention mechanism is:

[0076] a i =[a i,1 ,...,a i,j ,...,a i,T

[0077] According to the attention vector a i generate a sequence window w i of speech feature frames for the current time step, where:

[0078] w i =f(a i )

[0079] Introduce the unidirectional time selection limitation, such that the time point position of the sequence window of speech feature frames conforms to the monotonically increasing property. When calculating the position information of the current window w i , add the position information of the window w i-1 of the previous time step, so that w i appears later than w i-1 :

[0080] w i =f(a i ,w i-1 )

[0081] And add the limitation of a finite time span. When calculating the position information of the current window w i , add the window length limitation parameter L w , so that the window shape does not present a scattered state:

[0082] w i =f(a i ,w i-1 ,L w ).

[0083] Step S30: Multiply the dynamic window signal by the original acoustic feature signal to generate local acoustic features for detecting the language information at the current time step.

[0084] In this step, multiply the dynamic window signal by the original acoustic feature signal to generate local acoustic features for detecting the language information at the current time step, where the local acoustic features are:

[0085]

[0086] ​Wherein, X is the original acoustic feature input, and X' is the sequence of acoustic features selected in real time, serving as the feature input for language identification at the current time step.

[0087] Step S40: Regularize local acoustic features of any length into fixed-dimensional features, and use the fixed-dimensional features as the input for the neural network language classification.

[0088] In this step, the local acoustic features of any length are regularized into fixed-dimensional features. The regularization methods include long-term information statistics and recurrent neural network; among them, referring to Figure 2 As shown, the regularization method of long-term information statistics includes:

[0089] Calculate the mean and variance statistics of the frame sequence selected in real time;

[0090] Combine the mean, variance statistics with the frame sequence selected in real time as the fixed-dimensional features.

[0091] Among them, the calculation formula of the language feature vector formed by long-term information statistics is as follows:

[0092] v t = f(Append(μ t ,σ t ))

[0093]

[0094]

[0095] Among them, X' i represents the i-th frame acoustic feature in the frame sequence selected in real time. The length of the frame sequence selected in real time is T, and μ t and σ t represent the mean and variance statistics respectively. Finally, X” t = [μ t ,σ t are combined as the input feature of the language classifier at the t-th time step; v t represents the language feature vector formed by long-term information statistics; T represents the length of the frame sequence selected in real time; t + T represents the cut-off time step position of T frames counted backward starting from the t-th time step.

[0096] In this embodiment, referring to Figure 3 As shown, the regularization method of the recurrent neural network includes:

[0097] Generate fixed-dimensional features based on the recurrent neural network;

[0098] Use the output of the recurrent neural network as the fixed-dimensional features;

[0099] Among them, the calculation formula of the fixed-dimensional feature is as follows

[0100] s t = f(U X′ t + W s t-1 )

[0101] X″ t = softmax(V s t )

[0102] Among them, U, W, and V are the input mapping matrix, the hidden layer mapping matrix, and the output mapping matrix respectively. Finally, X″ t is used as the input feature of the language classifier at the t-th time step; s t represents the hidden layer state of the recurrent neural network at the t-th time step; X' t represents the input feature at the t-th time step; s t-1 represents the hidden layer state of the recurrent neural network at the (t - 1)-th time step, which is used to participate in the calculation of the hidden layer state at the current time step; V s t represents the result of the linear transformation of the hidden layer state s t after passing through the output mapping matrix V, and it is the intermediate representation before output.

[0103] Step S50: Input the fixed-dimensional feature into the deep neural network classifier, and output the language probability value corresponding to the speech feature at the current time step.

[0104] In this step, when inputting the fixed-dimensional feature into the deep neural network classifier and outputting the language probability value corresponding to the speech feature at the current time step, the probability value that the speech feature at the current time step is judged to be the corresponding language is determined according to the fixed-dimensional language feature; among them, a deep neural network is used as the language classifier, and the output of the neural network language classifier has two output endpoints, and the output results respectively represent the probability value that the speech feature at the current time step is judged to be the corresponding language.

[0105] Among them, as shown in Figure 4 the deep neural network classifier is a multi-level structure, including:

[0106] The input layer is used to receive the fixed-dimensional feature;

[0107] The hidden layer uses a feedforward neural network or a bidirectional LSTM to extract high-level language features;

[0108] The output layer outputs the probability value of the language corresponding to the speech feature at the current time step through Softmax.

[0109] In this embodiment, the language classification process of the neural network language classifier is as follows:

[0110] h′t = W i X″ t

[0111] h t = f(h′ t )

[0112]

[0113] where X″ t is a fixed - dimensional language feature generated by a long - term information statistical method or a recurrent neural network method and serves as the input feature of the language classifier at the t - th time step. W i is the input - layer mapping matrix, and W o is the output - layer mapping matrix. p(l t |X″ t ) is the probability of being judged as the corresponding language. p(l t |X″ t ) is normalized by softmax to ensure that the sum of the probabilities of the two languages is 1. Among them, represents the probability of being judged as the corresponding language ; represents the probability of being judged as the corresponding language ; h' t represents the intermediate variable after the input - layer transformation; h t is the output vector of the last hidden layer, which is the highest - level feature of the language information and can be transmitted to the decoder as a language - information vector to assist in the recognition of bilingual mixed - speech content. f() is the hidden - layer function, which varies according to the type of neural network in the hidden layer, and can be a feed - forward neural network or a recurrent neural network type.

[0114] Step S60: Based on the dynamic - window information and the language - probability values, output the start and end times of each speech segment in the mixed - speech stream and the corresponding language labels, and determine the time points of language switching.

[0115] In some embodiments, the language - information dynamic detection method further includes a step of adjusting the window - length limit parameter:

[0116] Obtain the mixed dataset and the corresponding labeled text and language labels;

[0117] Generate a dynamic window according to the window - length limit parameter and calculate the recognition accuracy;

[0118] Determine the optimization interval of the window - length limit parameter according to the recognition accuracy.

[0119] In this embodiment, the step of adjusting the window - length limit parameter includes:

[0120] In real-time time-domain selection, a dynamic window length limit parameter is adopted;

[0121] The window length limit parameter is dynamically adjusted according to real-time confidence feedback.

[0122] Specifically, the workflow of dynamic language information detection includes the following steps:

[0123] Select a dynamic window, multiply the window signal by the original feature signal to obtain the feature input for language recognition at the current time step;

[0124] Regularize the feature input of any length into a fixed-dimensional language feature;

[0125] Use a deep neural network as a language classifier to perform speech recognition on the fixed-dimensional language feature to obtain the probability of the recognized language.

[0126] Determine an appropriate window length limit parameter L w Interval:

[0127] Obtain a mixed dataset and the annotation text and language labels of the speech data in the mixed dataset, where the mixed dataset includes first-language speech data and second-language speech data;

[0128] Obtain the original feature signal of the mixed dataset;

[0129] According to the window length limit parameter L w , select a dynamic window, multiply the window signal by the original feature signal to obtain the feature input for language recognition at the current time step;

[0130] Regularize the feature input of any length into a fixed-dimensional language feature;

[0131] Use a deep neural network as a language classifier to perform speech recognition on the fixed-dimensional language feature to obtain the probability of the recognized language, obtaining the recognition result corresponding to the window length limit parameter L w ;

[0132] Compare the recognition result with the mixed dataset to obtain the recognition accuracy corresponding to the window length limit parameter L w ;

[0133] According to the recognition accuracy corresponding to the window length limit parameter L w , determine the window length limit parameter L w interval.

[0134] Furthermore, for the adjustment of the window length limit parameter L w , the window length limit parameter L wIt can ensure that the window has time concentration, but it may also reduce the accuracy of language identification. Therefore, it is necessary to determine an appropriate window length limit parameter L w interval.

[0135] Obtain a mixed dataset and the annotation text and language labels of the speech data in the mixed dataset, where the mixed dataset includes first-language speech data and second-language speech data;

[0136] Obtain the original feature signal of the mixed dataset;

[0137] According to the window length limit parameter L w , select a dynamic window, multiply the window signal by the original feature signal to obtain the feature input for language identification at the current time step;

[0138] Regularize the feature input of any length into a fixed-dimensional language feature;

[0139] Use a deep neural network as a language classifier to perform speech recognition on the fixed-dimensional language feature to obtain the probability of the recognized language, and obtain the recognition result corresponding to the window length limit parameter L w ;

[0140] Compare the recognition result with the mixed dataset to obtain the recognition accuracy corresponding to the window length limit parameter L w ;

[0141] According to the recognition accuracy corresponding to the window length limit parameter L w , determine the window length limit parameter L w interval.

[0142] In real-time time domain selection, within the interval of the window length limit parameter L w , adopt a dynamic length limit parameter L w .

[0143] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.

[0144] It should be understood that although the above is described in a certain order, these steps are not necessarily executed in the above order successively. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, a part of the steps of this embodiment may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0145] The above are exemplary embodiments disclosed by the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed by the present invention as defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein do not need to be executed in any specific order. In addition, although the elements disclosed by the embodiments of the present invention can be described or claimed in individual form, they can also be understood as multiple unless clearly limited to the singular.

[0146] It should be understood that, as used herein, unless the context clearly supports an exception, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations including one or more of the related listed items. The above serial numbers of the disclosed embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0147] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features between the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A method for dynamic detection of language information based on deep neural network, characterized in that: The method comprises the following steps: Obtaining a mixed language speech stream to be detected and extracting original acoustic features; Generate the attention vector of the current time step based on the attention mechanism, and dynamically select the speech feature frame sequence window of the current time step by combining the one-way time selection and the restriction of a specific time span; Multiply the dynamic window signal with the original acoustic feature signal to generate the local acoustic feature of the language information detected at the current time step; Regularize local acoustic features of arbitrary length into fixed-dimensional features, so that the fixed-dimensional features can be used as input for neural network language classification; Input fixed-dimensional features to the deep neural network classifier and output the language probability value corresponding to the speech features at the current time step; Based on the dynamic window information and language probability value, the start and end time of each voice segment in the mixed voice stream and the corresponding language label are output, and the time point of language switching is determined.

2. The method for dynamic detection of language information based on deep neural network according to claim 1, characterized in that: Generate the attention vector of the current time step based on the attention mechanism; also include: Generate and select the speech feature frame sequence window of the current time step according to the attention vector; A one-way time selection restriction is introduced, and the time point position of the speech feature frame sequence window conforms to the monotonically increasing property. When calculating the current window position information, the window position information of the previous time step is added, so that the current window position information appears later than the window position information of the previous time step. A finite time span restriction is added, and a window length restriction parameter is added when calculating the current window position information so that the window shape does not appear fragmented.

3. The method for dynamic detection of language information based on deep neural network according to claim 2, characterized in that: The dynamic window signal is multiplied by the original acoustic feature signal to generate the local acoustic feature of the language information detected at the current time step, wherein the local acoustic feature is: Where X is the original acoustic feature input, X' is the acoustic feature sequence selected in real time as the feature input for language recognition at the current time step, and w i is the speech feature frame sequence window at the current time step.

4. The method for dynamic detection of language information based on a deep neural network according to any one of claims 1 to 3, characterized in that: The local acoustic features of any length are regularized into fixed-dimensional features. The regularization methods include long-term information statistics and recurrent neural networks. Among them, the regularization methods of long-term information statistics include: Calculate the mean and variance statistics of the frame sequence selected in real time; The mean and variance statistics are combined with the frame sequence selected in real time as fixed-dimensional features.

5. The method for dynamic detection of language information based on deep neural network according to claim 4, characterized in that: The calculation formula of the language feature vector constructed by long-term information statistics is as follows: v i =f(Append(μ t ,s t )) Among them, X' i represents the acoustic features of the i-th frame in the frame sequence selected in real time. The length of the frame sequence selected in real time is T, μ t and σ t Represent the mean and variance statistics respectively; finally, X" t =[μ t ,σ t ] are combined as the input features of the language classifier at the tth time step; v t Represents the language feature vector constructed by long-term information statistics; T represents the length of the frame sequence selected in real time; t+T represents the cutoff time step position of T frames after the tth time step.

6. The method for dynamic detection of language information based on deep neural network according to claim 4, characterized in that: Regularization methods for recurrent neural networks include: Generate fixed-dimensional features based on recurrent neural networks; Use the output of the recurrent neural network as fixed-dimensional features; Among them, the calculation formula for fixed dimension features is as follows s t =f(UX′ t +Ws t-1 ) X″ t =soft max(Vs t ) Among them, U, W and V are the input mapping matrix, hidden layer mapping matrix and output mapping matrix respectively; finally, X" t As the input feature of the language classifier at the tth time step; s t represents the hidden layer state of the recurrent neural network at the tth time step; X' t represents the input feature of the tth time step; s t-1 Represents the hidden layer state of the recurrent neural network at the t-1th time step, and is used to participate in the calculation of the hidden layer state at the current time step; Vs t Represents the hidden layer state s t The result after linear transformation of the output mapping matrix V is the intermediate representation before output.

7. The method for dynamic detection of language information based on deep neural network according to claim 1, characterized in that: When inputting fixed-dimensional features into a deep neural network classifier and outputting the language probability value corresponding to the speech features at the current time step, the probability value of the speech features at the current time step being judged as the corresponding language is determined based on the fixed-dimensional language features; wherein a deep neural network is used as a language classifier, and the output of the neural network language classifier used has two output endpoints, and the output results respectively represent the probability values ​​of the speech features at the current time step being judged as the corresponding language.

8. The method for dynamic detection of language information based on deep neural network according to claim 7, characterized in that: The deep neural network classifier is a multi-layer structure, including: Input layer, used to receive fixed-dimensional features; Hidden layer, using feedforward neural network or bidirectional LSTM to extract high-level language features; The output layer outputs the probability value of the language corresponding to the speech feature at the current time step through Softmax.

9. The method for dynamic language information detection based on deep neural network according to claim 1, characterized in that: The method for dynamic detection of language information also includes the step of adjusting the window length limit parameter: Get the mixed dataset and the corresponding annotated text and language labels; Generate a dynamic window and calculate the recognition accuracy based on the window length limit parameter; The optimization interval of the window length limit parameter is determined according to the recognition accuracy.

10. The method for dynamic detection of language information based on deep neural network according to claim 9, characterized in that: The step of adjusting the window length limit parameter comprises: In real-time time domain selection, dynamic window length limit parameters are used; The window length limit parameter is dynamically adjusted according to real-time confidence feedback.

Citation Information

Patent Citations

  • A slot filling method for extracting semantic features based on a dynamic window self-attention mechanism

    CN109918503A

  • Streaming phonetic transcription system based on self-attention mechanism

    CN110473529A

  • Multi-language continuous voice stream voice content identification method and system

    CN112489622A

  • Speech function automatic evaluation system and method based on speech recognition

    CN113496696A

  • Multilingual speech synthesis and cross-language voice cloning

    US20200380952A1