Language Identification Method, Device, Electronic Device and Computer Readable Storage Medium
Through the dual language recognition network architecture, combined with the first language and the second language recognition network, the problem of instability in language recognition in the prior art is solved, and high-accurate language recognition in different languages is achieved.
Patent Information
- Application Number
- CN202210701565.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-17
AI Technical Summary
The existing language recognition technology has unstable recognition performance in different languages, especially in multilingual situations, and is difficult to accurately recognize.
The dual language recognition network architecture is adopted, and the first language recognition network is used to initially recognize monolingual audio, and whether the detection results meet the preset requirements; if not, the second language recognition network is used for re-recognition, and the second language recognition network has stronger recognition ability for multilingual audio.
It improves the accuracy of language recognition and can achieve accurate language recognition in different languages, especially in the case of mixed monolingual and multilingual, which can improve the accuracy of recognition.
Smart Images

Figure CN115223543B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of language identification, and in particular to a language identification method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Language identification, also known as language recognition, refers to the process by which a machine automatically determines the language category to which a speech segment belongs. Currently, the mainstream language identification method is the TV (Total Varbility) system, etc. This method uses training corpora to train a full-variability space covering various environments and channels, maps the speech to be measured into a language model vector with a fixed and unified dimension, and then compares the similarity with multiple pre-set language model vectors to be determined, so as to determine the language category of the speech to be measured.
[0003] However, the mainstream language identification technology mainly targets audio data in specific language situations (such as a single language). When facing audio data in other language situations, the recognition performance may jitter greatly or even fail to operate normally. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a language identification method, apparatus, electronic device, and computer-readable storage medium that can accurately identify the language of audio in different language situations.
[0005] To solve the above technical problem, a technical solution adopted by this application is: to provide a language identification method, the method includes: using a first language identification network to perform language identification on the audio to be identified to obtain an initial language identification result; detecting whether the initial language identification result meets a preset identification requirement; in response to the initial language identification result not meeting the preset identification requirement, using a second language identification network to perform language identification on the audio to be identified to obtain a target language identification result; wherein, the first language identification network has a stronger ability to identify audio in a first language situation than the second language identification network, and the second language identification network has a stronger ability to identify audio in a second language situation than the first language identification network.
[0006] Wherein, the first language situation is a single language, the second language situation is a multi-language, the first language identification network can identify a single language, and the second language identification network can identify at least one language.
[0007] Wherein, the initial language identification result includes an initial language existing in the audio to be identified and a confidence score corresponding to the initial language; the preset identification requirement includes that the confidence score meets a preset score requirement.
[0008] Among them, the second language recognition network is a hidden Markov model, which is composed of several spliced Gaussian mixture models, and each Gaussian mixture model is used to recognize a language.
[0009] Among them, the language recognition method further includes the following training steps for the first language recognition network: obtaining a first sample audio and a second sample audio; among them, the first sample audio is labeled with the true language information existing in the first sample audio, and the second sample audio is not labeled; performing random masking processing on the second sample audio to obtain a third sample audio; using the first language recognition network to perform language recognition on the first sample audio, the second sample audio, and the third sample audio, and correspondingly obtaining a first sample language recognition result, a second sample language recognition result, and a third sample language recognition result; adjusting the network parameters of the first language recognition network based at least on a first difference between the first sample language recognition result and the true language information, and a second difference between the second sample language recognition result and the third sample language recognition result.
[0010] Among them, there are several first sample audios. Adjusting the network parameters of the first language recognition network based at least on a first difference between the first sample language recognition result and the true language information, and a second difference between the second sample language recognition result and the third sample language recognition result includes: obtaining a first average distance between the feature representation of the current first sample audio and the feature representations of each positive sample audio, and a second average distance between the feature representation of the current first sample audio and the feature representations of each negative sample audio; among them, a positive sample audio is a first sample audio with the same language as the current first sample audio, a negative sample audio is a first sample audio with a different language from the current first sample audio, and the feature representation is extracted during the process of the first language recognition network performing language recognition on the corresponding first sample audio; adjusting the network parameters of the first language recognition network based on the first difference, the second difference, and the difference between the first average distance and the second average distance.
[0011] Among them, before using the first language recognition network to perform language recognition on the audio to be recognized and obtaining an initial language recognition result, the language recognition method further includes: extracting features from the audio to be recognized to obtain the target acoustic features of the audio to be recognized; or, extracting features from the initial audio to obtain the target acoustic features of the initial audio, and extracting the target acoustic features with a preset time length from the target acoustic features of the initial audio as the target acoustic features of the audio to be recognized; using the first language recognition network to perform language recognition on the audio to be recognized and obtaining an initial language recognition result includes: using the first language recognition network to perform language recognition on the target acoustic features to obtain an initial language recognition result; using the second language recognition network to perform language recognition on the audio to be recognized and obtaining a target language recognition result includes: using the second language recognition network to perform language recognition on the target acoustic features to obtain a target language recognition result.
[0012] Among them, performing feature extraction on the initial audio to obtain the target acoustic features of the initial audio, or performing feature extraction on the audio to be recognized to obtain the target acoustic features of the audio to be recognized, includes: obtaining the shifted differential cepstrum features of the data to be extracted; using a feature extraction network to perform feature extraction on the shifted differential cepstrum features of the data to be extracted to obtain bottleneck features, and taking the bottleneck features as the target acoustic features of the data to be extracted; where the data to be extracted is the initial audio or the audio to be recognized.
[0013] Among them, the feature extraction network includes a deep neural network layer and a bottleneck network layer connected in sequence, and the bottleneck features are output by the bottleneck network layer; and / or, the feature extraction network includes an output layer in the training stage, and the output layer is used to predict the bottleneck features extracted by the feature extraction network to obtain the corresponding language recognition result; the language recognition method further includes: obtaining a fourth sample audio and a fifth sample audio; where the fourth sample audio is labeled with the true language information existing in the fourth sample audio, and the fifth sample audio is not labeled; performing random masking processing on the fifth sample audio to obtain a sixth sample audio; using the feature extraction network in the training stage to process the fourth sample audio, the fifth sample audio, and the sixth sample audio respectively to obtain the fourth sample language recognition result, the fifth sample language recognition result, and the sixth sample language recognition result; adjusting the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information, and the difference between the fifth sample language recognition result and the sixth sample language recognition result; where after the feature extraction network is trained, the output layer is removed.
[0014] Among them, the target language recognition result includes at least one target language existing in the audio to be recognized and the time interval corresponding to each target language in the audio to be recognized; and / or, the audio to be recognized is a speech segment with a preset time length in the initial audio, and the initial audio includes several audios to be recognized corresponding to different time periods; after using the second language recognition network to perform language recognition on the audio to be recognized to obtain the target language recognition result, the language recognition method further includes: combining the target language recognition results of each audio to be recognized included in the initial audio to obtain the language recognition result of the initial audio.
[0015] To solve the above technical problems, another technical solution adopted in this application is: to provide a language recognition device, which includes: a first recognition module, a detection module, and a second recognition module; the first recognition module is used to perform language recognition on the audio to be recognized by using a first language recognition network to obtain an initial language recognition result; the detection module is used to detect whether the initial language recognition result meets a preset recognition requirement; the second recognition module is used to, in response to the initial language recognition result not meeting the preset recognition requirement, perform language recognition on the audio to be recognized by using a second language recognition network to obtain a target language recognition result; wherein, the first language recognition network has a stronger recognition ability for audio in the first language situation than the second language recognition network, and the second language recognition network has a stronger recognition ability for audio in the second language situation than the first language recognition network.
[0016] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above language recognition method.
[0017] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer-readable storage medium, which is used to store program instructions, and the program instructions can be executed to implement the above language recognition method.
[0018] In the above embodiment, when the initial language recognition result obtained by performing language recognition on the audio to be recognized by using the first language recognition network does not meet the preset recognition requirement, the second language recognition network is used to perform language recognition on the audio to be recognized to obtain a target language recognition result. Therefore, when the initial language recognition result does not meet the preset recognition requirement, using the second language recognition network to perform language recognition on the audio to be recognized again can improve the accuracy of language recognition; in addition, since the first language recognition network has a stronger recognition ability for audio in the first language situation than the second language recognition network, and the second language recognition network has a stronger recognition ability for audio in the second language situation than the first language recognition network, by combining the first language recognition network and the second language recognition network, accurate language recognition can be performed on audio in different language situations. Description of the Drawings
[0019] Figure 1 is a schematic flowchart of an embodiment of the language recognition method provided by this application;
[0020] Figure 2 is a schematic structural diagram of an embodiment of the first language recognition network provided by this application;
[0021] Figure 3 is a schematic structural diagram of an embodiment of the feature extraction network provided by this application;
[0022] Figure 4 It is a schematic structural diagram of an embodiment of the GMM model provided by this application;
[0023] Figure 5 It is a schematic flowchart of an embodiment of training the first language recognition network provided by this application;
[0024] Figure 6 is Figure 5 It is a schematic flowchart of an embodiment of step S54 shown;
[0025] Figure 7 It is a schematic flowchart of an embodiment of obtaining target acoustic features provided by this application;
[0026] Figure 8 It is a schematic flowchart of an embodiment of training a feature extraction network provided by this application;
[0027] Figure 9 It is a schematic structural diagram of an embodiment of the language recognition device provided by this application;
[0028] Figure 10 It is a schematic structural diagram of an embodiment of the electronic device provided by this application;
[0029] Figure 11 It is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Specific Embodiments
[0030] The following will combine the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of this application.
[0031] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures, interfaces, and technologies are presented to thoroughly understand this application.
[0032] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0033] Please refer to Figure 1 , Figure 1It is a schematic flowchart of an embodiment of the language recognition method provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. For example, Figure 1 as shown, this embodiment includes:
[0034] Step S11: Use the first language recognition network to perform language recognition on the audio to be recognized, and obtain an initial language recognition result.
[0035] The method of this embodiment is used to perform language recognition on the audio to be recognized to obtain a language recognition result corresponding to the audio to be recognized. In one embodiment, the audio to be recognized is any audio that needs to be subjected to language recognition, and can specifically be obtained from local storage or cloud storage. For example, the audio to be recognized can be an audio corresponding to a voice call, a recording playback, etc., or can also be an audio generated by means such as speech synthesis, voice conversion, or human imitation. It can be understood that in other embodiments, it can also be obtained by collecting the current environmental sound through a voice collection device, which is not specifically limited herein.
[0036] In this embodiment, the first language recognition network is used to perform language recognition on the audio to be recognized to obtain an initial language recognition result. Among them, the number of language categories involved in the audio to be recognized is not limited. For example, the audio to be recognized is audio data involving only a single language; or, the audio to be recognized is audio data involving 2, 3, 4, or multiple mixed languages. In addition, the specific language types involved in the audio to be recognized are not limited. For example, the audio to be recognized is mixed-language audio data involving multiple languages such as Chinese, English, and German; or, the audio to be recognized is audio data involving only a single language such as Chinese, English, or German.
[0037] In one embodiment, the first language recognition network can recognize a single language, that is to say, the first language recognition network is a single-language recognition network. In a specific embodiment, when the first language recognition network is a single-language recognition network, the initial language recognition result obtained by performing language recognition on the audio to be recognized includes an initial language existing in the audio to be recognized. Alternatively, in other specific embodiments, the initial language recognition result may also include other relevant information such as an initial language existing in the audio to be recognized and the confidence score corresponding to the initial language, which is not specifically limited herein. It can be understood that in other embodiments, the first language recognition network can recognize at least one language, that is to say, the first language recognition network is a mixed-language recognition network, which is not specifically limited herein. In a specific embodiment, when the first language recognition result is a mixed-language recognition network, the initial language recognition result obtained by performing language recognition on the audio to be recognized includes at least one language existing in the audio to be recognized. Alternatively, in other specific embodiments, the initial language recognition result may also include other relevant information such as at least one language existing in the audio to be recognized, the confidence scores corresponding to at least one language, and the time intervals corresponding to at least one language in the audio to be recognized, which is not specifically limited herein.
[0038] Among them, the specific model structure of the first language recognition network is not limited and can be specifically set according to actual usage needs. Exemplarily, as Figure 2 shown, Figure 2 is a schematic structural diagram of an embodiment of the first language recognition network provided by the present application. The first language recognition network includes a convolutional neural network (CNN) layer, a bidirectional long short-term memory (BILSTM) network layer, an attention (ATTENTION) layer, and an output (OUTPUT) layer connected in sequence. Among them, the CNN layer has a strong feature transformation ability; the BILSTM network layer has the ability to model the temporal sequence of related features; the ATTENTION layer is used to enhance the feature information effective for language classification; the OUTPUT layer is actually a linear fully connected layer used to classify language categories, and the number of its nodes is the number of language categories.
[0039] In one embodiment, the audio to be recognized is a voice segment of a preset time length in the initial audio, and the initial audio includes a plurality of audios to be recognized corresponding to different time periods. That is to say, the obtained initial audio will be divided into a plurality of voice segments corresponding to different time periods, and each voice segment will be used as the audio to be recognized for language identification respectively, so as to perform language identification on the initial audio in a segmented manner, which can make the language identification result more accurate and improve the accuracy of language identification. Among them, the preset time length is not limited and can be specifically set according to actual usage needs. For example, the preset time lengths of the audios to be recognized corresponding to different time periods included in the initial audio are the same. Of course, the preset time lengths of the audios to be recognized corresponding to different time periods included in the initial audio can also be all different, or the preset time lengths of some audios to be recognized are the same, and the preset time lengths of some audios to be recognized are different. It can be understood that in other embodiments, the initial audio may not be divided into voice segments, and the entire initial audio may be directly used as the audio to be recognized.
[0040] In a specific embodiment, the initial audio can be processed by a sliding window to obtain a plurality of audios to be recognized corresponding to different time periods. Specifically, the initial audio is processed by a sliding window to obtain a plurality of sliding window segments of a preset time length, and each sliding window segment corresponds to an audio to be recognized in a time period. Among them, it should be noted that if the preset time length of the audio to be recognized in a certain time period corresponding to the sliding window segment is insufficient, it can be filled with 0. Of course, in other specific embodiments, the initial audio can also be divided into voice segments of each preset time length by other methods.
[0041] In order to improve the accuracy of the initial language identification result obtained by the first language identification network for identifying the language of the audio to be recognized, in one embodiment, before using the first language identification network to identify the language of the audio to be recognized and obtain the initial language identification result, preprocessing such as noise removal and discontinuous sound removal is performed on the audio to be recognized, so that noise or discontinuous sound in the audio to be recognized can be filtered out, which is beneficial to the accurate identification of subsequent languages.
[0042] Since the first language recognition network actually performs language recognition on the audio to be recognized by performing language recognition on the acoustic features extracted from the audio to be recognized. Therefore, in one embodiment, before using the first language recognition network to perform language recognition on the audio to be recognized and obtaining the initial language recognition result, acoustic feature extraction is performed. Among them, when the audio to be recognized is a voice segment corresponding to a certain time period in the initial audio, the specific method for extracting the acoustic features of the audio to be recognized, which is specifically a voice segment, is as follows: First, perform feature extraction on the initial audio to obtain the target acoustic features of the initial audio; then, extract the target acoustic features with a preset time length from the target acoustic features of the initial audio as the target acoustic features of the audio to be recognized. When the audio to be recognized is the complete initial audio, directly perform feature extraction on the audio to be recognized to obtain the target acoustic features of the audio to be recognized.
[0043] At this time, the first language recognition network specifically performs language recognition on the target acoustic features of the audio to be recognized to obtain the initial language recognition result corresponding to the audio to be recognized.
[0044] In one embodiment, a feature extraction network can be used to perform feature extraction on the audio to be recognized or the initial audio to obtain the target acoustic features of the audio to be recognized or the initial audio. It can be understood that in other embodiments, the target acoustic features can also be extracted from the audio to be recognized or the initial audio by using a feature extraction algorithm, etc., and no specific limitation is made here.
[0045] In one embodiment, the target acoustic feature is the shifted differential cepstrum (SDC) feature of the audio to be recognized or the initial audio. In order to improve the accuracy of the initial language recognition result obtained by the first language recognition network performing language recognition on the target acoustic features, in other embodiments, the target acoustic feature is the bottleneck (Bottleneck Network, BN) feature, and the BN feature has stronger language representativeness and anti-noise ability. Among them, the specific structure of the feature extraction network is not limited and can be specifically set according to actual usage needs. Exemplarily, as Figure 3 shown, Figure 3 is a schematic structural diagram of an embodiment of the feature extraction network provided by the present application. The feature extraction network includes a deep neural network (DNN) layer and a bottleneck network layer connected in sequence, and the BN feature is output by the bottleneck network layer. Specifically, the DNN layer is a multi-layer linear fully connected layer with a large number of nodes between layers. For example, 512 or 1024 is taken, etc.; the bottleneck network layer is a single-layer linear fully connected layer, and the number of its nodes is less than that in the DNN layer. For example, 56 is taken.
[0046] In a specific embodiment, when the target acoustic feature is the BN feature, the feature extraction network can be directly used to extract features from the audio to be recognized or the initial audio, so as to obtain the BN feature of the audio to be recognized or the initial audio. It can be understood that in other specific embodiments, when the target acoustic feature is the BN feature, the SDC feature of the audio to be recognized or the initial audio can also be obtained first, and then the feature extraction network is used to extract the SDC feature of the audio to be recognized or the initial audio to obtain the BN feature of the audio to be recognized or the initial audio.
[0047] Step S12: Detect whether the initial language recognition result meets the preset recognition requirements.
[0048] In this embodiment, it is detected whether the initial language recognition result meets the preset recognition requirements. Since the first language recognition network has a stronger language recognition ability for the audio in the first language situation than the second language recognition network, and the first language recognition network has a weaker language recognition ability for the audio in the second language situation than the second language recognition network, that is, the first language recognition network has a stronger language recognition ability for the audio to be recognized in the first language situation, and the credibility or accuracy of the initial language recognition result obtained by performing language recognition on the audio to be recognized in the first language situation is relatively high, and the initial language recognition result correspondingly meets the preset recognition requirements; therefore, by detecting whether the initial language recognition result meets the preset recognition requirements, the accuracy or credibility of the initial language recognition result can be determined, so as to determine whether it is necessary to use the second language recognition network to perform language recognition on the audio to be recognized.
[0049] Among them, when the initial language recognition result does not meet the preset recognition requirements, it indicates that the first language recognition network has a weak recognition ability for the audio to be recognized in the current language situation, resulting in a low credibility or accuracy of the initial language recognition result obtained by performing language recognition on the audio to be recognized, so step S13 is executed at this time; while when the initial language recognition result meets the preset recognition requirements, it indicates that the first language recognition network has a strong recognition ability for the audio to be recognized in the current language situation, making the obtained initial language recognition result have a high credibility or accuracy, so the initial language recognition result is used as the final language recognition result of the audio to be recognized.
[0050] In one embodiment, the first language scenario is a single language, and the second language scenario is multiple languages. The first language recognition network can recognize a single language, and the second language recognition network can recognize at least one language. That is to say, the first language recognition network has a stronger ability to recognize single-language audio than the second language recognition network, and the second language recognition network has a stronger ability to recognize multi-language audio than the first language recognition network. Therefore, when the audio to be recognized is multi-language, due to the weak ability of the first language recognition network to recognize multi-language audio, the accuracy of the initial language recognition result obtained is relatively low, that is, the initial language recognition result does not meet the preset recognition requirements. At this time, step S13 is executed to use the second language recognition network to recognize the language of the multi-language audio to be recognized. Since the second language recognition network has a strong ability to recognize multi-language audio, the accuracy of language recognition can be improved. When the audio to be recognized is a single language, since the first language recognition network has a strong ability to recognize single-language audio, the accuracy of the initial language recognition result obtained is relatively high, that is, the initial language recognition result meets the preset recognition requirements. Therefore, the initial language recognition result is used as the final language recognition result of the audio to be recognized. It can be understood that in other embodiments, the first language scenario can also be multiple languages, the second language scenario is a single language, the first language recognition network can recognize at least one language, and the second language recognition network can recognize a single language. Among them, the specific language categories that the first language recognition network can recognize and the specific language categories that the second language recognition network can recognize are not limited and can be specifically set according to actual usage needs.
[0051] In a specific embodiment, the first language scenario is a single language, the second language scenario is multiple languages, the first language recognition network can recognize a single language, and the second language recognition network can recognize at least one language. At this time, the initial language recognition result obtained by using the first language recognition network to recognize the language of the audio to be recognized includes an initial language existing in the audio to be recognized and a confidence score corresponding to the initial language. It can be understood that in other specific embodiments, the initial language recognition result obtained by using the first language recognition network to recognize the language of the audio to be recognized may also include the time interval corresponding to the initial language. Since the first language recognition network is a single-language recognition network, the time interval corresponding to the initial language is the time interval of the audio to be recognized.
[0052] Optionally, the preset recognition requirements are not specifically limited and can be specifically set according to actual usage needs. For example, when the first language scenario is a single language and the initial language recognition result includes an initial language existing in the audio to be recognized and a confidence score corresponding to the initial language, the preset recognition requirement can be that the confidence score meets the preset score requirement. Among them, the preset score requirement is not limited. For example, the preset score requirement is greater than or equal to 0.9.
[0053] Exemplarily, take the case where the first language scenario is a single language, the first language recognition network can recognize a single language, the initial language recognition result includes an initial language existing in the audio to be recognized and the confidence score corresponding to the initial language, the preset recognition requirement is that the confidence score meets the preset score requirement, and the preset score requirement is greater than or equal to 0.9 as an example. Among them, using the first language recognition network to perform language recognition on the audio to be recognized A that only includes Chinese, the obtained initial language recognition result includes an initial language existing in the audio to be recognized as the Chinese language and the confidence score corresponding to the initial language is 0.98; since the first language recognition network has a strong recognition ability for single-language audio, the obtained confidence score is relatively high and meets the preset score requirement, so the initial language recognition result meets the preset recognition requirement. Therefore, the initial language recognition result is used as the language recognition result of the audio to be recognized A, that is, the language existing in the audio to be recognized A is Chinese. When using the first language recognition network to perform language recognition on the audio to be recognized B that includes English and Chinese, since the first language recognition result can only recognize a single language, the obtained initial language recognition result includes an initial language existing in the audio to be recognized and the confidence score corresponding to the initial language is 0.7; since the first language recognition network has a weak recognition ability for multi-language audio, the obtained confidence score is relatively low and does not meet the preset score requirement, so the initial language recognition result does not meet the preset recognition requirement, that is, the accuracy of the initial language recognition result is relatively low. Therefore, step S13 is executed.
[0054] Step S13: In response to the initial language recognition result not meeting the preset recognition requirement, use the second language recognition network to perform language recognition on the audio to be recognized to obtain the target language recognition result.
[0055] In this embodiment, in response to the initial language recognition result not meeting the preset recognition requirement, use the second language recognition network to perform language recognition on the audio to be recognized to obtain the target language recognition result. When the initial language recognition result does not meet the preset recognition requirement, it indicates that the first language recognition network has a weak recognition ability for the audio to be recognized in the current language scenario. Therefore, at this time, use the second language recognition network to perform language recognition on the audio to be recognized again to obtain the target language recognition result, thereby improving the accuracy of language recognition of the audio to be recognized. By combining the first language recognition network and the second language recognition network, it is possible to perform language recognition on audio in various language scenarios (such as single language or multi-language); in addition, since the first language recognition network has a stronger recognition ability for audio in the first language scenario than the second language recognition network, and the second language recognition network has a stronger recognition ability for audio in the second language scenario than the first language recognition network, the language recognition method provided in this application can accurately perform language recognition on audio in various language scenarios.
[0056] In one embodiment, the second language recognition network specifically performs language recognition on the target acoustic features of the audio to be recognized, and obtains the target language recognition result corresponding to the audio to be recognized. Among them, the relevant content about the target acoustic features is as described above, and will not be elaborated here.
[0057] Exemplarily, taking the case where the first language is a single language, the first language recognition network can recognize a single language, the initial language recognition result includes an initial language existing in the audio to be recognized and the confidence score corresponding to the initial language, the preset recognition requirement is that the confidence score is not less than 0.9, the second language is a multi-language, and the second language recognition network can recognize at least one language as an example. Using the first language recognition network to perform language recognition on the audio C to be recognized, which includes English and Chinese. Since the first language recognition network can only recognize a single language, the obtained initial language recognition result includes an initial language existing in the audio to be recognized and the confidence score corresponding to the initial language is 0.7, and the confidence score does not meet the preset recognition requirement. Therefore, at this time, the second language recognition network is used to perform language recognition on the audio C to be recognized, and the target language recognition result is obtained. Since the second language recognition network has a strong recognition ability for multi-language audio, the accuracy of the target language recognition result corresponding to the multi-language audio C to be recognized is relatively high, improving the accuracy of language recognition of the audio to be recognized.
[0058] In one embodiment, the target language recognition result includes at least one target language existing in the audio to be recognized. It can be understood that in other embodiments, the target language recognition result includes at least one target language existing in the audio to be recognized and the time interval corresponding to each target language in the audio to be recognized, so that the position information where each target language appears alternately can be determined through the time interval corresponding to each target language in the audio to be recognized, which has practical business application value. For example, the target language recognition result includes 3 target languages existing in the audio to be recognized, namely target language A, target language B, and target language C; the time interval corresponding to target language A is the time interval formed by the time corresponding to the first audio frame to the time corresponding to the eighth audio frame in the audio to be recognized, the time interval corresponding to target language B is the time interval formed by the time corresponding to the second audio frame to the time corresponding to the fifth audio frame and the time corresponding to the eighth audio frame to the time corresponding to the fifteenth audio frame in the audio to be recognized, and the time interval corresponding to target language C is the time interval formed by the time corresponding to the sixteenth audio frame to the time corresponding to the eighteenth audio frame in the audio to be recognized.
[0059] Since when the audio to be recognized is a speech segment of a preset time length in the initial audio, it is also necessary to combine the language recognition results of several audios to be recognized corresponding to different time periods included in the initial audio in order to obtain the language recognition result of the initial audio. Therefore, in one embodiment, after using the second language recognition network to perform language recognition on the audio to be recognized and obtaining the target language recognition result, it is also necessary to combine the target language recognition results of each audio to be recognized included in the initial audio to obtain the language recognition result of the initial audio. It should be noted that if a certain audio to be recognized only includes a single language, that is, when the initial language recognition result obtained by using the first language recognition network to perform language recognition on a certain audio to be recognized meets the preset recognition requirements, the initial language recognition result is used as the target language recognition result of the audio to be recognized.
[0060] In one embodiment, the second language recognition network is a Hidden Markov Model (HMM). The HMM is composed of several Gaussian Mixed Model (GMM) models spliced together. Each GMM model is used to recognize a language. Among them, the GMM model corresponding to each language is used as a state of the HMM. The log-domain self-jump penalty factor is set to penalty1, and the state transition penalty factor is set to penalty2. The viterbi decoding is used to dynamically plan the language category and the corresponding time interval in the audio to be recognized.
[0061] In a specific embodiment, as Figure 4 shown, Figure 4 FIG. is a schematic structural diagram of an embodiment of the GMM model provided by the present application. The training process of the GMM model corresponding to each language specifically includes: setting the Gaussian mixture system of the GMM model to M, and using K-means clustering to obtain the initial GMM model; then, using the EM algorithm (Expectation-Maximization algorithm) to perform iterative training on the labeled sample audio data to train the GMM model; then, iterating repeatedly until the GMM model converges. Among them, M is an integer value. For example, M is taken as 1024. Generally speaking, the more labeled sample audio data there is, the larger the M value can be set.
[0062] In the above embodiments, when the initial language recognition result obtained by using the first language recognition network to recognize the audio to be recognized does not meet the preset recognition requirements, the second language recognition network is used to recognize the audio to be recognized to obtain the target language recognition result. Therefore, when the initial language recognition result does not meet the preset recognition requirements, using the second language recognition network to recognize the audio to be recognized again can improve the accuracy of language recognition; in addition, since the first language recognition network has a stronger recognition ability for audio in the first language situation than the second language recognition network, and the second language recognition network has a stronger recognition ability for audio in the second language situation than the first language recognition network, by combining the first language recognition network and the second language recognition network, accurate language recognition can be performed on audio in different language situations.
[0063] Please refer to Figure 2 and Figure 5 , Figure 5 which is a schematic flowchart of an embodiment for training the first language recognition network provided by the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 5 the process sequence shown. As Figure 5 shown, the first language recognition network can recognize a single language. The first language recognition network includes a CNN layer, a BILSTM network layer, an ATTENTION layer, and an OUTPUT layer connected in sequence. The training steps of the first language recognition network specifically include:
[0064] Step S51: Obtain the first sample audio and the second sample audio.
[0065] The method of this embodiment is used to train the first language recognition network based on the first sample audio and the second sample audio, so that the trained first language recognition network can more accurately recognize the language existing in the audio data. Therefore, in this embodiment, the first sample audio and the second sample audio are obtained, where the first sample audio is marked with the true language information existing in the first sample audio, and the second sample audio is not marked with the true language information existing in the second sample audio. That is to say, in the training of the first language recognition network, a large amount of unsupervised sample audio data is added, so that the trained first language recognition network can more accurately recognize the language of the audio data and has anti-interference ability, that is, the trained first language recognition network can also have a good language recognition effect in the presence of interference, improving the robustness of the first language recognition network.
[0066] In one embodiment, the first sample audio and the second sample audio can be specifically obtained from local storage or cloud storage. It can be understood that in other embodiments, they can also be obtained by collecting the current environmental sound through a voice collection device, and no specific limitation is made here.
[0067] In one embodiment, there are several first sample audios and second sample audios respectively. Subsequently, several first sample audios and several second sample audios can be input into the first language recognition network simultaneously. That is, in a subsequent training of the first language recognition network, the first language recognition network is trained using a batch of sample audios, which improves the training efficiency of the first language recognition network. Among them, the number of first sample audios and second sample audios input into the first language recognition network in batches is not limited and can be specifically set according to actual usage needs. It can be understood that in other embodiments, the first sample audio and the second sample audio are each one, that is, in a subsequent training of the first language recognition network, the first language recognition network is trained using a single sample audio.
[0068] Among them, the specific languages involved in the first sample audio and the second sample audio are not limited. For example, the languages involved in the first sample audio and the second sample audio are Chinese, English, German, etc. It should be noted that when the first language recognition network is subsequently trained using a batch of first sample audios and second sample audios, the specific languages involved in each first sample audio can be the same or different, and the specific languages involved in each second sample audio can be the same or different.
[0069] In one embodiment, the first sample audio and the second sample audio can be preprocessed to filter out invalid sounds such as noise and intermittent sounds, so as to improve the quality of the first sample audio and the second sample audio, thereby making the language recognition effect of the first language recognition network trained based on the first sample audio and the second sample audio better.
[0070] In one embodiment, the first sample audio and the second sample audio can also be augmented, such as speed change processing, so that the number of samples used to train the first language recognition network is more sufficient and more diverse, thereby enabling the first language recognition network completed in subsequent training to have better generalization ability.
[0071] Step S52: Perform random masking processing on the second sample audio to obtain a third sample audio.
[0072] In this embodiment, the second sample audio is randomly masked to obtain the third sample audio. Specifically, a random tfmask is applied to the second sample audio to obtain the third sample audio. That is to say, the difference between the third sample audio and the second sample audio lies in whether the tfmask is applied, which is equivalent to creating an interference on the second sample audio to obtain the third sample audio. After adjusting the network parameters of the first language recognition network based on the differences between the second sample audio and the third sample audio and training until convergence, the trained and converged first language recognition network can have a good language recognition effect in the presence of interference.
[0073] In one embodiment, when the first language recognition network is trained once using a single second sample audio subsequently, random audio frames of this second sample audio can be randomly masked at this time to obtain the third sample audio.
[0074] In one embodiment, when the first language recognition network is trained once using a batch of second sample audios subsequently, random audio frames of any second sample audio in the batch of second sample audios can be randomly masked at this time to obtain a batch of third sample audios.
[0075] Step S53: Use the first language recognition network to perform language recognition on the first sample audio, the second sample audio, and the third sample audio, and correspondingly obtain the first sample language recognition result, the second sample language recognition result, and the third sample language recognition result.
[0076] In this embodiment, the first language recognition network is used to perform language recognition on the first sample audio, the second sample audio, and the third sample audio, and correspondingly obtain the first sample language recognition result, the second sample language recognition result, and the third sample language recognition result. That is to say, when the first language recognition network is used to perform language recognition on the first sample audio, the second sample audio, and the third sample audio, the language prediction result corresponding to the first sample audio, i.e., the first sample language recognition result, the language prediction result corresponding to the second sample audio, i.e., the second sample language recognition result, and the language prediction result corresponding to the third sample audio, i.e., the third sample language recognition result, will be obtained.
[0077] In one embodiment, the first sample audio, the second sample audio, and the third sample audio are single audio data. That is, when the first language recognition network is trained once using a single first sample audio, a single second sample audio, and a single third sample audio, the first sample language recognition result obtained at this time only includes the language recognition result corresponding to this single first sample audio, the second sample language recognition result only includes the language recognition result corresponding to this single second sample audio, and the third sample language recognition result only includes the language recognition result corresponding to this single third sample audio.
[0078] In one embodiment, when the first sample audio, the second sample audio, and the third sample audio are several, that is, when performing one training of the first language recognition network using a batch of the first sample audio, the second sample audio, and the third sample audio, the obtained first sample language recognition network includes the language recognition results corresponding to each first sample audio, the second sample language recognition results include the language recognition results corresponding to each second sample audio, and the third sample language recognition results include the language recognition results corresponding to each third sample audio.
[0079] Step S54: Adjust the network parameters of the first language recognition network based at least on the first difference between the first sample language recognition results and the true language information, and the second difference between the second sample language recognition results and the third sample language recognition results.
[0080] In this embodiment, the network parameters of the first language recognition network are adjusted based at least on the first difference between the first sample language recognition results and the true language information existing in the first sample audio marked on the first sample audio, and the second difference between the second sample language recognition results and the third sample language recognition results.
[0081] In one embodiment, the network parameters of the first language recognition network can be adjusted based on the first difference between the first sample language recognition result and the true language information existing in the first sample audio annotated on the first sample audio, and the second difference between the second sample language recognition result and the third sample language recognition result. Since the first sample audio is annotated with the true language information existing in the first sample audio, adjusting the network parameters of the first language recognition network based on the first difference between the first sample language recognition result and the true language information annotated on the first sample audio can minimize the difference between the first sample language recognition result and the true language information, so that the first sample language recognition result predicted by using the first language recognition network approaches the true language information, driving the first language recognition network to identify the language of the first sample audio as accurately as possible. That is, adjusting the network parameters of the first language recognition network based on the first difference between the first sample language recognition result and the true language information annotated on the first sample audio can improve the recognition accuracy of the first language recognition network for the language of the sample audio. In addition, since the third sample audio can be regarded as sample audio data with added interference based on the second sample audio, adjusting the network parameters of the first language recognition network based on the difference between the second sample language recognition result and the third sample language recognition result can minimize the difference between the second sample language recognition result and the third sample language recognition result, so that the second sample language recognition result predicted by using the first language recognition network approaches the third sample language recognition result, driving the first language recognition network to have a basically consistent language recognition effect in the presence of interference as in the absence of interference. Therefore, adjusting the network parameters of the first language recognition network based on the first difference between the first sample language recognition result and the true language information and the second difference between the second sample language recognition result and the third sample language recognition result can enable the first language recognition network after subsequent adjustment of the network parameters to training convergence to have high language recognition accuracy and maintain a good language recognition effect in the presence of interference.
[0082] In a specific embodiment, there are several first sample audios and second sample audios. In each round of training of the first language recognition network, a batch of first sample audios and second sample audios are used. Specifically, the first loss of the first language recognition network can be obtained by combining the first loss function and the first difference between the first sample language recognition result and the true language information; the second loss of the first language recognition network can be obtained by combining the second loss function and the second difference between the second sample language recognition result and the third sample language recognition result; then, according to the first loss and the second loss of the first language recognition network, the total loss of the first language recognition network is obtained; then, using the obtained total loss of the first language recognition network, the network parameters of the first language recognition network are adjusted; the above steps are used to iteratively train the first language recognition network, and finally the first language recognition network with network convergence is obtained, and at this time, the training of the first language recognition network is completed. The specific formula of the total loss function of the first language recognition network is as follows:
[0083]
[0084] Wherein, represents the total loss function of the first language recognition network; represents the first loss function of the first language recognition network; represents the second loss function of the first language recognition network; represents a batch of first sample audios; represents a batch of second sample audios; represents the nth first sample audio in a batch of first sample audios; represents the nth second sample audio in a batch of second sample audios; represents the nth third sample audio in a batch of third sample audios; represents the true language information of the nth first sample audio in a batch of first sample audios; represents the number of first sample audios included in a batch of first sample audios, the number of second sample audios included in a batch of second sample audios, and the number of third sample audios included in a batch of third sample audios; is denoted as the output of the BILSTM network layer, which is actually the splicing of multiple features. Among them, ; represents the characterization feature of the sample audio. Since the contributions of the respective features obtained by splicing to the subsequent language classification are different, splicing them into an embedding for language classification should not be the optimal solution. Therefore, it is input to the ATTENTION layer for weighting to obtain the weighted characterization feature ; ; Represents the classification function corresponding to the OUTPUT layer; Represents calculating the Euclidean distance between two features; Represents the softmax function; Represents adjustable parameters. Here, the specific values of the adjustable parameters are not limited. For example, is 0.2.
[0085] Among them, the specific formula for obtaining the representative feature of the sample audio by weighting each feature is as follows:
[0086]
[0087]
[0088] Among them, Represents the representative feature of the sample audio; Represents each feature output by the BILSTM network layer; Represents the weight corresponding to each feature.
[0089] In order to further improve the language recognition effect of the first language recognition network during training convergence, that is, to improve the language recognition robustness of the first language recognition network. In other embodiments, as Figure 6 shown, Figure 6 is Figure 5 A schematic flowchart of an embodiment of step S54 shown. When there are several first sample audios, it is also possible to adjust the network parameters of the first language recognition network based on the first difference between the first sample language recognition result and the true language information existing in the first sample audio marked on the first sample audio, the second difference between the second sample language recognition result and the third sample language recognition result, and the difference between the first average distance between the representative feature of the current first sample audio and the representative features of each positive sample audio and the second average distance between the representative feature of the current first sample audio and the representative features of each negative sample audio. Specifically, it includes the following sub-steps:
[0090] Step S541: Obtain the first average distance between the representative feature of the current first sample audio and the representative features of each positive sample audio and the second average distance between the representative feature of the current first sample audio and the representative features of each negative sample audio.
[0091] In this embodiment, the first average distance between the characterization features of the current first sample audio and the characterization features of each positive sample audio, and the second average distance between the characterization features of the current first sample audio and the characterization features of each negative sample audio are obtained. Herein, the positive sample audio is the first sample audio with the same language as the current first sample audio, and the negative sample audio is the first sample audio with a different language from the current first sample audio. The characterization features are extracted during the language identification process of the corresponding first sample audio by the first language identification network. In one embodiment, as Figure 2 shown, the characterization features are the outputs of the ATTENTION layer of the first language identification network.
[0092] It should be noted that since a number of first sample audios are used for one training of the first language identification network, that is, since a batch of first sample audios are used for one training of the first language identification network, each first sample audio in the batch of first sample audios needs to be used as the current first sample audio respectively, and the first average value of the characterization features of it and the characterization features of each positive sample audio, as well as the second average distance between the characterization features of it and the characterization features of each negative sample audio are calculated.
[0093] Specifically, first, the distance between the characterization features of the current first sample audio and the characterization features of each positive sample audio, and the distance between the characterization features of the current first sample audio and the characterization features of any negative sample audio are obtained. Among them, the specific formula for calculating the distance between the characterization features of the current first sample audio and the characterization features of any other sample audio is as follows:
[0094]
[0095] Among them, represents the characterization features of the current first sample audio; represents the characterization features of any other first sample audio in the batch of first sample audios except the current first sample audio; represents the distance between the characterization features of the current first sample audio and the characterization features of any other first sample audio.
[0096] Secondly, the average value of the sum of the distances between the characterization features of the current first sample audio and the characterization features of each positive sample audio is calculated as the first average distance, denoted as , where represents the current first sample audio; and, the average value of the sum of the distances between the characterization features of the current first sample audio and the characterization features of each negative sample audio is calculated as the second average distance, denoted as , where represents the current first sample audio.
[0097] Step S542: Adjust the network parameters of the first language recognition network based on the first difference, the second difference, and the difference between the first average distance and the second average distance.
[0098] In this embodiment, based on the first difference between the first sample language recognition result and the true language information existing in the first sample audio marked on the first sample audio, the second difference between the second sample language recognition result and the third sample language recognition result, and the difference between the first average distance between the feature representations of the current first sample audio and the feature representations of each positive sample audio and the second average distance between the feature representations of the current first sample audio and the feature representations of each negative sample audio, the network parameters of the first language recognition network are adjusted.
[0099] Specifically, the first loss of the first language recognition network can be obtained by combining the first loss function and the first difference between the first sample language recognition result and the true language information; the second loss of the first language recognition network can be obtained by combining the second loss function and the second difference between the second sample language recognition result and the third sample language recognition result; the third loss of the first language recognition network can be obtained by combining the third loss function and the difference between the first average distance and the second average distance; then, according to the first loss, the second loss, and the third loss of the first language recognition network, the total loss of the first language recognition network is obtained; then, using the obtained total loss of the first language recognition network, the network parameters of the first language recognition network are adjusted; the first language recognition network is iteratively trained using the above steps, and finally, the first language recognition network with network convergence is obtained, and at this time, the training of the first language recognition network is completed. Among them, the specific formula of the total loss function of the first language recognition network is as follows:
[0100]
[0101] Among them, represents the total loss function of the first language recognition network; represents the first loss function of the first language recognition network; represents the second loss function of the first language recognition network; represents the third loss function of the first language recognition network; represents the first average distance; represents the second average distance; represents a margin parameter value used to control the dispersion degree between the positive sample audio and the negative sample audio. Among them, its specific value is not limited. For example, the value is 0.2; and represent adjustable parameters. Among them, the specific values of the adjustable parameters are not limited. For example, is 0.2, is 1.
[0102] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of an embodiment for obtaining target acoustic features provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 7 the process sequence shown. As shown in Figure 7 , BN features are extracted from the SDC features of the data to be extracted by using a feature extraction network as the target acoustic features, specifically including:
[0103] Step 71: Obtain the shifted differential cepstrum features of the data to be extracted.
[0104] In this embodiment, first, the shifted differential cepstrum features of the data to be extracted are obtained. In one embodiment, the SDC features can be extracted from the data to be extracted by using a relevant extraction network. It can be understood that in other embodiments, the SDC features can also be extracted from the data to be extracted by using relevant extraction algorithms.
[0105] Among them, the data to be extracted is the initial audio or the audio to be recognized.
[0106] Step 72: Use the feature extraction network to perform feature extraction on the shifted differential cepstrum features of the data to be extracted to obtain bottleneck features, and use the bottleneck features as the target acoustic features of the data to be extracted.
[0107] In this embodiment, the feature extraction network is used to perform feature extraction on the shifted differential cepstrum features of the data to be extracted to obtain the bottleneck features of the data to be extracted, and the bottleneck features are used as the target acoustic features of the data to be extracted. Since the BN features are more representative of languages and more resistant to noise, using the BN features as the target acoustic features of the data to be extracted can make the initial language recognition result obtained by the subsequent first language recognition network based on the target acoustic features of the data to be extracted more accurate.
[0108] Please refer to Figure 3 and Figure 8 , Figure 8 which is a schematic flowchart of an embodiment for training the feature extraction network provided by this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 8 the process sequence shown. As shown in Figure 8 , the feature extraction network further includes an output layer in the training stage, that is, the feature extraction network includes a deep neural network layer, a bottleneck network layer, and an output layer connected in sequence in the training stage. The output layer is used to predict the bottleneck features extracted by the feature extraction network to obtain the corresponding language recognition result. Among them, the output layer is fully connected to the bottleneck network layer, and the number of its nodes is the same as the number of language categories. The training steps of the feature extraction network specifically include:
[0109] Step S81: Obtain a fourth sample audio and a fifth sample audio.
[0110] The method of this embodiment is used to train a feature extraction network based on a fourth sample audio and a fifth sample audio, so that the features extracted by the trained feature extraction network are more linguistically representative and noise-resistant. Therefore, in this embodiment, a fourth sample audio and a fifth sample audio are obtained, where the fourth sample audio is labeled with the true language information existing in the fourth sample audio, and the fifth sample audio is not labeled with the true language information existing in the fifth sample audio. That is to say, in the training of the feature extraction network, a large amount of unsupervised sample audio data is added, so that the features extracted by the trained feature extraction network from the audio data are noise-resistant, that is, the trained feature extraction network can accurately extract linguistically representative features even in the presence of interference.
[0111] In one embodiment, the fourth sample audio and the fifth sample audio can be specifically obtained from local storage or cloud storage. It can be understood that in other embodiments, they can also be obtained by collecting the current environmental sounds through a voice collection device, which is not specifically limited here.
[0112] In one embodiment, the fourth sample audio and the fifth sample audio each include several. Subsequently, several fourth sample audios and several fifth sample audios can be input into the feature extraction network at the same time. That is, in the subsequent training of the feature extraction network once, the feature extraction network is trained using a batch of sample audios, which improves the efficiency of training the feature extraction network. Among them, the number of the fourth sample audios and the fifth sample audios input into the feature extraction network in batches is not limited and can be specifically set according to actual usage needs. It can be understood that in other ways, the fourth sample audio and the fifth sample audio are each one, that is, in the subsequent training of the feature extraction network once, the feature extraction network is trained using a single sample audio.
[0113] Among them, the specific languages involved in the fourth sample audio and the fifth sample audio are not limited. For example, the languages involved in the fourth sample audio and the fifth sample audio are Chinese, English, German, etc. It should be noted that when the feature extraction network is subsequently trained using a batch of the fourth sample audios and the fifth sample audios, the specific languages involved in each of the fourth sample audios can be the same or different, and the specific languages involved in each of the fifth sample audios can be the same or different.
[0114] In one embodiment, the fourth sample audio and the fifth sample audio can be preprocessed to filter out invalid sounds such as noise and intermittent sounds, so as to improve the quality of the fourth sample audio and the fifth sample audio, and thus make the effect of the subsequent feature extraction network trained based on the fourth sample audio and the fifth sample audio better.
[0115] In one embodiment, the fourth sample audio and the fifth sample audio can also be augmented, such as by changing the speed, to make the number of samples for training the feature extraction network more sufficient and more diverse, so that the subsequent trained feature extraction network has better generalization ability.
[0116] Step S82: Randomly mask the fifth sample audio to obtain a sixth sample audio.
[0117] In this embodiment, the fifth sample audio is randomly masked to obtain a sixth sample audio. Specifically, a random tfmask is applied to the fifth sample audio to obtain the sixth sample audio. That is to say, the difference between the fifth sample audio and the sixth sample audio lies in whether the tfmask is applied, which is equivalent to creating an interference on the fifth sample audio to obtain the sixth sample audio, so that after the network parameters of the feature extraction network are adjusted based on the differences between the fifth sample audio and the sixth sample audio and the training converges, the trained feature extraction network can accurately extract language-representative features even in the presence of interference.
[0118] In one embodiment, when the feature extraction network is trained once using a single fifth sample audio subsequently, at this time, the random audio frames of this fifth sample audio can be randomly masked to obtain a sixth sample audio.
[0119] In one embodiment, when the feature extraction network is trained once using a batch of fifth sample audios subsequently, at this time, the random audio frames of any fifth sample audio in the batch of fifth sample audios can be randomly masked to obtain a batch of sixth sample audios.
[0120] Step S83: Use the feature extraction network in the training stage to process the fourth sample audio, the fifth sample audio, and the sixth sample audio respectively, and obtain the fourth sample language recognition result, the fifth sample language recognition result, and the sixth sample language recognition result correspondingly.
[0121] In this embodiment, the feature extraction network in the training phase is used to process the fourth sample audio, the fifth sample audio, and the sixth sample audio respectively, and the fourth sample language recognition result, the fifth sample language recognition result, and the sixth sample language recognition result are obtained correspondingly. That is to say, when using the feature extraction network to perform language recognition on the fourth sample audio, the fifth sample audio, and the sixth sample audio, the language prediction result corresponding to the fourth sample audio, that is, the fourth sample language recognition result, the language prediction result corresponding to the fifth sample audio, that is, the fifth sample language recognition result, and the language prediction result corresponding to the sixth sample audio, that is, the sixth sample language recognition result, will be obtained.
[0122] In one embodiment, the fourth sample audio, the fifth sample audio, and the sixth sample audio are single audio data. That is, when using the single fourth sample audio, the fifth sample audio, and the sixth sample audio for one training of the feature extraction network, the fourth sample language recognition result obtained at this time only includes the language recognition result corresponding to the single fourth sample audio, the fifth sample language recognition result only includes the language recognition result corresponding to the single fifth sample audio, and the sixth sample language recognition result only includes the language recognition result corresponding to the single sixth sample audio.
[0123] In one embodiment, when there are several fourth sample audios, fifth sample audios, and sixth sample audios, that is, when using a batch of fourth sample audios, fifth sample audios, and sixth sample audios for one training of the feature extraction network, the fourth sample language recognition result obtained at this time includes the language recognition results corresponding to each fourth sample audio, the fifth sample language recognition result includes the language recognition results corresponding to each fifth sample audio, and the sixth sample language recognition result includes the language recognition results corresponding to each sixth sample audio.
[0124] Step S84: Adjust the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information, and the difference between the fifth sample language recognition result and the sixth sample language recognition result.
[0125] In this embodiment, based on the differences between the fourth sample language recognition result and the true language information existing in the fourth sample audio marked on the fourth sample audio, and the differences between the fifth sample language recognition result and the sixth sample language recognition result, the network parameters of the feature extraction network are adjusted. Since the fourth sample audio is marked with the true language information existing in the fourth sample audio, adjusting the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information marked on the fourth sample audio can minimize the difference between the fourth sample language recognition result and the true language information, so that the fourth sample language recognition result predicted by using the feature extraction network approaches the true language information, driving the feature extraction network to recognize the language of the fourth sample audio as accurately as possible. That is, adjusting the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information marked on the fourth sample audio can improve the recognition accuracy of the feature extraction network for the language of the sample audio. Among them, after the feature extraction network is trained, the output layer of the feature extraction network is removed. Therefore, adjusting the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information is equivalent to improving the ability of the feature extraction network to mine the hidden language representation information in the sample audio data features. In addition, since the fifth sample audio can be regarded as the sample audio data with interference added on the basis of the sixth sample audio, adjusting the network parameters of the feature extraction network based on the difference between the fifth sample language recognition result and the sixth sample language recognition result can minimize the difference between the fifth sample language recognition result and the sixth sample language recognition result, so that the fifth sample language recognition result predicted by using the feature extraction network approaches the sixth sample language recognition result, driving the feature extraction network to have a basically consistent language recognition effect in the case of interference as in the case of no interference. Among them, after the feature extraction network is trained, the output layer of the feature extraction network is removed. Therefore, adjusting the network parameters of the feature extraction network based on the difference between the fifth sample language recognition result and the sixth sample language recognition result is equivalent to improving the ability of the feature extraction network to mine the hidden language representation information in the sample audio data features in the case of interference. Therefore, adjusting the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information and the difference between the fifth sample language recognition result and the sixth sample language recognition result can enable the feature extraction network after subsequent training convergence to have a high ability to mine the hidden language representation information in the sample audio data features, and can maintain a good ability to mine the hidden language representation information in the sample audio data in the case of interference.
[0126] In a specific embodiment, there are several fourth sample audios and fifth sample audios. In each round of training of the feature extraction network, a batch of fourth sample audios and fifth sample audios are used. Specifically, the first loss of the feature extraction network can be obtained by combining the fourth loss function and the difference between the fourth sample language recognition result and the true language information; the second loss of the feature extraction network can be obtained by combining the fifth loss function and the difference between the fifth sample language recognition result and the sixth sample language recognition result; then, according to the first loss and the second loss of the feature extraction network, the total loss of the feature extraction network is obtained; then, using the obtained total loss of the feature extraction network, the network parameters of the feature extraction network are adjusted; the above steps are used to iteratively train the feature extraction network, and finally a feature extraction network with network convergence is obtained, and at this time, the training of the feature extraction network is completed. The specific formula of the total loss function of the feature extraction network is as follows:
[0127]
[0128]
[0129] Wherein, represents the total loss function of the feature extraction network; represents the first loss function of the feature extraction network; represents the second loss function of the feature extraction network; represents a batch of fourth sample audios; represents a batch of fifth sample audios; represents the nth fourth sample audio in a batch of fourth sample audios; represents the nth fifth sample audio in a batch of fifth sample audios; represents the nth sixth sample audio in a batch of sixth sample audios; represents the true language information of the nth fourth sample audio in a batch of fourth sample audios; represents the number of fourth sample audios included in a batch of fourth sample audios, the number of fifth sample audios included in a batch of fifth sample audios, and the number of sixth sample audios included in a batch of sixth sample audios; is denoted as the output of the bottleneck network layer; represents the classification function corresponding to the output layer; represents calculating the Euclidean distance between two features; represents the softmax function; represents adjustable parameters. Among them, the specific values of the adjustable parameters are not limited. For example, is 0.2.
[0130] Please refer to Figure 9 ,Figure 9 It is a schematic structural diagram of an embodiment of the language identification device provided by this application. The language identification device 90 includes a first identification module 91, a detection module 92, and a second identification module 93. The first identification module 91 is used to perform language identification on the audio to be identified by using a first language identification network, and obtain an initial language identification result; the detection module 92 is used to detect whether the initial language identification result meets a preset identification requirement; the second identification module 93 is used to respond to the situation that the initial language identification result does not meet the preset identification requirement, and perform language identification on the audio to be identified by using a second language identification network, and obtain a target language identification result; wherein, the first language identification network has a stronger identification ability for audio in the first language situation than the second language identification network, and the second language identification network has a stronger identification ability for audio in the second language situation than the first language identification network.
[0131] Wherein, the above-mentioned first language situation is a single language, the second language situation is a multi-language, the first language identification network can identify a single language, and the second language identification network can identify at least one language.
[0132] Wherein, the above-mentioned initial language identification result includes an initial language existing in the audio to be identified and a confidence score corresponding to the initial language; the preset identification requirement includes that the confidence score meets a preset score requirement.
[0133] Wherein, the above-mentioned second language identification network is a hidden Markov model, and the hidden Markov model is composed of a plurality of Gaussian mixture models spliced together, and each Gaussian mixture model is used to identify and obtain a language.
[0134] Wherein, the language identification device 90 further includes a training module 94, and the training steps of the training module 94 for the first language identification network specifically include: obtaining a first sample audio and a second sample audio; wherein, the first sample audio is marked with the true language information existing in the first sample audio, and the second sample audio is not marked; performing random masking processing on the second sample audio to obtain a third sample audio; using the first language identification network to perform language identification on the first sample audio, the second sample audio, and the third sample audio, and correspondingly obtaining a first sample language identification result, a second sample language identification result, and a third sample language identification result; adjusting the network parameters of the first language identification network at least based on a first difference between the first sample language identification result and the true language information, and a second difference between the second sample language identification result and the third sample language identification result.
[0135] Among them, there are several of the above first sample audios. The training module 94 is used to adjust the network parameters of the first language recognition network at least based on the first difference between the first sample language recognition result and the true language information, and the second difference between the second sample language recognition result and the third sample language recognition result. Specifically, it includes: obtaining the first average distance between the representation features of the current first sample audio and the representation features of each positive sample audio, and the second average distance between the representation features of the current first sample audio and the representation features of each negative sample audio; where the positive sample audio is a first sample audio with the same language as the current first sample audio, and the negative sample audio is a first sample audio with a different language from the current first sample audio. The representation features are extracted during the language recognition process of the corresponding first sample audio by the first language recognition network; based on the first difference, the second difference, and the difference between the first average distance and the second average distance, adjust the network parameters of the first language recognition network.
[0136] Among them, the language recognition device 90 further includes a feature extraction module 95. The feature extraction module 95 is used to, before using the first language recognition network to perform language recognition on the audio to be recognized and obtaining the initial language recognition result, specifically including: performing feature extraction on the audio to be recognized to obtain the target acoustic features of the audio to be recognized; or, performing feature extraction on the initial audio to obtain the target acoustic features of the initial audio, and extracting the target acoustic features with a preset time length from the target acoustic features of the initial audio as the target acoustic features of the audio to be recognized; The first recognition module 91 is used to use the first language recognition network to perform language recognition on the audio to be recognized to obtain the initial language recognition result, specifically including: using the first language recognition network to perform language recognition on the target acoustic features to obtain the initial language recognition result; The second recognition module 93 is used to use the second language recognition network to perform language recognition on the audio to be recognized to obtain the target language recognition result, specifically including: using the second language recognition network to perform language recognition on the target acoustic features to obtain the target language recognition result.
[0137] Among them, the feature extraction module 95 is used to perform feature extraction on the initial audio to obtain the target acoustic features of the initial audio, or the feature extraction module 95 is used to perform feature extraction on the audio to be recognized to obtain the target acoustic features of the audio to be recognized, specifically including: obtaining the shifted differential cepstral features of the data to be extracted; using the feature extraction network to perform feature extraction on the shifted differential cepstral features of the data to be extracted to obtain the bottleneck features, and taking the bottleneck features as the target acoustic features of the data to be extracted; where the data to be extracted is the initial audio or the audio to be recognized.
[0138] Among them, the above-mentioned feature extraction network includes a deep neural network layer and a bottleneck network layer connected in sequence, and the bottleneck features are output by the bottleneck network layer; and / or, the feature extraction network includes an output layer in the training stage, and the output layer is used to predict the bottleneck features extracted by the feature extraction network to obtain the corresponding language recognition result; the language recognition method further includes: obtaining a fourth sample audio and a fifth sample audio; among them, the fourth sample audio is marked with the true language information existing in the fourth sample audio, and the fifth sample audio is not marked; randomly masking the fifth sample audio to obtain a sixth sample audio; using the feature extraction network in the training stage to process the fourth sample audio, the fifth sample audio and the sixth sample audio respectively, and correspondingly obtaining the fourth sample language recognition result, the fifth sample language recognition result and the sixth sample language recognition result; based on the difference between the fourth sample language recognition result and the true language information, and the difference between the fifth sample language recognition result and the sixth sample language recognition result, adjusting the network parameters of the feature extraction network; among them, after the feature extraction network is trained, the output layer is removed.
[0139] Among them, the target language recognition result includes at least one target language existing in the audio to be recognized and the time interval corresponding to each target language in the audio to be recognized; and / or, the audio to be recognized is a voice segment of a preset time length in the initial audio, and the initial audio includes several audios to be recognized corresponding to different time periods; the language recognition device 90 further includes a combining module 96, and the combining module 96 is used for, after using the second language recognition network to perform language recognition on the audio to be recognized to obtain the target language recognition result, specifically including: combining the target language recognition results of each audio to be recognized included in the initial audio to obtain the language recognition result of the initial audio.
[0140] Please refer to Figure 10 , Figure 10 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. The electronic device 100 includes a memory 101 and a processor 102 that are coupled to each other, and the processor 102 is configured to execute program instructions stored in the memory 101 to implement the steps of any of the above-mentioned language recognition method embodiments. In a specific implementation scenario, the electronic device 100 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 100 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0141] Specifically, the processor 102 is used to control itself and the memory 101 to implement the steps of any of the above-described method embodiments for language recognition. The processor 102 may also be referred to as a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 102 may be implemented jointly by integrated circuit chips.
[0142] Please refer to Figure 11 , Figure 11 FIG. is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 110 of the embodiment of the present application stores program instructions 111, and when the program instructions 111 are executed, the methods provided by any of the embodiments of the language recognition method of the present application and any non-conflicting combinations are implemented. Among them, the program instructions 111 may form a program file and be stored in the above-mentioned computer-readable storage medium 110 in the form of a software product, so that a computer device (which may be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium 110 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.
[0143] If the technical solution of the present application involves personal information, before the product applying the technical solution of the present application processes personal information, it has clearly informed the personal information processing rules and obtained the independent consent of the individual. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the device for processing personal information, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0144] The above are only the implementation manners of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.
Claims
1. A language identification method, characterized in that, The method includes: Performing language identification on the audio to be identified by using a first language identification network to obtain an initial language identification result; Detecting whether the initial language identification result meets a preset identification requirement; In response to the initial language identification result not meeting the preset identification requirement, performing language identification on the audio to be identified by using a second language identification network to obtain a target language identification result; wherein, the first language identification network has a stronger ability to identify audio in a first language situation than the second language identification network, the second language identification network has a stronger ability to identify audio in a second language situation than the first language identification network, the first language situation is a single language, the second language situation is a multi-language, the first language identification network can identify a single language, and the second language identification network can identify at least one language.
2. The method according to claim 1, characterized in that, The initial language identification result includes an initial language existing in the audio to be identified and a confidence score corresponding to the initial language; the preset identification requirement includes that the confidence score meets a preset score requirement.
3. The method according to claim 1, characterized in that, The second language identification network is a hidden Markov model, and the hidden Markov model is composed of a plurality of Gaussian mixture models spliced together, and each Gaussian mixture model is used to identify a language.
4. The method according to claim 1, wherein The method further includes the following training steps for the first language identification network: Obtaining a first sample audio and a second sample audio; wherein, the first sample audio is labeled with the true language information existing in the first sample audio, and the second sample audio is not labeled; Performing masking processing on random audio frames of the second sample audio to obtain a third sample audio; Performing language identification on the first sample audio, the second sample audio, and the third sample audio by using the first language identification network to correspondingly obtain a first sample language identification result, a second sample language identification result, and a third sample language identification result; Adjusting the network parameters of the first language identification network based at least on a first difference between the first sample language identification result and the true language information, and a second difference between the second sample language identification result and the third sample language identification result.
5. The method according to claim 4, wherein There are a plurality of the first sample audios, and the adjusting the network parameters of the first language identification network based at least on the first difference between the first sample language identification result and the true language information, and the second difference between the second sample language identification result and the third sample language identification result includes: Obtaining a first average distance between the characterization feature of the current first sample audio and the characterization features of each positive sample audio, and a second average distance between the characterization feature quantity of the current first sample audio and the characterization features of each negative sample audio; wherein, the positive sample audio is the first sample audio having the same language as the current first sample audio, the negative sample audio is the first sample audio having a different language from the current first sample audio, and the characterization feature is extracted during the process of the first language identification network performing language identification on the corresponding first sample audio. Adjust the network parameters of the first language recognition network based on the first difference, the second difference, and the difference between the first average distance and the second average distance.
6. The method according to claim 1, wherein Before using the first language recognition network to perform language recognition on the audio to be recognized and obtaining the initial language recognition result, the method further includes: Performing feature extraction on the audio to be recognized to obtain the target acoustic features of the audio to be recognized; or, performing feature extraction on the initial audio to obtain the target acoustic features of the initial audio, and extracting the target acoustic features with a preset time length from the target acoustic features of the initial audio as the target acoustic features of the audio to be recognized; The using the first language recognition network to perform language recognition on the audio to be recognized and obtaining the initial language recognition result includes: Using the first language recognition network to perform language recognition on the target acoustic features to obtain the initial language recognition result; The using the second language recognition network to perform language recognition on the audio to be recognized and obtaining the target language recognition result includes: Using the second language recognition network to perform language recognition on the target acoustic features to obtain the target language recognition result.
7. The method according to claim 6, wherein The performing feature extraction on the initial audio to obtain the target acoustic features of the initial audio, or the performing feature extraction on the audio to be recognized to obtain the target acoustic features of the audio to be recognized includes: Obtaining the shifted differential cepstral features of the data to be extracted; Using the feature extraction network to perform feature extraction on the shifted differential cepstral features of the data to be extracted to obtain bottleneck features, and using the bottleneck features as the target acoustic features of the data to be extracted; Wherein, the data to be extracted is the initial audio or the audio to be recognized.
8. The method according to claim 7, wherein The feature extraction network includes a deep neural network layer and a bottleneck network layer connected in sequence, and the bottleneck features are output by the bottleneck network layer; And / or, the feature extraction network includes an output layer in the training stage, and the output layer is used to predict the bottleneck features extracted by the feature extraction network to obtain the corresponding language recognition result; the method further includes: Obtaining a fourth sample audio and a fifth sample audio; wherein, the fourth sample audio is labeled with the true language information existing in the fourth sample audio, and the fifth sample audio is not labeled; Performing masking processing on random audio frames of the fifth sample audio to obtain a sixth sample audio; Using the feature extraction network in the training stage to process the fourth sample audio, the fifth sample audio, and the sixth sample audio respectively, and correspondingly obtaining a fourth sample language recognition result, a fifth sample language recognition result, and a sixth sample language recognition result; Adjusting the network parameters of the feature extraction network based on the difference between the fourth sample language recognition result and the true language information, and the difference between the fifth sample language recognition result and the sixth sample language recognition result; wherein, after the feature extraction network is trained, the output layer is removed.
9. The method according to claim 1, wherein The target language recognition result includes at least one target language existing in the audio to be recognized and the time interval corresponding to each target language in the audio to be recognized; and / or, the audio to be recognized is a voice segment with a preset time length in the initial audio, and the initial audio includes a plurality of the audios to be recognized corresponding to different time periods; After using the second language recognition network to perform language recognition on the audio to be recognized and obtaining the target language recognition result, the method further includes: Combining the target language recognition results of the audios to be recognized included in the initial audio to obtain the language recognition result of the initial audio.
10. A language identification device, characterized in that, The device includes: A first recognition module, configured to use a first language recognition network to perform language recognition on the audio to be recognized to obtain an initial language recognition result; A detection module, configured to detect whether the initial language recognition result meets a preset recognition requirement; A second recognition module, configured to, in response to the initial language recognition result not meeting the preset recognition requirement, use a second language recognition network to perform language recognition on the audio to be recognized to obtain a target language recognition result; wherein, the first language recognition network has a stronger recognition ability for audios in a first language situation than the second language recognition network, the second language recognition network has a stronger recognition ability for audios in a second language situation than the first language recognition network, the first language situation is a single language, the second language situation is a multi-language, the first language recognition network can recognize a single language, and the second language recognition network can recognize at least one language.
11. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the language recognition method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the language recognition method according to any one of claims 1-9.
Citation Information
Patent Citations
Method and device for realizing call forwarding according to language
CN104580762A
Language recognition method, device, translator, medium and equipment
CN109147769A
Language recognition method and device, and device for language recognition
CN110930978A