Language recognition method and device, storage medium and electronic device

CN119380698BActive Publication Date: 2026-09-08SHENZHEN HEYTAP TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310935106.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2026-09-08
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

同时,使用语音识别技术的用户也越来越多,这些用户可能来自不同的地区,使用不同的语种,因此针对用户输入的待识别音频,需要识别待识别音频所属的语种类别,但是相关技术中,如果需要识别多个语种的待识别音频,需要切换使用多个语种识别模型,且占用的计算资源较大

Benefits of technology

[0018]In this embodiment, the language recognition method uses a language recognition model to identify the language of the audio to be recognized. The language recognition model can reuse the acoustic model in the speech recognition model. The high-dimensional acoustic features extracted from the audio by the acoustic model are used as the input of the language recognition model. It is not necessary to set up a separate layer with similar function to the acoustic model of the speech recognition model to extract high-dimensional acoustic features, nor is it necessary to train this layer. This reduces the occupation of computing resources and avoids cost waste. More computing resources can be reserved for training other layers in the language recognition model to improve the performance of the language recognition model and improve the accuracy of language recognition. Furthermore, in this embodiment, the high-dimensional acoustic features are converted into phoneme sequences through a phoneme conversion layer for subsequent language identification, rather than using text information. This means that regardless of the language, even if it's a less common language, the audio to be identified can have its high-dimensional acoustic features extracted using an acoustic model and converted into phoneme sequences through the phoneme conversion layer. Therefore, even if training sample resources for some less common languages ​​are limited, the phoneme conversion layer can be trained using training samples from major languages ​​such as Chinese and English, making the trained phoneme conversion layer applicable to less common languages ​​as well. Additionally, the language feature extraction layer and language identification layer of the language identification model can also be trained using training samples from major languages ​​such as Chinese and English, making them applicable to major languages. When language identification is needed for audio in a less common language, the training samples from the less common language can be input into the language feature extraction layer and language identification layer trained with major languages ​​for training with a small amount of data, achieving language identification results close to those of major languages. Therefore, the embodiments of this application do not require training a separate language recognition model for each language. A single language recognition model can be used to recognize multiple languages. When recognizing different languages, there is no need to switch the corresponding language recognition model. Moreover, especially for less common languages, there is no need to train a separate language recognition model using training samples for the less common language. With fewer training sample resources, a high language recognition accuracy can still be obtained through the language recognition model of the embodiments of this application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380698B_ABST
    Figure CN119380698B_ABST
Patent Text Reader

Abstract

A language recognition method, device, storage medium and electronic equipment. The method comprises: obtaining to-be-recognized audio that needs to be subjected to language recognition; inputting the to-be-recognized audio into an acoustic model of a speech recognition model, extracting high-dimensional acoustic features of the to-be-recognized audio through the acoustic model; inputting the high-dimensional acoustic features into a phoneme conversion layer of the language recognition model, converting the high-dimensional acoustic features into a phoneme sequence of the to-be-recognized audio through the phoneme conversion layer; inputting the phoneme sequence into a language feature extraction layer of the language recognition model, extracting first language features of the to-be-recognized audio through the language feature extraction layer; inputting the first language features into a language recognition layer of the language recognition model, performing language recognition through the language recognition layer, and obtaining a language recognition result of the to-be-recognized audio. The language recognition method of the present application uses one language recognition model to recognize to-be-recognized audio of multiple languages, and can reduce the occupation of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech technology, and in particular to a language recognition method, apparatus, storage medium and electronic device. Background Technology

[0002] With the rapid development of speech recognition technology, it has been widely applied in various scenarios, such as voice input, voice search, voice translation, and smart homes. At the same time, the number of users of speech recognition technology is also increasing. These users may come from different regions and speak different languages. Therefore, for the audio input to be recognized, it is necessary to identify the language category of the audio. However, in related technologies, if it is necessary to recognize audio in multiple languages, multiple language recognition models need to be switched, which consumes a large amount of computing resources. Summary of the Invention

[0003] This application provides a language recognition method, apparatus, storage medium, and electronic device. The language recognition method can use the same language recognition model to recognize audio in multiple languages ​​and can reduce the consumption of computing resources.

[0004] In a first aspect, embodiments of this application provide a language identification method, the method comprising:

[0005] Obtain the audio to be identified, which requires language recognition;

[0006] The audio to be recognized is input into the acoustic model of the speech recognition model, and the high-dimensional acoustic features of the audio to be recognized are extracted through the acoustic model.

[0007] The high-dimensional acoustic features are input into the phoneme conversion layer of the language recognition model, and the high-dimensional acoustic features are converted into the phoneme sequence of the audio to be recognized through the phoneme conversion layer.

[0008] The phoneme sequence is input into the language feature extraction layer of the language recognition model, and the first language feature of the audio to be recognized is extracted through the language feature extraction layer.

[0009] The first language feature is input into the language recognition layer of the language recognition model, and language recognition is performed through the language recognition layer to obtain the language recognition result of the audio to be recognized.

[0010] Secondly, embodiments of this application provide a language recognition device, the device comprising:

[0011] The acquisition module is used to acquire the audio to be identified for language recognition.

[0012] The first extraction module is used to input the audio to be recognized into the acoustic model of the speech recognition model, and extract the high-dimensional acoustic features of the audio to be recognized through the acoustic model.

[0013] The conversion module is used to input the high-dimensional acoustic features into the phoneme conversion layer of the language recognition model, and convert the high-dimensional acoustic features into a phoneme sequence of the audio to be recognized through the phoneme conversion layer;

[0014] The second extraction module is used to input the phoneme sequence into the language feature extraction layer of the language recognition model, and extract the first language feature of the audio to be recognized through the language feature extraction layer;

[0015] The recognition module is used to input the first language feature into the language recognition layer of the language recognition model, perform language recognition through the language recognition layer, and obtain the language recognition result of the audio to be recognized.

[0016] The storage medium provided in this application stores a computer program that, when loaded by a processor, executes the steps in the pronunciation skill detection method provided in this application.

[0017] The electronic device provided in this application includes a processor and a memory, the memory storing a computer program, and the processor loading the computer program to execute the steps in the pronunciation skill detection method provided in this application.

[0018] In this embodiment, the language recognition method uses a language recognition model to identify the language of the audio to be recognized. The language recognition model can reuse the acoustic model in the speech recognition model. The high-dimensional acoustic features extracted from the audio by the acoustic model are used as the input of the language recognition model. It is not necessary to set up a separate layer with similar function to the acoustic model of the speech recognition model to extract high-dimensional acoustic features, nor is it necessary to train this layer. This reduces the occupation of computing resources and avoids cost waste. More computing resources can be reserved for training other layers in the language recognition model to improve the performance of the language recognition model and improve the accuracy of language recognition. Furthermore, in this embodiment, the high-dimensional acoustic features are converted into phoneme sequences through a phoneme conversion layer for subsequent language identification, rather than using text information. This means that regardless of the language, even if it's a less common language, the audio to be identified can have its high-dimensional acoustic features extracted using an acoustic model and converted into phoneme sequences through the phoneme conversion layer. Therefore, even if training sample resources for some less common languages ​​are limited, the phoneme conversion layer can be trained using training samples from major languages ​​such as Chinese and English, making the trained phoneme conversion layer applicable to less common languages ​​as well. Additionally, the language feature extraction layer and language identification layer of the language identification model can also be trained using training samples from major languages ​​such as Chinese and English, making them applicable to major languages. When language identification is needed for audio in a less common language, the training samples from the less common language can be input into the language feature extraction layer and language identification layer trained with major languages ​​for training with a small amount of data, achieving language identification results close to those of major languages. Therefore, the embodiments of this application do not require training a separate language recognition model for each language. A single language recognition model can be used to recognize multiple languages. When recognizing different languages, there is no need to switch the corresponding language recognition model. Moreover, especially for less common languages, there is no need to train a separate language recognition model using training samples for the less common language. With fewer training sample resources, a high language recognition accuracy can still be obtained through the language recognition model of the embodiments of this application. Attached Figure Description

[0019] The technical solution and its beneficial effects will become apparent from the following detailed description of specific embodiments of this application, in conjunction with the accompanying drawings.

[0020] Figure 1 This is a schematic diagram of the first language identification method provided in the embodiments of this application.

[0021] Figure 2 This is a first structural block diagram of the language recognition model provided in the embodiments of this application.

[0022] Figure 3This is a schematic diagram of the second language recognition method provided in the embodiments of this application.

[0023] Figure 4 This is a first structural block diagram of the language recognition model provided in the embodiments of this application.

[0024] Figure 5 This is a schematic diagram of the training process of the language recognition model provided in the embodiments of this application.

[0025] Figure 6 This is a schematic diagram illustrating the training of the phoneme conversion layer of the language recognition model provided in this application embodiment.

[0026] Figure 7 This is a training diagram of the language feature extraction layer, feature pooling layer, and language recognition layer of the language recognition model provided in the embodiments of this application.

[0027] Figure 8 This is a training diagram of the phoneme conversion layer, language feature extraction layer, feature pooling layer, and language recognition layer of the language recognition model provided in the embodiments of this application.

[0028] Figure 9 This is a schematic diagram of the first structure of the language recognition device provided in the embodiments of this application.

[0029] Figure 10 This is a second structural schematic diagram of the language recognition device provided in the embodiments of this application.

[0030] Figure 11 This is a schematic diagram of a first structure of an electronic device provided in an embodiment of this application.

[0031] Figure 12 This is a schematic diagram of a second structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0032] It should be noted that the principles of this application are illustrated by example in a suitable computing environment. The following description is based on the specific embodiments of this application exemplified, and should not be considered as limiting other specific embodiments not detailed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0033] The relational terms such as "first" and "second" used in the following embodiments of this application are only used to distinguish one object or operation from another, and are not intended to limit the actual order of these objects or operations. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0034] Artificial intelligence (AI) is the theory, methods, technology, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0035] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include machine learning (ML), with deep learning (DL) being a relatively new research direction within ML. It has been introduced into machine learning to bring it closer to its original goal: artificial intelligence. Currently, deep learning is mainly applied in fields such as computer vision and natural language processing.

[0036] Deep learning learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly aids in the interpretation of data such as text, images, and sound. Using deep learning techniques and corresponding training datasets, it is possible to train network models that perform different functions. For example, a deep learning network for gender classification can be trained based on one training dataset, while a deep learning network for image optimization can be trained based on another training dataset.

[0037] This application introduces deep learning into language recognition of multilingual audio, providing a language recognition method, a language recognition device, a storage medium, and an electronic device. The language recognition method can be executed by an electronic device, which can be any device equipped with a processor and possessing processing capabilities, such as mobile electronic devices with processors like smartphones, tablets, PDAs, and laptops, or fixed electronic devices with processors like desktop computers, televisions, and servers. Please refer to [link / reference]. Figure 1 , Figure 1 This is a schematic diagram of the first language identification method provided in the embodiments of this application.

[0038] In step 101, obtain the audio to be identified that requires language recognition.

[0039] It is understandable that users may come from different regions of the world and use different languages. Therefore, the audio to be identified obtained in this embodiment can be a segment of audio in a specific language, such as Chinese or a foreign language, such as English, French, or Japanese; it can also be Mandarin or a dialect, such as Sichuanese, Minnan, Northeastern Mandarin, or Cantonese. The audio to be identified can be collected through devices such as a microphone; this application does not specifically limit the method of acquiring the audio to be identified.

[0040] In step 102, the audio to be recognized is input into the acoustic model of the speech recognition model, and the high-dimensional acoustic features of the audio to be recognized are extracted through the acoustic model.

[0041] Please refer to Figure 2 , Figure 2 This is a first structural block diagram of the language recognition model provided in this application embodiment. The language recognition method in this application embodiment is based on the language recognition model to identify the language of the audio to be recognized. The language recognition model is built on the acoustic model in the speech recognition model. That is, the acoustic model in the speech recognition model can be reused in this application embodiment. The speech recognition model can be, for example, a recurrent neural network transducer (RNN-T), an attention-based coding structure, etc.

[0042] Acoustic models can extract features from one-dimensional audio to obtain high-dimensional acoustic features, such as two-dimensional high-dimensional acoustic features. The one-dimensional audio to be identified is, for example, the audio to be identified in the time dimension, while the two-dimensional high-dimensional acoustic features include both time and frequency dimensions. The high-dimensional acoustic features can be, for example, short-time Fourier transform features, thereby obtaining the frequency components of the audio to be identified in different time frames, and obtaining the high-dimensional acoustic features of the audio to be identified in the time-frequency (T*F) two-dimensional plane.

[0043] In related technologies, when using a speech recognition model for speech recognition, if the audio input by the user may be in multiple languages, the speech recognition model usually also includes a language recognition model as a sub-module to identify the language of the audio to be recognized. The language recognition model and the speech recognition model are built independently. That is, the speech recognition model has an acoustic model, while the language recognition model may need to be set up separately with a similar function to the acoustic model of the speech recognition model and be trained accordingly, thus resulting in a waste of computing resources and costs. This application embodiment reuses the acoustic model in the speech recognition model to extract high-dimensional acoustic features of the audio to be recognized. The extracted high-dimensional acoustic features can be used in both the speech recognition model and the language recognition model. That is, the high-dimensional acoustic features extracted by the acoustic model in the speech recognition model can be used as input to the language recognition model for subsequent language recognition. It is not necessary to set up a separate layer with similar function to the acoustic model of the speech recognition model to extract high-dimensional acoustic features, nor is it necessary to train this layer. This reduces the occupation of computing resources, avoids cost waste, and allows more resources to be reserved for training other layers in the language recognition model, so as to improve the performance of the language recognition model and improve the accuracy of language recognition.

[0044] In 103, high-dimensional acoustic features are input into the phoneme conversion layer of the language recognition model, and the high-dimensional acoustic features are converted into a phoneme sequence of the audio to be recognized through the phoneme conversion layer.

[0045] In the embodiment of the present application, the high-dimensional acoustic features are input into the phoneme conversion layer of the language recognition model, and the high-dimensional acoustic features are converted into the phoneme sequence of the audio to be recognized through the phoneme conversion layer. It can be understood that a phoneme is the smallest speech unit divided according to the natural attributes of speech. From the perspective of acoustic properties, a phoneme is the smallest speech unit divided from the perspective of sound quality. One articulation movement forms one phoneme. For example, phonemes can be international phonetic alphabet, Roman phonetic alphabet, pinyin, etc. The phoneme sequence can be understood as being formed by arranging phonemes in the chronological order of the audio to be recognized, and the phoneme sequence takes phonemes as units. For example, taking the Chinese expression "ni hao" (meaning "hello") as an example, its corresponding phoneme sequence is "ni" and "hao". It can be understood that through arranging the phoneme conversion layer in the embodiment of the present application, the audio to be recognized "ni hao" is converted into the phoneme sequence "ni" and "hao", instead of the words "ni" and "hao". In the related art, the acoustic model performs feature extraction on the audio to be recognized, and the obtained high-dimensional acoustic features can be used to convert the substantial content information of the audio to be recognized. For example, the audio to be recognized "ni hao" can be converted into "ni" and "hao", that is, the high-dimensional acoustic features extracted by the acoustic model only retain common content information, and delete some other features such as accent, intonation and other features. It can be understood that for users from different regions, even if the content information in the audio to be recognized is the same, there will still be different accents, intonation and other features. In the embodiment of the present application, continuing to take the audio to be recognized "ni hao" as an example, if it is Mandarin, "ni hao" is converted into the phoneme sequence "ni" and "hao"; if it is Sichuan dialect, due to different pronunciation habits, the phoneme sequence converted from "ni hao" can be "li" and "hao". Therefore, different phoneme sequences can be used to distinguish accent information of different languages, and the embodiment of the present application outputs a finer-grained phoneme sequence through the phoneme conversion layer, which is more conducive to language recognition.

[0046] At step 104, the phoneme sequence is input into the language feature extraction layer of the language recognition model, and the first language feature of the audio to be recognized is extracted through the language feature extraction layer.

[0047] Furthermore, in this embodiment, the phoneme sequence is input into the language feature extraction layer. The language feature extraction layer extracts the first language feature of the audio to be identified. It can be understood that phonemes are universal for different languages. That is, phonemes can be used to annotate audio to be identified in different languages. The first language feature can be understood as the correlation feature between the phoneme sequence and the language. For example, taking Pinyin as an example, if "ni" is input into the language feature extraction layer, the language feature extraction layer will extract that "n" and "i" are often combined together to form "ni" in Mandarin. As another example, if "li" is input into the language feature extraction layer, the language feature extraction layer can extract that "l" and "i" are often combined together to form "li" in Sichuan dialect. It can also be understood that Mandarin and Sichuan dialect may have different phoneme combinations for the same word due to different accents. The first language feature obtained by the language feature extraction layer can extract the first language feature related to the accent of different languages. That is, the first language feature obtained by the language feature extraction layer can reflect the pronunciation habits of different languages ​​to a certain extent.

[0048] For example, continuing with the phoneme-based pinyin example, if the input audio to be recognized is "hello", then the phoneme sequence obtained by the phoneme conversion layer is, for example, "halou". This "halou" is then input into the language feature extraction layer. The language feature extraction layer extracts the phoneme sequence because "h", "a", "l", "o", and "u" often combine to form "halou" in English. This can be understood as the difference in phoneme sequences between Chinese and English being due to different linguistic systems. Different languages ​​correspond to certain specific or commonly used phoneme sequences, or in other words, a certain phoneme sequence has a higher probability of appearing in a certain language. Therefore, through the correlation obtained by the language feature extraction layer, we can extract the first language features related to different languages ​​and linguistic aspects such as semantics. Regardless of whether the difference in phoneme sequences is due to pronunciation habits or semantics, different languages ​​correspond to certain commonly used or commonly used phoneme sequences. Therefore, the extracted first language features can be used to characterize the probability of use and occurrence of phoneme sequences in different languages.

[0049] In step 105, the first language feature is input into the language recognition layer of the language recognition model, and language recognition is performed through the language recognition layer to obtain the language recognition result of the audio to be recognized.

[0050] In this embodiment, the first language feature is input into the language recognition layer of the language recognition model for language recognition. Since the first language feature can be used to characterize the probability of the phoneme sequence of the audio to be recognized being used in different languages, that is, it can identify the probability that the audio to be recognized belongs to different languages. For example, in the example above, the phoneme sequence "ni" is more likely to be Mandarin, the phoneme sequence "li" is more likely to be Sichuan dialect, and the phoneme sequence "halou" is more likely to be English, thus obtaining the language recognition result of the audio to be recognized.

[0051] It should be noted that in this embodiment, the high-dimensional acoustic features are converted into phoneme sequences through a phoneme conversion layer for subsequent language identification, rather than using words directly. That is, text information is not used for subsequent language identification. It is understandable that different languages, such as Chinese and English, are not interchangeable. Therefore, language identification through text results in a coarser overall modeling granularity, which is not conducive to language identification. For example, if the phoneme conversion layer is not used to convert the features into phoneme sequences, and instead text information is used for language identification, then for each minor language, there must be a sufficient number of training samples corresponding to that minor language in order to use the text information for subsequent minor language identification. The collection and annotation of training samples for minor languages ​​is relatively difficult, thus requiring a large amount of computing resources and resulting in high costs.

[0052] However, regardless of the language, even less commonly spoken languages ​​such as Cambodian, Vietnamese, Korean, Burmese, Urdu, Lao, Arabic, Persian, and Hungarian, phonemes are universal. This means that regardless of the language, the audio to be recognized can be converted into a phoneme sequence after extracting high-dimensional acoustic features using an acoustic model. Therefore, even if training sample resources are limited for some less commonly spoken languages, training samples from major languages, such as Chinese and English, can be used to train the phoneme conversion layer, making the resulting layer applicable to less commonly spoken languages ​​as well. For example, using Pinyin as the phoneme, the phoneme conversion layer can be trained with Chinese training samples. When the phoneme conversion layer receives audio from other languages, it can convert the audio from those languages ​​into a combination of Pinyin, i.e., into a phoneme sequence, thus achieving the sharing of phoneme conversion layers between different languages. Furthermore, the language feature extraction layer and language recognition layer can first be trained using training samples from major languages, such as Chinese and English, to enable them to be applicable to major languages. When language recognition is required for audio in a minor language, the training samples from the minor language can be input into the language feature extraction layer and language recognition layer trained with major languages ​​for small-scale training, achieving the same language recognition effect as for major languages. For example, in this embodiment, the language feature extraction layer and language recognition layer trained with major languages ​​have essentially trained about 90% of their parameters. For minor languages, only the final 10% or so of training is needed, with adjustments made to the relevant parameters, without requiring separate large-scale training for each minor language.

[0053] Therefore, this application embodiment does not require training a separate language recognition model for each language. One language recognition model can be used to recognize multiple languages. When recognizing different languages, there is no need to switch the corresponding language recognition model. Especially for less common languages, there is no need to train a separate language recognition model using training samples for the less common language. It can still achieve a high language recognition accuracy by using fewer training sample resources. It is understood that the obtained language recognition results can be used as language / ethnic / regional information for downstream natural language processing tasks to assist in making responses that are more suitable for user groups.

[0054] Please see Figure 3 , Figure 3 This is a schematic diagram of the second language recognition method provided in the embodiments of this application.

[0055] In step 201, obtain the audio to be identified that requires language recognition.

[0056] It is understandable that users may come from different regions of the world and use different languages. Therefore, the audio to be identified obtained in this embodiment can be a segment of audio in a specific language, such as Chinese or a foreign language, such as English, French, or Japanese; it can also be Mandarin or a dialect, such as Sichuanese, Minnan, Northeastern Mandarin, or Cantonese. The audio to be identified can be collected through devices such as a microphone; this application does not specifically limit the method of acquiring the audio to be identified.

[0057] In step 202, the audio to be recognized is input into the acoustic model of the speech recognition model, and the high-dimensional acoustic features of the audio to be recognized are extracted through the acoustic model.

[0058] Please refer to Figure 4 , Figure 4 This is a second structural block diagram of the language recognition model provided in the embodiments of this application. The embodiments of this application can obtain the acoustic model of the speech recognition model. It can also be understood that the embodiments of this application can reuse the acoustic model in the speech recognition model. That is, the language recognition method of the embodiments of this application performs language recognition on the audio to be recognized based on the language recognition model, and the language recognition model is built on the acoustic model in the speech recognition model. The speech recognition model can be, for example, a Recurrent Neural Network Transducer (RNN-T).

[0059] Acoustic models can extract features from one-dimensional audio samples to obtain high-dimensional acoustic features, such as two-dimensional high-dimensional acoustic features. The one-dimensional audio sample may be the audio sample in the time dimension, while the two-dimensional high-dimensional acoustic features may include both time and frequency dimensions. These high-dimensional acoustic features may be short-time Fourier transform features, thereby obtaining the frequency components of the audio sample in different time frames and obtaining the high-dimensional acoustic features of the audio sample in the time-frequency (T*F) two-dimensional plane.

[0060] In the related art, when a speech recognition model is used for speech recognition, if the audio to be recognized input by the user may be in multiple languages, a language recognition model is usually included as a sub-module in the speech recognition model to perform language recognition on the audio to be recognized. Wherein, the construction of language recognition and the speech recognition model are independent, that is, the speech recognition model has an acoustic model, while the language recognition model may need to separately set a layer with similar functions to the acoustic model of the speech recognition model and perform corresponding training, thereby resulting in waste of computing resources and cost. In the embodiments of the present application, by reusing the acoustic model in the speech recognition model, high-dimensional acoustic features of the audio to be recognized are extracted, and the extracted high-dimensional acoustic features can be used in both the speech recognition model and the language recognition model, that is, the high-dimensional acoustic features extracted by the acoustic model in the speech recognition model can be used as the input of the language recognition model for subsequent language recognition. It is not necessary to separately set a layer with similar functions to the acoustic model of the speech recognition model to extract high-dimensional acoustic features, nor is it necessary to train the layer, thereby reducing the occupation of computing resources, avoiding waste of cost, and reserving more resources for training other layers in the language recognition model, so as to improve the performance of the language recognition model and the accuracy of language recognition.

[0061] In step 203, the high-dimensional acoustic features are input into a phoneme conversion layer of a language recognition model, and the high-dimensional acoustic features are converted into a phoneme sequence of the audio to be recognized through the phoneme conversion layer.

[0062] In the embodiments of the present application, the high-dimensional acoustic features are input into a phoneme conversion layer of a language recognition model, and the high-dimensional acoustic features are converted into a phoneme sequence of the audio to be recognized through the phoneme conversion layer. It can be understood that a phoneme is the smallest speech unit divided according to the natural attributes of speech. From the perspective of acoustic properties, a phoneme is the smallest speech unit divided from the perspective of sound quality. One pronunciation action forms one phoneme. For example, phonemes can be International Phonetic Alphabet, Roman phonetic notation, pinyin, etc. A phoneme sequence can be understood as being formed by arranging phonemes in the chronological order of the audio to be recognized, and the phoneme sequence takes phonemes as units. For example, taking the Chinese word "ni hao" meaning hello as an example, the corresponding phoneme sequence is "ni", "hao". It can be understood that in the embodiments of the present application, by arranging the phoneme conversion layer, the audio to be recognized "ni hao" is converted into the phoneme sequence "ni", "hao", instead of the characters "ni", "hao".

[0063] In the related art, an acoustic model performs feature extraction on audio to be recognized, and the obtained high-dimensional acoustic features can be used to convert the substantial content information of the audio to be recognized. For example, the audio to be recognized "ni hao" (hello) can be converted into "ni" and "hao", that is, the high-dimensional acoustic features extracted by the acoustic model only retain common content information and delete some other features, such as accent, tone and other features. It can be understood that for users in different regions, even if the content information in the audio to be recognized is the same, there will still be different accents, tones and other features. In the embodiment of the present application, for example, taking the audio to be recognized "ni hao" (hello) as an example, if it is Mandarin, "ni hao" is converted into a phoneme sequence "ni", "hao"; if it is Sichuan dialect, due to different pronunciation habits, the phoneme sequence converted from "ni hao" can be "li", "hao". Therefore, different phoneme sequences can be used to distinguish accent information of different languages, and the embodiment of the present application outputs a more fine-grained phoneme sequence through the phoneme conversion layer, which is more conducive to language recognition.

[0064] At 204, the phoneme sequence is input into the language feature extraction layer of the language recognition model, and the first language feature of the audio to be recognized is extracted through the language feature extraction layer.

[0065] Further, in the embodiment of the present application, the phoneme sequence is input into the language feature extraction layer, and the language feature extraction layer extracts the first language feature of the audio to be recognized. It can be understood that, for different languages, phonemes can be universal, that is, phonemes can be used to label audio to be recognized in different languages, and the first language feature can be understood as an association feature between the phoneme sequence and the language. For example, taking phonemes as pinyin, when "ni" is input into the language feature extraction layer, the language feature extraction layer extracts that "n" and "i" are often combined to form "ni" in Mandarin; when "li" is input into the language feature extraction layer, the language feature extraction layer can extract that "l" and "i" are often combined to form "li" in Sichuan dialect. It can also be understood that Mandarin and Sichuan dialect may have different phoneme combinations for the same word due to different accents. The first language feature obtained by the language feature extraction layer can extract the first language feature related to accents of different languages, that is, the first language feature obtained by the language feature extraction layer can reflect pronunciation habits of different languages to a certain extent.

[0066] For example, continuing with the phoneme-based pinyin example, if the input audio to be recognized is "hello", then the phoneme sequence obtained by the phoneme conversion layer is, for example, "halou". This "halou" is then input into the language feature extraction layer. The language feature extraction layer extracts the phoneme sequence because "h", "a", "l", "o", and "u" often combine to form "halou" in English. This can be understood as the difference in phoneme sequences between Chinese and English being due to different linguistic systems. Different languages ​​correspond to certain specific or commonly used phoneme sequences, or in other words, a certain phoneme sequence has a higher probability of appearing in a certain language. Therefore, through the correlation obtained by the language feature extraction layer, we can extract the first language features related to different languages ​​and linguistic aspects such as semantics. Regardless of whether the difference in phoneme sequences is due to pronunciation habits or semantics, different languages ​​correspond to certain commonly used or commonly used phoneme sequences. Therefore, the extracted first language features can be used to characterize the probability of use and occurrence of phoneme sequences in different languages.

[0067] In step 205, the first language features are input into the feature pooling layer of the language recognition model to perform temporal pooling on the first language features, thereby reducing the dimensionality of the first language features to obtain one-dimensional second language features.

[0068] In this embodiment, features are extracted from one-dimensional audio to be identified using an acoustic model to obtain high-dimensional acoustic features, such as two-dimensional high-dimensional acoustic features. For the phoneme conversion layer, the input is the two-dimensional high-dimensional acoustic features, and the output is also a two-dimensional phoneme sequence. For the language feature extraction layer, the input is a two-dimensional phoneme sequence, and the output is also a two-dimensional first language feature. In this embodiment, the two-dimensional first language feature is input to the feature pooling layer to perform temporal pooling on the first language feature, reducing its dimensionality to obtain a one-dimensional second language feature. It can be understood that the feature pooling layer can convert first language features (T*F) of different durations and frequencies into one-dimensional second language features (1*F). Therefore, regardless of the duration of the audio to be identified, such as 10 seconds, 20 seconds, or 40 seconds, the second language feature, as the input parameter for language identification, is unrelated to the duration, but only related to the second language feature of a fixed duration. This facilitates subsequent unified language identification for audio of different durations in real-world scenarios. The pooling layer may include at least one of the following: global average pooling layer, maximum pooling layer, and minimum pooling layer.

[0069] In step 206, the second language feature is input into the language recognition layer, and language recognition is performed through the language recognition layer. The language recognition layer outputs the language distribution probability corresponding to the audio to be recognized, and generates the language recognition result based on the language distribution probability.

[0070] The phoneme sequence is processed by a language feature extraction layer to extract the first language feature, and then the first language feature is reduced in dimensionality by a feature pooling layer to obtain a one-dimensional second language feature. The second language feature can be used to reflect the probability that the phoneme sequence belongs to different languages. Moreover, the second language feature is used as the input of the language recognition layer and is not related to different durations. This makes it convenient for the language recognition layer to perform unified language recognition work on audio to be recognized of different durations in real-world scenarios. It can avoid the problem that the language recognition layer will reduce the accuracy of language recognition due to the variable duration of the input parameter, thereby improving the reliability and accuracy of language recognition.

[0071] One-dimensional second language features are input into the language recognition layer. The language recognition layer outputs the language distribution probability corresponding to the audio to be recognized and generates the language recognition result based on the language distribution probability. For example, regardless of the number of second language features in the audio to be recognized, the language recognition layer can output feature values ​​corresponding to the number of languages ​​the language recognition model can recognize. These feature values ​​can be language distribution probabilities. If the current language recognition method can be used to recognize three languages: Chinese, English, and French, then the language recognition layer can output three feature values. The distribution of these three feature values ​​indicates the distribution probability of the audio to be recognized for each of these three languages. The higher the probability value, the greater the probability that the audio to be recognized belongs to that language; the lower the probability value, the less likely the audio to be recognized belongs to that language. Thus, the language recognition result can be determined based on the language distribution probability. It should be noted that the feature values ​​can also be other values, and other values ​​can also be used to characterize the language distribution probability.

[0072] The language identification method in this application embodiment is based on a language identification model to identify the language of the audio to be identified. The language identification model includes a phoneme conversion layer, a language feature extraction layer, a feature pooling layer, and a language identification layer. To ensure accurate language identification, the phoneme conversion layer, language feature extraction layer, feature pooling layer, and language identification layer in the language identification model can be trained using deep learning based on the basic phoneme conversion layer, basic language feature extraction layer, basic feature pooling layer, and basic language identification layer. Please refer to [reference needed]. Figure 5 , Figure 5 This is a schematic diagram of the training process of the language recognition model provided in the embodiments of this application.

[0073] In step 301, a first training sample is obtained, which includes a first sample audio to be identified and a first sample phoneme sequence corresponding to the first sample audio to be identified.

[0074] In step 302, the basic phoneme conversion layer is trained based on the first sample audio to be identified and the first sample phoneme sequence to obtain the phoneme conversion layer.

[0075] The acoustic model in this embodiment can reuse the acoustic model from the speech recognition model. When training the language recognition model, this embodiment does not require training the acoustic model separately; instead, it can directly utilize the acoustic model from the speech recognition model without affecting its performance. The language recognition model in this embodiment includes a phoneme conversion layer, used to convert the high-dimensional acoustic features extracted by the acoustic model into a phoneme sequence of the audio to be recognized. To ensure accurate conversion, the phoneme conversion layer of the language recognition model can be trained. Please refer to... Figure 6 , Figure 6 This diagram illustrates the training of the phoneme conversion layer in the language recognition model provided in this embodiment. Specifically, a first training sample is obtained, comprising a first sample audio to be recognized and a corresponding first sample phoneme sequence. The first sample audio is input into an acoustic model, which outputs a first high-dimensional acoustic feature matching the first sample audio. This first high-dimensional acoustic feature is then input into a basic phoneme conversion layer, which converts the first high-dimensional acoustic feature into a first phoneme sequence of the first sample audio. The basic phoneme conversion layer can, for example, employ a single-layer temporal model structure, such as a recurrent neural network (RNN) or a self-attention mechanism layer.

[0076] Based on the difference information between the first phoneme sequence and the first sample phoneme sequence, a first loss value is obtained. This first loss value is then used to train the basic phoneme conversion layer, resulting in a phoneme conversion layer. Specifically, the first phoneme sequence and the first sample phoneme sequence can be input into the first loss function, which determines the first loss value. This value is then backpropagated to the basic phoneme conversion layer to adjust its parameters, resulting in the trained phoneme conversion layer. The first loss function can be, for example, the Connectionist Temporal Classification (CTC) loss function.

[0077] This application embodiment can convert high-dimensional acoustic features into phoneme sequences through a trained phoneme conversion layer for subsequent language recognition. Regardless of the language, even for audio in less commonly spoken languages, high-dimensional acoustic features can be extracted using an acoustic model and then accurately converted into phoneme sequences by the trained phoneme conversion layer. Since training sample resources are limited for some less commonly spoken languages, training samples from major languages, such as Chinese and English, can be used to train the phoneme conversion layer. This allows the trained phoneme conversion layer to be applicable to less commonly spoken languages, achieving the sharing of phoneme conversion layers between different languages.

[0078] In step 303, a second training sample is obtained. The second training sample includes the high-dimensional acoustic features of the first sample of the audio to be identified and the language recognition result of the first sample.

[0079] In step 304, the phoneme conversion layer is frozen. Based on the high-dimensional acoustic features of the first sample and the language recognition result of the first sample, the basic language feature extraction layer, the basic feature pooling layer, and the basic language recognition layer are trained to obtain the language feature extraction layer, the feature pooling layer, and the language recognition layer.

[0080] The language recognition model comprises four neural network layers: a phoneme conversion layer, a language feature extraction layer, a feature pooling layer, and a language recognition layer. The trained phoneme conversion layer can already output relatively accurate phoneme sequences based on the high-dimensional acoustic features output by the acoustic model; therefore, the phoneme conversion layer can be frozen, meaning its parameters are not trained. To ensure the language recognition model accurately identifies the language, the language feature extraction layer, feature pooling layer, and language recognition layer can also be trained. Please refer to [reference needed]. Figure 7 , Figure 7 This is a training diagram of the language feature extraction layer, feature pooling layer, and language recognition layer of the language recognition model provided in the embodiments of this application.

[0081] Specifically, a second training sample is obtained. The second training sample includes the high-dimensional acoustic features of the first sample of the audio to be identified and the language recognition result of the first sample. For example, the second sample of the audio to be identified can be obtained first; the second sample of the audio to be identified is input into a multiplexed acoustic model, and the acoustic model outputs the high-dimensional acoustic features of the first sample of the audio to be identified. Then, the high-dimensional acoustic features of the first sample of the audio to be identified and the language recognition result of the first sample are used as the second training sample. It should be clarified that the acoustic model does not need to be trained. In this embodiment, the high-dimensional acoustic features of the first sample output by the acoustic model can be used to train the language recognition model.

[0082] After obtaining the second training sample, the high-dimensional acoustic features of the first sample are input into the phoneme conversion layer, which converts the high-dimensional acoustic features of the first sample into the second phoneme sequence of the audio to be identified in the second sample. The second phoneme sequence is input into the basic language feature extraction layer, which extracts the third language features of the audio to be identified in the second sample. The third language features are input into the basic feature pooling layer, which reduces the dimensionality of the third language features to obtain a one-dimensional fourth language feature. The fourth language features are input into the basic language recognition layer, which performs language recognition to obtain the first predicted language recognition result of the audio to be identified in the second sample. Based on the difference between the first predicted language recognition result and the first sample language recognition result, a second loss value is obtained. The basic language feature extraction layer, the basic feature pooling layer, and the basic language recognition layer are trained based on the second loss value to obtain the language feature extraction layer, the feature pooling layer, and the language recognition layer. The process involves inputting the first predicted language identification result and the first sample language identification result into a second loss function. This second loss function determines the second loss value, which is then backpropagated to the basic language feature extraction layer, the basic feature pooling layer, and the basic language identification layer to adjust their parameters, resulting in the trained language feature extraction layer, feature pooling layer, and language identification layer. The second loss function can be, for example, the cross-entropy (CE) loss function. The basic language feature extraction layer and the basic language identification layer can be constructed using temporal model structures, such as recurrent neural networks (RNNs) or self-attention mechanisms.

[0083] In step 305, a third training sample is obtained. The third training sample includes the high-dimensional acoustic features of the second sample of the audio to be identified and the language recognition result of the second sample.

[0084] In 306, based on the high-dimensional acoustic features of the second sample and the language recognition results of the second sample, the phoneme conversion layer, language feature extraction layer, feature pooling layer and language recognition layer are trained to obtain a new phoneme conversion layer, a new language feature extraction layer, a new feature pooling layer and a new language recognition layer.

[0085] Furthermore, after training the phoneme conversion layer, language feature extraction layer, feature pooling layer, and language recognition layer separately, this embodiment of the application can further retrain the trained phoneme conversion layer, language feature extraction layer, feature pooling layer, and language recognition layer as a whole model. Please refer to [link / reference]. Figure 8 , Figure 8This diagram illustrates the training of the phoneme conversion layer, language feature extraction layer, feature pooling layer, and language recognition layer of the language recognition model provided in this embodiment. Specifically, a third training sample is obtained. The third training sample includes the high-dimensional acoustic features of the second sample of the audio to be recognized and the language recognition result of the second sample. For example, the third sample of the audio to be recognized can be obtained first; the third sample of the audio to be recognized can be input into the acoustic model, the acoustic model can output the high-dimensional acoustic features of the second sample of the audio to be recognized, and then the high-dimensional acoustic features of the second sample of the audio to be recognized and the language recognition result of the second sample can be used as the third training sample.

[0086] After obtaining the third training sample, the high-dimensional acoustic features of the second sample are input into the phoneme conversion layer, which converts the high-dimensional acoustic features of the second sample into the third phoneme sequence of the audio to be identified in the third sample. The third phoneme sequence is input into the language feature extraction layer, which extracts the fifth language feature of the audio to be identified in the third sample. The fifth language feature is input into the feature pooling layer, which reduces the dimensionality of the fifth language feature to obtain a one-dimensional sixth language feature. The sixth language feature is input into the language recognition layer, which performs language recognition to obtain the second predicted language recognition result of the audio to be identified in the third sample. Based on the difference between the second predicted language recognition result and the second sample language recognition result, a third loss value is obtained. The phoneme conversion layer, language feature extraction layer, feature pooling layer, and language recognition layer are trained based on the third loss value to obtain a new phoneme conversion layer, a new language feature extraction layer, a new feature pooling layer, and a new language recognition layer. Specifically, the second predicted language identification result and the second sample language identification result can be input into a third loss function. The third loss function determines the third loss value, and the gradient is backpropagated to the phoneme conversion layer, language feature extraction layer, feature pooling layer, and language identification layer to adjust the parameters of the phoneme conversion layer, language feature extraction layer, feature pooling layer, and language identification layer as a whole, resulting in a new phoneme conversion layer, a new language feature extraction layer, a new feature pooling layer, and a new language identification layer after training. The third loss function can be, for example, the cross-entropy (CE) loss function. The first, second, and third loss functions can be selected as needed, and can be the same or different; this embodiment does not limit this.

[0087] It should be noted that this application does not limit the number of times the language recognition model is trained separately or as a whole. This application can perform iterative training multiple times until a preset stopping condition is met to obtain the language recognition model of this application. It is understood that after separate training and overall training, the language recognition model can replace the phoneme conversion layer, the language feature extraction layer, the feature pooling layer, and the language recognition layer with a new one. Furthermore, this application can perform only the separate training or overall training on each level of the language recognition model as described above, and this application does not limit this approach.

[0088] It should be noted that the language recognition model in this application embodiment can reuse the acoustic model in the speech recognition model, and take the high-dimensional acoustic features output by the acoustic model as input. It does not need to set up a separate layer with similar function to the acoustic model of the speech recognition model to extract high-dimensional acoustic features, nor does it need to train the layer. Moreover, the feature pooling layer of the language recognition model is used to reduce the dimensionality of the first language features output by the language feature extraction layer. It may have no computational parameters or only a few computational parameters. The language recognition model mainly includes the training of three neural network layers during the training process. Therefore, the training process of the language recognition model occupies less computing resources.

[0089] Even if training sample resources for some less common languages ​​are limited in this embodiment, the phoneme conversion layer can still be trained using training samples from major languages, such as Chinese and English, so that the trained phoneme conversion layer can also be applied to less common languages. For example, taking Pinyin as the phoneme, the phoneme conversion layer can be trained using Chinese training samples. When the phoneme conversion layer receives audio to be recognized in other languages, it can also convert the audio to be recognized in other languages ​​into a combination of Pinyin, that is, into a phoneme sequence, realizing the sharing of phoneme conversion layers between different languages. In addition, the language feature extraction layer and the language recognition layer can also be trained using training samples from major languages, such as Chinese and English, so that the language feature extraction layer and the language recognition layer can be used for major languages. When it is necessary to perform language recognition on audio to be recognized in a less common language, the training samples of the less common language can be input into the language feature extraction layer and the language recognition layer trained in the major language for training with a small amount of data, so as to obtain the same language recognition effect as the major language. For example, the language feature extraction and language recognition layers trained on major languages ​​have essentially trained about 90% of their parameters. For minor languages, only the final 10% of the training is needed, with adjustments made to the relevant parameters. This eliminates the need for separate large-scale training for each minor language. Therefore, this embodiment does not require training a separate language recognition model for each language. One language recognition model can be used to recognize multiple languages. When recognizing different languages, there is no need to switch to a different language recognition model. Especially for minor languages, a separate language recognition model is unnecessary, allowing for the use of fewer training samples while still achieving high language recognition accuracy through a shared model. This reduces computational resources and the cost of the language recognition model. The obtained language recognition results can be used as language / ethnic / regional information for downstream natural language processing tasks, assisting in providing responses more suitable for the target user group.

[0090] In one embodiment, the language recognition model in this application includes four neural network layers: a phoneme extraction layer, a language feature extraction layer, a feature pooling layer, and a language recognition layer. However, it can be configured with more than four neural network layers as needed, such as multiple phoneme extraction layers, and / or multiple language feature extraction layers, and / or multiple feature pooling layers, and / or multiple language recognition layers, etc., to achieve better language recognition results. Furthermore, the language recognition model may also include other layers, such as nonlinear activation functions and normalization layers, etc., but this application does not limit this aspect.

[0091] Please see Figure 9 and Figure 10 , Figure 9 This is a schematic diagram of a first structure of the language recognition device provided in the embodiments of this application. Figure 10 This is a second structural schematic diagram of the language recognition device provided in an embodiment of this application. The language recognition device 400 may include an acquisition module 401, a first extraction module 402, a conversion module 403, a second extraction module 404, and a recognition module 405.

[0092] The acquisition module 401 is used to: acquire the audio to be identified that needs to be language-identified;

[0093] The first extraction module 402 is used to: input the audio to be recognized into the acoustic model of the speech recognition model, and extract the high-dimensional acoustic features of the audio to be recognized through the acoustic model;

[0094] The conversion module 403 is used to: input high-dimensional acoustic features into the phoneme conversion layer of the language recognition model, and convert the high-dimensional acoustic features into a phoneme sequence of the audio to be recognized through the phoneme conversion layer;

[0095] The second extraction module 404 is used to: input the phoneme sequence into the language feature extraction layer of the language recognition model, and extract the first language feature of the audio to be recognized through the language feature extraction layer;

[0096] The recognition module 405 is used to: input the first language feature into the language recognition layer of the language recognition model, perform language recognition through the language recognition layer, and obtain the language recognition result of the audio to be recognized.

[0097] In one embodiment, the language recognition device further includes a pooling module 406, which can be used to: input the first language feature into the feature pooling layer to perform temporal pooling on the first language feature, and reduce the dimensionality of the first language feature to obtain a one-dimensional second language feature; the recognition module 405 can also be used to: input the second language feature into the language recognition layer of the language recognition model, and perform language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized.

[0098] In one embodiment, the recognition module 405 can also be used to: input the second language feature into the language recognition layer, perform language recognition through the language recognition layer, output the language distribution probability corresponding to the audio to be recognized by the language recognition layer, and generate the language recognition result based on the language distribution probability.

[0099] In one embodiment, the language recognition device 400 further includes a training module 407, which can be used to: acquire a first training sample, the first training sample including a first sample audio to be recognized and a first sample phoneme sequence corresponding to the first sample audio to be recognized; and train the basic phoneme conversion layer according to the first sample audio to be recognized and the first sample phoneme sequence to obtain the phoneme conversion layer.

[0100] In one embodiment, the training module 407 can be used to: acquire a second training sample, wherein the second training sample includes the high-dimensional acoustic features of the first sample of the audio to be identified and the language recognition result of the first sample; freeze the phoneme conversion layer; and train the basic language feature extraction layer, the basic feature pooling layer and the basic language recognition layer according to the high-dimensional acoustic features of the first sample and the language recognition result of the first sample to obtain the language feature extraction layer, the feature pooling layer and the language recognition layer.

[0101] In one implementation, the training module 407 can be used to: acquire a third training sample, the third training sample including the high-dimensional acoustic features of the second sample of the audio to be identified and the language recognition result of the second sample; unfreeze the phoneme conversion layer, and train the phoneme conversion layer, the language feature extraction layer, the feature pooling layer and the language recognition layer according to the high-dimensional acoustic features of the second sample and the language recognition result of the second sample, so as to obtain a new phoneme conversion layer, a new language feature extraction layer, a new feature pooling layer and a new language recognition layer.

[0102] In one implementation, the training module 407 can be used to: acquire a third sample audio to be identified; input the third sample audio to be identified into an acoustic model, and the acoustic model outputs the second sample high-dimensional acoustic features of the third sample audio to be identified.

[0103] This application provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed on a computer, it causes the computer to perform the process in the language recognition method provided in this embodiment.

[0104] This application also provides an electronic device, including a memory and a processor, wherein the processor executes the process in the language recognition method provided in this embodiment by calling a computer program stored in the memory.

[0105] For example, the aforementioned electronic device could be a mobile terminal such as a tablet or smartphone. See also... Figure 11 , Figure 11 This is a schematic diagram of a first structure of an electronic device provided in an embodiment of this application.

[0106] The electronic device 500 may include components such as a memory 502 and a processor 501. Those skilled in the art will understand that... Figure 11 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0107] Memory 502 can be used to store applications and data. The applications stored in memory 502 contain executable code. Applications can be composed of various functional modules. Processor 501 executes various functional applications and data processing by running the applications stored in memory 502. Processor 501 is electrically connected to memory 502.

[0108] The processor 501 is the control center of the electronic device. It connects various parts of the electronic device through various interfaces and lines. By running or executing the application program stored in the memory 502 and calling the data stored in the memory 502, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0109] In this embodiment, the processor 501 in the electronic device loads the executable code corresponding to the process of one or more applications into the memory 502 according to the following instructions, and the processor 501 runs the applications stored in the memory 502 to perform the following: acquiring the audio to be recognized for language identification; inputting the audio to be recognized into the acoustic model of the speech recognition model, and extracting high-dimensional acoustic features of the audio to be recognized through the acoustic model; inputting the high-dimensional acoustic features into the phoneme conversion layer of the language recognition model, and converting the high-dimensional acoustic features into a phoneme sequence of the audio to be recognized through the phoneme conversion layer; inputting the phoneme sequence into the language feature extraction layer of the language recognition model, and extracting the first language feature of the audio to be recognized through the language feature extraction layer; inputting the first language feature into the language recognition layer of the language recognition model, and performing language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized.

[0110] Please see Figure 12 , Figure 12 This is a second structural schematic diagram of the electronic device provided in an embodiment of this application. The electronic device 500 also includes components such as a display screen 503, a radio frequency circuit 504, a control circuit 505, an input unit 506, an audio circuit 507, a sensor 508, and a power supply 509. The processor 501 is electrically connected to the display screen 503, the radio frequency circuit 504, the control circuit 505, the input unit 506, the audio circuit 507, the sensor 508, and the power supply 509.

[0111] The display screen 503 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic devices, which can be composed of images, text, icons, videos, and any combination thereof.

[0112] The radio frequency circuit 504 is used to transmit and receive radio frequency signals to communicate with network devices or other electronic devices via wireless communication.

[0113] The control circuit 505 is electrically connected to the display screen 503 and is used to control the display screen 501 to display information.

[0114] The input unit 506 can be used to receive input numeric, character information, or user characteristic information (such as fingerprints), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. The input unit 506 may include a fingerprint recognition module.

[0115] Audio circuit 507 provides an audio interface between the user and electronic device via a speaker and microphone. Audio circuit 507 includes a microphone. The microphone is electrically connected to processor 501. The microphone is used to receive voice information input by the user.

[0116] Sensor 508 is used to collect information about the external environment. Sensor 508 may include one or more sensors such as an ambient light sensor, an accelerometer, and a gyroscope.

[0117] The power supply 509 is used to supply power to the various components of the electronic device 500. In some embodiments, the power supply 509 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system.

[0118] Although not shown in the figure, the electronic device 500 may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0119] In this embodiment, the processor 501 in the electronic device 500 loads the executable code corresponding to the process of one or more applications into the memory 502 according to the following instructions, and the processor 501 runs the applications stored in the memory 502 to perform the following: acquiring the audio to be recognized for language identification; inputting the audio to be recognized into the acoustic model of the speech recognition model, and extracting high-dimensional acoustic features of the audio to be recognized through the acoustic model; inputting the high-dimensional acoustic features into the phoneme conversion layer of the language recognition model, and converting the high-dimensional acoustic features into a phoneme sequence of the audio to be recognized through the phoneme conversion layer; inputting the phoneme sequence into the language feature extraction layer of the language recognition model, and extracting the first language feature of the audio to be recognized through the language feature extraction layer; inputting the first language feature into the language recognition layer of the language recognition model, and performing language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized.

[0120] In one implementation, in the process of inputting the first language features into the language recognition layer of the language recognition model and performing language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized, the processor 501 executes: inputting the first language features into the feature pooling layer to perform temporal pooling on the first language features, reducing the dimensionality of the first language features to obtain a one-dimensional second language feature; inputting the second language features into the language recognition layer of the language recognition model and performing language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized.

[0121] In one implementation, in the process of inputting the second language features into the language recognition layer of the language recognition model, performing language recognition through the language recognition layer, and obtaining the language recognition result of the audio to be recognized, the processor 501 executes: inputting the second language features into the language recognition layer, performing language recognition through the language recognition layer, outputting the language distribution probability corresponding to the audio to be recognized by the language recognition layer, and generating the language recognition result based on the language distribution probability.

[0122] In one implementation, before inputting the high-dimensional acoustic features into the phoneme conversion layer of the language recognition model, the processor 501 performs the following: acquiring a first training sample, the first training sample including a first sample audio to be recognized and a first sample phoneme sequence corresponding to the first sample audio to be recognized; training the basic phoneme conversion layer based on the first sample audio to be recognized and the first sample phoneme sequence to obtain the phoneme conversion layer.

[0123] In one implementation, after obtaining the phoneme conversion layer, the processor 501 performs the following: acquiring a second training sample, wherein the second training sample includes the high-dimensional acoustic features of the first sample of the audio to be identified and the language recognition result of the first sample; freezing the phoneme conversion layer; and training the basic language feature extraction layer, the basic feature pooling layer and the basic language recognition layer based on the high-dimensional acoustic features of the first sample and the language recognition result of the first sample to obtain the language feature extraction layer, the feature pooling layer and the language recognition layer.

[0124] In one implementation, after obtaining the language feature extraction layer, the feature pooling layer, and the language recognition layer, the processor 501 executes: acquiring a third training sample, the third training sample including the high-dimensional acoustic features of the second sample of the audio to be recognized and the language recognition result of the second sample; unfreezing the phoneme conversion layer, and training the phoneme conversion layer, the language feature extraction layer, the feature pooling layer, and the language recognition layer based on the high-dimensional acoustic features of the second sample and the language recognition result of the second sample, so as to obtain a new phoneme conversion layer, a new language feature extraction layer, a new feature pooling layer, and a new language recognition layer.

[0125] In one implementation, before acquiring the third training sample, the processor 501 performs the following: acquiring the third sample audio to be identified; inputting the third sample audio to be identified into the acoustic model, and the acoustic model outputting the second sample high-dimensional acoustic features of the third sample audio to be identified.

[0126] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not detailed in a particular embodiment, please refer to the detailed description of the language recognition method above; they will not be repeated here. The language recognition device provided in this application embodiment belongs to the same concept as the language recognition method in the above embodiments. Any method provided in the language recognition method embodiments can be run on the language recognition device; its specific implementation process is detailed in the language recognition method embodiments, and will not be repeated here.

[0127] It should be noted that, regarding the language recognition method of this application embodiments, those skilled in the art will understand that all or part of the process of implementing the language recognition method of this application embodiments can be accomplished by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium, such as a memory, and executed by at least one processor. During execution, it can include the process of the language recognition method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), etc.

[0128] For the language recognition device of this application embodiment, its functional modules can be integrated into a processing chip, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0129] The language recognition method, language recognition device, storage medium, and electronic device provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A language identification method, characterized in that, The method includes: A first training sample is obtained, which includes a first sample audio to be identified and a first sample phoneme sequence corresponding to the first sample audio to be identified; a basic phoneme conversion layer is trained based on the first sample audio to be identified and the first sample phoneme sequence to obtain the phoneme conversion layer, wherein the phoneme conversion layer trained with the first training sample of a major language can be applied to minor languages. A second training sample is obtained, which includes the high-dimensional acoustic features of the first sample of the audio to be identified and the language recognition result of the first sample; the phoneme conversion layer is frozen, and the basic language feature extraction layer, the basic feature pooling layer and the basic language recognition layer are trained according to the high-dimensional acoustic features of the first sample and the language recognition result of the first sample to obtain the language feature extraction layer, the feature pooling layer and the language recognition layer. In this process, a large amount of data is first trained using the second training sample of a major language, and then a small amount of data is trained using the second training sample of a minor language. Obtain the audio to be identified, which requires language recognition; The audio to be recognized is input into the acoustic model of the speech recognition model, and the high-dimensional acoustic features of the audio to be recognized are extracted through the acoustic model. The high-dimensional acoustic features are input into the phoneme conversion layer of the language recognition model. The phoneme conversion layer converts the high-dimensional acoustic features into a phoneme sequence of the audio to be recognized. The phoneme sequence is composed of phonemes common to different languages ​​arranged in the chronological order of the audio to be recognized. The phoneme sequence is input into the language feature extraction layer of the language recognition model, and the first language feature of the audio to be recognized is extracted through the language feature extraction layer. The first language feature is used to characterize the probability of use or occurrence of the phoneme sequence in different languages. The first language feature is input into the feature pooling layer of the language recognition model to perform temporal pooling on the first language feature and reduce the dimensionality of the first language feature to obtain a one-dimensional second language feature. The second language feature is input into the language recognition layer of the language recognition model, and language recognition is performed through the language recognition layer to obtain the language recognition result of the audio to be recognized.

2. The language identification method according to claim 1, characterized in that, The step of inputting the second language feature into the language recognition layer of the language recognition model, and performing language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized includes: The second language feature is input into the language recognition layer, and language recognition is performed through the language recognition layer. The language recognition layer outputs the language distribution probability corresponding to the audio to be recognized, and generates the language recognition result based on the language distribution probability.

3. The language identification method according to claim 1, characterized in that, After obtaining the language feature extraction layer, the feature pooling layer, and the language recognition layer, the method further includes: Obtain a third training sample, which includes the high-dimensional acoustic features of the second sample of the audio to be identified and the language recognition result of the second sample. Unfreeze the phoneme conversion layer, and train the phoneme conversion layer, the language feature extraction layer, the feature pooling layer, and the language recognition layer based on the high-dimensional acoustic features of the second sample and the language recognition result of the second sample to obtain a new phoneme conversion layer, a new language feature extraction layer, a new feature pooling layer, and a new language recognition layer.

4. The language identification method according to claim 3, characterized in that: Before obtaining the third training sample, the method further includes: Obtain the third sample audio to be identified; The third sample audio to be identified is input into the acoustic model, and the acoustic model outputs the second sample high-dimensional acoustic features of the third sample audio to be identified.

5. A language recognition device, characterized in that, The device includes: The training module is used to acquire a first training sample, which includes a first sample audio to be identified and a first sample phoneme sequence corresponding to the first sample audio to be identified; to train a basic phoneme conversion layer based on the first sample audio to be identified and the first sample phoneme sequence to obtain the phoneme conversion layer, wherein the phoneme conversion layer trained with the first training sample of a major language can be applied to minor languages; and to acquire a second training sample, which includes the first sample high-dimensional acoustic features of the second sample audio to be identified and the first sample language recognition result; to freeze the phoneme conversion layer, and to train a basic language feature extraction layer, a basic feature pooling layer, and a basic language recognition layer based on the first sample high-dimensional acoustic features and the first sample language recognition result to obtain the language feature extraction layer, the feature pooling layer, and the language recognition layer, wherein training is first performed with a large amount of data using the second training sample of a major language, and then training is performed with a small amount of data using the second training sample of a minor language; The acquisition module is used to acquire the audio to be identified for language recognition. The first extraction module is used to input the audio to be recognized into the acoustic model of the speech recognition model, and extract the high-dimensional acoustic features of the audio to be recognized through the acoustic model. The conversion module is used to input the high-dimensional acoustic features into the phoneme conversion layer of the language recognition model, and convert the high-dimensional acoustic features into a phoneme sequence of the audio to be recognized through the phoneme conversion layer. The phoneme sequence is composed of phonemes common to different languages ​​arranged in the time order of the audio to be recognized. The second extraction module is used to input the phoneme sequence into the language feature extraction layer of the language recognition model, and extract the first language feature of the audio to be recognized through the language feature extraction layer. The first language feature is used to characterize the probability of use or occurrence of the phoneme sequence in different languages. The recognition module is used to input the first language feature into the feature pooling layer of the language recognition model to perform temporal pooling on the first language feature and reduce the dimensionality of the first language feature to obtain a one-dimensional second language feature; input the second language feature into the language recognition layer of the language recognition model, and perform language recognition through the language recognition layer to obtain the language recognition result of the audio to be recognized.

6. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed on a computer, it causes the computer to perform the method as described in any one of claims 1 to 4.

7. An electronic device, comprising a memory and a processor, characterized in that, The processor executes the method as described in any one of claims 1 to 4 by invoking a computer program stored in the memory.

Citation Information

Patent Citations

  • Method and device for speech recognition in adaptive language, and equipment

    CN109817213A

  • Mixed language speech recognition method, device and system and storage medium

    CN115394287A