Language Classification Method, Apparatus, and Computer-Readable Storage Medium
By using the trained target acoustic model and target language classification model to process the spectrum characteristics of the audio, obtain the phoneme sequence and determine the language, the problem of low accuracy of traditional language classification schemes is solved, and higher language classification accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202210743472.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-28
AI Technical Summary
The traditional language classification scheme has low accuracy in language classification because the audio spectrum characteristics are disturbed by noise such as tone and accompaniment.
By obtaining the spectral characteristics of the audio to be classified, the trained target acoustic model is called to process it based on the phoneme dictionary to obtain the phoneme sequence, and then the trained target language classification model is called to process the phoneme sequence to determine the language.
This method directly uses phoneme sequences to classify languages by excluding unrelated information such as tone and accompaniment, which improves the accuracy of language classification and adapts to multiple languages, improving classification efficiency.
Smart Images

Figure CN115132170B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a language classification method, apparatus, and computer-readable storage medium. Background Art
[0002] The language classification technology refers to an artificial intelligence technology that determines the language to which a text belongs through audio. In the music field, the language classification technology can identify the language category of the lyrics through music, and this technology can be applied to music library management, song recommendation, etc. The identified language can provide conditions for subsequent operations such as judging the interests of listeners.
[0003] Traditional language classification schemes use an end-to-end model to classify and identify the language of audio. This end-to-end model receives the spectral features of the audio and directly outputs the language category of the audio text after processing it. However, due to factors such as the audio recording environment, the spectral features of the audio are often interfered by noises such as tones and accompaniments, resulting in low accuracy of language classification. Summary of the Invention
[0004] Embodiments of the present application provide a language classification method, apparatus, and computer-readable storage medium, which can improve the accuracy of language classification.
[0005] In a first aspect, embodiments of the present application provide a language classification method, and the method includes:
[0006] Obtain the spectral features of the audio to be classified;
[0007] Call the trained target acoustic model to process the spectral features to obtain the phoneme sequence of the audio to be classified; the trained target acoustic model is a neural network model trained based on a phoneme dictionary, and the phoneme dictionary is used to indicate the correspondence between characters and phonemes of different languages;
[0008] Call the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs; the trained target language classification model is trained by the phoneme sequences of multiple training audios, and each of the training audios has a labeled preset language label, and the trained target language classification model records the correspondence between the phoneme sequence of the audio and the language to which the audio belongs.
[0009] In a possible implementation manner, the trained target language classification model includes a feature extraction sub-model and a language determination sub-model, and calling the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs includes:
[0010] Call the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector, where the phoneme feature vector is composed of multiple phoneme features of the phoneme sequence, and the phoneme features have a corresponding relationship with the language;
[0011] Call the language determination sub-model to process the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language, and determine the language to which the audio to be classified belongs according to the probabilities of the audio to be classified belonging to each language.
[0012] In a possible implementation, the trained feature extraction sub-model includes an embedding layer, a self-attention layer, and a batch normalization layer;
[0013] The step of calling the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector includes:
[0014] Call the embedding layer to perform vector encoding on the phoneme sequence to obtain a phoneme embedding vector;
[0015] Call the self-attention layer to process the phoneme embedding vector based on the correlation between each vector component in the phoneme embedding vector to obtain an initial feature vector;
[0016] Call the batch normalization layer to normalize the initial feature vector to obtain a phoneme feature vector.
[0017] In a possible implementation, the step of calling the language determination sub-model to process the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language includes:
[0018] Call the mapping rule in the language determination sub-model to map the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language.
[0019] In a possible implementation, the method further includes:
[0020] Obtain a first training audio set, where the first training audio set includes at least one first training audio of at least one language;
[0021] Obtain the spectral features of each first training audio and the first character sequence corresponding to the first training audio;
[0022] Call an initial acoustic model to process the spectral features of each first training audio to obtain the predicted phoneme sequence of each first training audio;
[0023] According to the phoneme dictionary and the predicted phoneme sequence of each first training audio, obtain the second character sequence of each first training audio;
[0024] Train the parameters in the initial acoustic model based on the first character sequence and the second character sequence to obtain the trained target acoustic model.
[0025] In a possible implementation, the method further includes:
[0026] Obtain the phoneme sequence to be processed for each first training audio, where the phoneme sequence to be processed is obtained by processing the spectral features of each first training audio using the trained target acoustic model;
[0027] Obtain the preset language to which each first training audio belongs;
[0028] Call the initial language classification model to process the phoneme sequence to be processed for each first training audio to obtain the predicted language to which each first training audio belongs;
[0029] Train the parameters in the initial language classification model based on the preset language and the predicted language to which each first training audio belongs to obtain the trained target language classification model.
[0030] In a possible implementation, the method further includes:
[0031] Obtain a second training audio set, where the second training audio set includes at least one second training audio of at least one language;
[0032] Obtain the spectral features and the preset language to which each second training audio in the second training audio set belongs;
[0033] Call the target acoustic model and the target language classification model in sequence to process the spectral features of each second training audio to obtain the predicted language to which each second training audio belongs;
[0034] Update the target acoustic model and the target language classification model based on the predicted language and the preset language to which each second training audio belongs.
[0035] In a possible implementation, the preset language to which each second training audio belongs is included in the preset languages to which the first training audios in the first training audio set belong; the number of second training audios included in the second training audio set is less than or equal to the number of first training audios included in the first training audio set.
[0036] In a possible implementation, the method further includes:
[0037] Obtain a third training audio set, where the third training audio set includes at least one third training audio of at least one language; the preset language to which the third training audio belongs is different from the preset language to which the first training audio belongs;
[0038] Obtain the preset language to which each third training audio in the third training audio set belongs;
[0039] Call the target acoustic model and the target language classification model in sequence to process each third training audio, and obtain the predicted language of each third training audio; based on the preset language and the predicted language to which each third training audio belongs, update the target language classification model.
[0040] In a second aspect, an embodiment of the present application provides a language classification device, where the device includes:
[0041] An acquisition unit, configured to acquire the spectral features of the audio to be classified;
[0042] A processing unit, configured to call the trained target acoustic model to process the spectral features, and obtain the phoneme sequence of the audio to be classified; the trained target acoustic model is a neural network model trained based on a phoneme dictionary, and the phoneme dictionary is used to indicate the correspondence between characters and phonemes of different languages;
[0043] The processing unit is further configured to call the trained target language classification model to process the phoneme sequence, and obtain the language to which the audio to be classified belongs; the trained target language classification model is trained by the phoneme sequences of multiple training audios, and each of the training audios has a labeled preset language label, and the trained target language classification model records the correspondence between the phoneme sequence of the audio and the language to which the audio belongs.
[0044] In a possible implementation manner, the trained target language classification model includes a feature extraction sub-model and a language determination sub-model. When the processing unit is used to call the trained target language classification model to process the phoneme sequence and obtain the language to which the audio to be classified belongs, it specifically includes:
[0045] Call the feature extraction sub-model to process the phoneme sequence, and obtain a phoneme feature vector, where the phoneme feature vector is composed of multiple phoneme features of the phoneme sequence, and the phoneme feature has a correspondence with the language;
[0046] Call the language determination sub-model to process the phoneme feature vector, obtain the probabilities that the audio to be classified belongs to each language, and determine the language to which the audio to be classified belongs according to the probabilities that the audio to be classified belongs to each language.
[0047] In a possible implementation, the trained feature extraction sub-model includes an embedding layer, a self-attention layer, and a batch normalization layer;
[0048] When the processing unit is used to call the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector, it specifically includes:
[0049] Call the embedding layer to perform vector encoding on the phoneme sequence to obtain a phoneme embedding vector;
[0050] Call the self-attention layer to process the phoneme embedding vector based on the correlation between each vector component in the phoneme embedding vector to obtain an initial feature vector;
[0051] Call the batch normalization layer to normalize the initial feature vector to obtain a phoneme feature vector.
[0052] In a possible implementation, when the processing unit is used to call the language determination sub-model to process the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language, it specifically includes:
[0053] Call the mapping rule in the language determination sub-model to map the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language.
[0054] In a possible implementation, the acquisition unit is further used for:
[0055] Acquire a first training audio set, where the first training audio set includes at least one first training audio of at least one language; acquire the spectral features of each first training audio and the first character sequence corresponding to the first training audio;
[0056] The processing unit is further used for:
[0057] Call the initial acoustic model to process the spectral features of each first training audio to obtain the predicted phoneme sequence of each first training audio;
[0058] According to the phoneme dictionary and the predicted phoneme sequence of each first training audio, obtain the second character sequence of each first training audio;
[0059] Train the parameters in the initial acoustic model based on the first character sequence and the second character sequence to obtain the trained target acoustic model.
[0060] In a possible implementation, the acquisition unit is further used for:
[0061] Obtain the to-be-processed phoneme sequence of each first training audio, where the to-be-processed phoneme sequence is obtained by processing the spectral features of each first training audio by invoking the trained target acoustic model;
[0062] Obtain the preset language to which each first training audio belongs;
[0063] The processing unit is further configured to:
[0064] Invoke the initial language classification model to process the to-be-processed phoneme sequence of each first training audio, and obtain the predicted language to which each first training audio belongs;
[0065] Train the parameters in the initial language classification model based on the preset language and the predicted language to which each first training audio belongs, and obtain the trained target language classification model.
[0066] In a possible implementation manner, the obtaining unit is further configured to:
[0067] Obtain a second training audio set, where the second training audio set includes at least one second training audio of at least one language;
[0068] Obtain the spectral features and the preset language to which each second training audio in the second training audio set belongs;
[0069] The processing unit is further configured to:
[0070] Successively invoke the target acoustic model and the target language classification model to process the spectral features of each second training audio, and obtain the predicted language to which each second training audio belongs;
[0071] Update the target acoustic model and the target language classification model based on the predicted language and the preset language to which each second training audio belongs.
[0072] In a possible implementation manner, the preset language to which each second training audio belongs is included in the preset languages to which the first training audios in the first training audio set belong; the number of second training audios included in the second training audio set is less than or equal to the number of first training audios included in the first training audio set.
[0073] In a possible implementation manner, the obtaining unit is further configured to:
[0074] Obtain a third training audio set, where the third training audio set includes at least one third training audio of at least one language; the preset language to which the third training audio belongs is different from the preset language to which the first training audio belongs;
[0075] Obtain the preset language to which each third training audio in the third training audio set belongs;
[0076] The processing unit is further configured to:
[0077] Successively call the target acoustic model and the target language classification model to process each third training audio, and obtain the predicted language of each third training audio; update the target language classification model based on the preset language and the predicted language to which each third training audio belongs.
[0078] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor, a memory, and a network interface. The processor is connected to the memory and the network interface; the network interface is used to provide network communication functions, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to call the program code to implement the method in the first aspect and the possible implementation manners of the first aspect.
[0079] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the method in the first aspect and the possible implementation manners of the first aspect are implemented.
[0080] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, the method in the first aspect and the possible implementation manners of the first aspect are implemented.
[0081] The embodiment of the present application identifies the language category to which an audio belongs based on the phoneme sequence corresponding to the audio. Compared with the traditional language classification scheme, phonemes exclude the interference of irrelevant information such as tones and accompaniments in the audio and are directly associated with the speech pronunciation actions in the audio, and can be used as a powerful basis for language classification (the speech pronunciation actions of different languages are different), improving the accuracy of language classification. Moreover, the trained target acoustic model in the present application is obtained based on a phoneme dictionary, and the phoneme dictionary can adapt to multiple languages, so that the trained target acoustic model can process audio in different languages, improving the efficiency of language classification. Description of the Drawings
[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0083] Figure 1 It is a schematic diagram of the principle of a language classification method provided by an embodiment of the present application;
[0084] Figure 2 It is a schematic flowchart of a language classification method provided by an embodiment of the present application;
[0085] Figure 3 It is a schematic flowchart of a training method for a language classification related model provided by an embodiment of the present application;
[0086] Figure 4 It is a schematic flowchart of the training process of an initial acoustic model provided by an embodiment of the present application;
[0087] Figure 5 It is a schematic flowchart of the training process of an initial language classification model provided by an embodiment of the present application;
[0088] Figure 6 It is a schematic diagram before and after the update of a language determination sub-model provided by an embodiment of the present application;
[0089] Figure 7 It is a schematic structural diagram of a language classification device provided by an embodiment of the present application;
[0090] Figure 8 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0091] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0092] In the description, claims and drawings of the present application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0093] To better understand the embodiments of the present application, the following first introduces the professional terms and related fields in the embodiments of the present application:
[0094] Phone: A phone is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the pronunciation actions in a syllable, one action constitutes one phone. For example, the speech "ma" contains two phones, "m" and "a", and the two phones have their respective corresponding pronunciation actions. The sounds produced by the same pronunciation action are the same phone, and the sounds produced by different pronunciation actions are different phones. That is to say, a phone is a symbolic representation of a pronunciation action. For example, in "ma" and "mi", the two "m"s correspond to the same pronunciation action and are the same phone, while "a" and "i" correspond to different pronunciation actions and are different phones. Usually, in Chinese speech, initials and finals are used as the symbolic representation (phones) of pronunciation actions, and in English speech, English phonetic symbols are used as the symbolic representation (phones) of pronunciation actions, etc.; in the embodiments of this application, a unified symbolic representation (phone) is adopted to represent the pronunciation actions of different language speeches (such as the International Phonetic Alphabet).
[0095] The embodiments of this application relate to the field of artificial intelligence (AI) technology. Among them, AI uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, AI technology is a comprehensive technology in computer science. It mainly produces a new intelligent machine that can react in a way similar to human intelligence by understanding the essence of intelligence, enabling the intelligent machine to have multiple functions such as perception, reasoning, and decision-making. AI technology mainly includes several major directions such as computer vision technology (CV), speech technology, natural language processing technology, and machine learning (ML).
[0096] Among them, the key technologies of speech technology are automatic speech recognition technology (ASR), speech synthesis technology, and voiceprint recognition technology. Speech technology enables the computer to listen, see, speak, and feel, and is one of the development directions of future human-computer interaction.
[0097] Machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent. Its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as deep learning, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0098] Based on the speech technology and machine learning technology in AI technology, an embodiment of the present application provides a language classification method. As Figure 1 shown, it is a schematic diagram of the principle of a language classification method provided by an embodiment of the present application. Among them, the audio to be classified first determines the spectral features through the spectral feature extraction module, then calls the trained target acoustic model to process the spectral features to obtain a phoneme sequence, and finally calls the trained target language classification model to process the phoneme sequence to obtain the language category to which the audio to be classified belongs. In this process, the trained target language classification model mainly judges the language to which the audio to be classified belongs based on the information contained in the phoneme sequence (such as the arrangement and combination of each phoneme in the phoneme sequence, the semantic connection of the characters corresponding to each phoneme, etc.), without considering irrelevant information such as the accompaniment and tone of the audio to be classified, and the audio to be classified with a short duration can also obtain a phoneme sequence with a sufficient number of phonemes (for example, an audio of 2 to 3 seconds can also identify a sequence containing 10 phonemes) after being processed by the trained target acoustic model. Based on this, the language classification method proposed in the present application can effectively improve the accuracy of language classification and provide conditions for subsequent processing based on the language classification result.
[0099] It should be noted that the above-mentioned trained target acoustic model and trained target language classification model can be trained by methods such as deep learning, and the entire method process can be implemented by a language classification device. The language classification device can be a terminal device, a device in the terminal device, or a device that can be used in combination with the terminal device. Exemplarily, the terminal device can be a device with data processing functions and input / output functions, including but not limited to handheld devices with wireless communication functions (such as smart phones, tablets), computing devices (such as personal computers (PC)), intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.
[0100] Next, the language classification method, device, and computer-readable storage medium provided by the embodiments of the present application will be described in detail respectively. Figures 2 to 7
[0101] See Figure 2 , which is a schematic flowchart of a language classification method provided by an embodiment of the present application. The method includes steps S201 to S203 and can be executed by the above-mentioned language classification device. Among them:
[0102] S201. Obtain the spectral features of the audio to be classified.
[0103] In some possible implementation manners, before obtaining the spectral features of the audio to be classified, it is also necessary to first obtain the audio signal of the audio to be classified. Exemplarily, a signal processing algorithm is used to convert the audio signal in the time domain to the frequency domain to obtain a spectrogram, and finally the spectral features are obtained based on the spectrogram. Exemplarily, the signal processing algorithm may be Fourier transform, discrete Fourier transform, fast Fourier transform, etc. The spectral features include but are not limited to features such as zero-crossing rate and short-time energy. It should be noted that when this application is applied to specific products and technologies, data such as the audio to be classified in the embodiments of this application need to be obtained with the permission or consent of the object first, and the collection, use, and processing of these data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0104] Optionally, the spectrogram can be further processed to obtain a mel spectrogram (the frequency distribution in the mel spectrogram is more in line with the linear perception of the human ear for frequency), and the mel spectral features are determined based on the mel spectrogram. Exemplarily, the mel spectral features may be mel frequency cepstrum coefficients (MFCC).
[0105] S202. Invoke the trained target acoustic model to process the spectral features to obtain the phoneme sequence of the audio to be classified; the trained target acoustic model is a neural network model trained based on a phoneme dictionary, and the phoneme dictionary is used to indicate the corresponding relationship between characters in different languages and phonemes.
[0106] Among them, this step is used to obtain the phoneme sequence of the audio to be classified based on the trained target acoustic model, and the subsequent steps are all based on the phoneme sequence obtained in this step. Specifically, the trained target acoustic model is a neural network model that obtains the pronunciation representation (that is, the phoneme sequence) of the audio to be classified based on the spectral features. Here, the pronunciation representation (phoneme sequence) is a sequence combination formed by arranging at least one phoneme. The phonemes in at least one phoneme may be the same phoneme or different phonemes, and one phoneme is the smallest pronunciation action in the pronunciation representation. For example, if the audio to be classified contains 5 pronunciation actions connected in chronological order, and the phonemes corresponding to the 5 pronunciation actions are "l1", "l2", "l3", "l2", "l1", then the phoneme sequence of the audio to be classified can be expressed as "l1,l2,l3,l2,l1".
[0107] The above-mentioned trained target acoustic model is trained based on a phoneme dictionary and an initial acoustic model, and the training method can refer to the embodiments shown below Figure 3 Exemplarily, the trained target acoustic model may be a neural network model with a self-attention mechanism (CNN-TDNN-F-SA) constructed based on a convolutional neural network and a factored delay neural network, etc. This application does not limit this.
[0108] The above-mentioned phoneme dictionary is a pronunciation dictionary suitable for multiple languages, which contains the correspondence between characters and phonemes in multiple languages. In other words, the pronunciation of characters in multiple languages can be represented by phonemes in the phoneme dictionary. This representation method can make the phoneme sequence output by the trained target acoustic model have a standard and unified data format, so that the trained target language classification model can process any phoneme sequence and improve the accuracy of language classification. In addition, the use of a phoneme dictionary containing unified phonemes reduces the number of phonemes, which can effectively alleviate the situation of data sparsity (if different languages use different phoneme sets, the number of phonemes will be large and the combination will be very complicated).
[0109] For example, the phonemes in the phoneme dictionary can be represented by the International Phonetic Alphabet (IPA), which includes 107 individual symbols, 56 diacritical marks and suprasegmental components, each symbol is a phoneme, the English phonetic alphabet is part of the International Phonetic Alphabet, and each consonant and vowel in Chinese can be determined in the International Phonetic Alphabet with symbols having the same pronunciation, and so on, without limitation here.
[0110] For example, Table 1 is a schematic diagram of a phoneme dictionary provided in an embodiment of the present application, where the phonemes in the phoneme dictionary are the phonemes in the set L = {l1, l2, ...}. As shown in Table 1, for the Chinese character "今日", the English character "today", and the Japanese character "きょう", these three characters have the same semantic meaning (all of which mean today), but have different pronunciations. They can be represented by the phonemes in the set L as "l1l2l3l4", "l 20 l 21 l 22 l 23 ", "l5l 15 l 25 l 35 ”.
[0111] Table 1
[0112] character phoneme today <![CDATA[l1l2l3l4]]> today <![CDATA[l 20 l 21 l 22 l 23 > きょう <![CDATA[l5l 15 l 25 l 35 >
[0113] S203, calling the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs; the trained target language classification model is trained by the phoneme sequences of multiple training audios, each training audio has a marked preset language label, and the trained target language classification model records the correspondence between the phoneme sequence of the audio and the language to which the audio belongs.
[0114] This step is used to obtain the language category to which the audio to be classified belongs based on the trained target language classification model.
[0115] In some possible implementation manners, the trained target language classification model includes a feature extraction sub-model and a language determination sub-model. Here, calling the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs specifically includes: calling the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector, where the phoneme feature vector is composed of multiple phoneme features of the phoneme sequence, and the phoneme features have a corresponding relationship with the language; calling the language determination sub-model to process the phoneme feature vector to obtain the probability that the audio to be classified belongs to each language, and determining the language to which the audio to be classified belongs according to the probability that the audio to be classified belongs to each language.
[0116] Specifically, the feature extraction sub-model may be composed of an input layer, an output layer, and at least one intermediate layer, and is used to extract a phoneme feature vector composed of multiple phoneme features from the phoneme sequence. The phoneme feature vector can be represented by a vector of a fixed length or can be converted into a vector of a fixed length.
[0117] In some possible implementation manners, the feature extraction sub-model includes an embedding layer, a self-attention layer, and a batch normalization layer after training. The above-mentioned manner of calling the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector specifically includes: calling the embedding layer to perform vector encoding on the phoneme sequence to obtain a phoneme embedding vector; calling the self-attention layer to process the phoneme embedding vector based on the relevance between each vector component in the phoneme embedding vector to obtain an initial feature vector; calling the batch normalization layer to normalize the initial feature vector to obtain a phoneme feature vector.
[0118] Among them, the embedding layer here can be understood as the input layer in the above-mentioned feature extraction sub-model, the self-attention layer here can be understood as the intermediate layer above, and the batch normalization layer here can be understood as the output layer above.
[0119] Specifically, when the embedding layer performs vector encoding on the phoneme sequence, it can look up the embedding vector representation matching each phoneme in a preset embedding dictionary matrix, and then use the embedding vector representation matching each phoneme as the vector component in the phoneme embedding vector. Optionally, each row of the preset embedding dictionary matrix is the embedding vector representation of a phoneme.
[0120] Optionally, to enable the trained language model to understand the position (i.e., the arrangement order) of the phonemes in the phoneme sequence, position vectors can also be used to represent the position of each phoneme in the phoneme sequence. The vector dimension of the position vector is the same as the vector dimension of the embedding vector, and it can be obtained by looking up in a preset position vector matrix. Adding the embedding vector corresponding to the phoneme in the phoneme sequence to the position vector can obtain the vector component corresponding to the phoneme in the phoneme embedding vector.
[0121] Furthermore, the self-attention layer learns the positional, semantic, and other correlations between the phonemes corresponding to the respective vector components in the phoneme embedding vector, and performs attention scoring on each vector component based on the learned correlations. The scoring results of all vector components can be represented by the initial feature vector. For example, in certain specific languages, the usage frequency of some phonemes is higher than that in other languages. Therefore, these phonemes can be regarded as the key phonemes of this specific language. That is to say, more attention can be paid to whether this phoneme exists in the phoneme sequence, and a higher attention score can be assigned to this phoneme. Based on the higher attention score, the subsequent language classification sub-model is more likely to determine that the language to which the audio to be classified belongs is this specific language.
[0122] Furthermore, the batch normalization layer can normalize the initial features to a standard normal distribution to obtain the phoneme feature vector. The batch normalization layer can alleviate the gradient explosion / vanishing phenomenon during the training of the feature extraction sub-model and improve the training effect of the feature extraction sub-model.
[0123] Optionally, a residual connection can also be added between the self-attention layer and the batch normalization layer in the feature extraction sub-model. Here, the residual connection means that during the processing of batch normalization, the vector obtained by superimposing the phoneme embedding vector and the initial feature vector will be processed. This method can further reduce the model complexity of the feature extraction sub-model and prevent the gradient explosion / vanishing phenomenon during training.
[0124] Each vector component in the phoneme feature vector determined according to the above process can represent a phoneme feature. The phoneme feature is used to represent the correlations in terms of position, semantics, etc. between the phonemes in the phoneme sequence, and has a corresponding relationship with the language of the audio to be classified.
[0125] Exemplarily, the phoneme feature vector can be expressed as (a1, a2, a3, a4, a5). "a1" is the first component in the vector representation (that is, representing a phoneme feature), "a2" is the second component in the vector representation, and so on. If there are three language categories, namely "Language A", "Language B", and "Language C", and the phoneme feature vector corresponding to Language A is: the components "a1" and "a2" both take the value of 1 and the remaining components both take the value of 0; the phoneme feature vector corresponding to Language B is: the components "a2" and "a3" both take the value of 1 and the remaining components both take the value of 0; the phoneme feature vector corresponding to Language C is: the components "a3" and "a4" both take the value of 1 and the remaining components both take the value of 0. Then, if the phoneme feature vector is (1, 1, 0, 0, 0), it can be determined that the language to which the audio to be classified belongs is Language A.
[0126] It should be noted that the above phoneme feature vectors are only examples. In specific implementation, the component values in the phoneme feature vectors may be any values, and the audio feature vectors may contain many components. Generally, a simple correspondence relationship like the above example cannot be directly obtained. Therefore, it is also necessary to use a language determination sub-model to process the phoneme feature vectors, map the values of the phoneme feature vectors to the probability values of the audio to be classified belonging to each language, and determine the language to which the audio to be classified belongs based on the probability values.
[0127] In some possible implementation manners, the above manner of calling the language determination sub-model to process the phoneme feature vectors to obtain the probabilities of the audio to be classified belonging to each language specifically includes: calling the mapping rules in the language determination sub-model to map the phoneme feature vectors to obtain the probabilities of the audio to be classified belonging to each language.
[0128] Exemplarily, the mapping rule can be represented by the following formula:
[0129] Y = WX + b
[0130] Wherein, "X" is the phoneme feature vector, "W" is the weight matrix, "b" is the bias, and "Y" is the probability value vector. Each component in "Y" is the probability of the audio to be classified belonging to each language. For example, when "Y" determined by the language determination sub-model is (0.4, 0.5, 0.1), "0.4", "0.5", and "0.1" are the probability values of the audio to be classified belonging to language A, language B, and language C in sequence. If the language is determined by taking the maximum probability value, the language to which the audio to be classified belongs is language B.
[0131] It should be noted that the language determination sub-model here can adopt a support vector machine (SVM) classifier, softmax classification, etc. The present application does not limit this. And the trained target language classification model (including the feature extraction sub-model and the language determination sub-model) can be obtained by training the initial language classification model. The training method can refer to the Figure 3 embodiments shown below and will not be described here. Exemplarily, the trained target language classification model can be a neural network model such as a machine translation model (Transformer), etc. The present application does not limit this. The trained target language classification model can learn the correlation between different phonemes and can also better process longer phoneme sequences.
[0132] In Figure 2In the corresponding embodiments, the present application can identify the language category to which an audio belongs based on the phoneme sequence corresponding to the audio. Compared with traditional language classification schemes, phonemes exclude the interference of irrelevant information such as tones and accompaniments in the audio and are directly associated with the speech pronunciation actions in the audio, and can serve as a strong basis for language classification (the speech pronunciation actions of different languages are different). It is applied to various language classification and recognition scenarios, improving the accuracy of language classification. Furthermore, the trained target acoustic model in the present application is obtained based on a phoneme dictionary, and this phoneme dictionary can adapt to multiple languages, so that the trained target acoustic model can process audio in different languages, improving the efficiency of language classification.
[0133] Next, Figure 2 the method for determining the trained target acoustic model and the trained target language classification model in the embodiments will be described. Refer to Figure 3 , which is a schematic flowchart of a method for training a language classification-related model provided by an embodiment of the present application. The related model here includes the models involved in the above method embodiments. This method is applied to the above language classification device and includes steps S301 to S303, where:
[0134] S301. Obtain a first training audio set, where the first training audio set includes at least one first training audio of at least one language.
[0135] Exemplarily, at least one language includes but is not limited to languages such as Chinese, English, Japanese, and Korean. Each language in at least one language can be represented by phonemes using the phoneme dictionary introduced in the corresponding embodiments above. Figure 2
[0136] Optionally, each first training audio in the first training audio set can be an audio processed through a preprocessing stage. The preprocessing stage can screen the duration, content, etc. of the first training audio to improve the accuracy of subsequent model training. Exemplarily, the shortest effective duration of the first training audio can be limited to 2 - 3 s (so as to at least identify a phoneme sequence containing 10 phonemes). Here, the effective duration refers to the duration containing speech content (that is, the duration other than the silent duration, pure accompaniment duration, etc.); and training audios in different scenarios (for example, lyrical, children's, etc. training audios) are preferably selected to improve the effectiveness and diversity of the training data, and thus improve the efficiency of subsequent model training.
[0137] S302. Train the parameters in the initial acoustic model based on the first training audio set to obtain a trained target acoustic model.
[0138] In some possible implementation manners, training the parameters in the initial acoustic model based on the first training audio set herein to obtain the trained target acoustic model specifically includes: obtaining the spectral features of each first training audio and the first character sequence corresponding to the first training audio; calling the initial acoustic model to process the spectral features of each first training audio to obtain the predicted phoneme sequence of each first training audio; obtaining the second character sequence of each first training audio according to the phoneme dictionary and the predicted phoneme sequence of each first training audio; training the parameters in the initial acoustic model based on the first character sequence and the second character sequence to obtain the trained target acoustic model.
[0139] Specifically, the above-mentioned first character sequence is the character sequence of the real text of the first training audio. If the first training audio is music, the first character sequence is the character sequence of the music lyrics. If the first training audio is a vocal segment in a film or television work, the first character sequence is the character sequence of the subtitle corresponding to the vocal segment, and so on. The above-mentioned second character sequence is the second character sequence obtained from the predicted phoneme sequence output by the initial acoustic model. The loss function of the initial acoustic model can be obtained from the first character sequence and the second character sequence. This loss function is used to represent the similarity degree between the first character sequence and the second character sequence. The higher the similarity degree between the first character sequence and the second character sequence, the more accurate the predicted phoneme sequence predicted by the initial acoustic model. Exemplarily, this loss function can be an absolute value loss function, a squared loss function, a cross-entropy loss function, etc., and the present application does not limit this; when the loss function is selected as the absolute value loss function, when the value of the loss function decreases to a stable state as the training progresses, it represents the end of the training stage.
[0140] Exemplarily, Figure 4 is a schematic diagram of the training process of an initial acoustic model provided by an embodiment of the present application. As Figure 4 shown, the spectral features of each obtained first training audio can be used as input data and input into the initial acoustic model to obtain the predicted phoneme sequence L1; then, through the phoneme dictionary (the phoneme dictionary stores the corresponding relationship between the phoneme set L and the characters), the second character sequence W2 corresponding to the predicted phoneme sequence L1 is obtained; then, the loss function is used to calculate the difference between the first character sequence W1 and the second character sequence W2, and the calculated result is the loss value; finally, the loss value is passed back to the initial acoustic model to train the parameters in the model.
[0141] Optionally, the second character sequence W2 corresponding to the predicted phoneme sequence L1 obtained through the phoneme dictionary can be understood as: according to the arrangement order of each phoneme in the predicted phoneme sequence L1, the phoneme combination that can constitute a character is searched in the phoneme dictionary from front to back to obtain the second character sequence W2. For example, if L1 includes 5 phonemes, in the order from front to back, the characters corresponding to the combination of the first three phonemes and the characters corresponding to the combination of the last two phonemes can be found in the phoneme dictionary, then W2 is a sequence containing the characters corresponding to these two combinations.
[0142] Furthermore, when searching, the principle of the largest number of phonemes can be followed. For example, if the combination of the first two phonemes in L1 is found to have a corresponding character, the combination of the first three phonemes in L1 also has a corresponding character. In this case, the characters corresponding to the combination of the first three phonemes can be used as the final search result. This process is similar to the situation of front and back nasals in Chinese, and can effectively avoid missing searches. Taking Chinese vowels as an example, if the predicted phoneme sequence is "yingw.....", although "yin" can find the characters "音", "引", etc. in the phoneme dictionary, "g" and "w" cannot form characters. Here, the principle of the largest number of phonemes should be followed to find the character corresponding to "ying" in the phoneme dictionary.
[0143] It should be noted that the above method of obtaining the spectral features of each first training audio is the same as Figure 2 The method of obtaining the frequency spectrum features of the audio to be classified in the corresponding embodiment is the same and will not be repeated here.
[0144] S303: Train the parameters in the initial language classification model based on the first training audio set and the trained target acoustic model to obtain a trained target language classification model.
[0145] In some possible implementations, the parameters in the initial language classification model are trained based on the first training audio set and the trained target acoustic model to obtain the trained target language classification model, specifically including: obtaining a to-be-processed phoneme sequence of each first training audio, where the to-be-processed phoneme sequence is obtained by calling the trained target acoustic model to process the spectral features of each first training audio; obtaining the preset language to which each first training audio belongs; calling the initial language classification model to process the to-be-processed phoneme sequence of each first training audio to obtain the predicted language to which each first training audio belongs; training the parameters in the initial language classification model based on the preset language and the predicted language to which each first training audio belongs to obtain the trained target language classification model.
[0146] Among them, the to-be-processed phoneme sequence of each of the above first training audios is obtained by calling the trained target acoustic model. That is to say, in this step, the trained target acoustic model obtained by training in the above step S302 is used to process each of the first training audios again to obtain the to-be-processed phoneme sequence.
[0147] Optionally, in some possible implementation manners, it is also possible to obtain the to-be-processed phoneme sequence of each of the first training audios annotated based on a phoneme dictionary. In this possible implementation manner, the accuracy of the obtained to-be-processed phoneme sequence is higher, which can make the effect of subsequent training of the initial language classification model better.
[0148] Specifically, the preset language to which each of the above first training audios belongs is the language of each of the first training audios in the real situation, and this language is also the language to which the first character sequence of the above first training audio belongs. In the process of training the initial language classification model, the embodiment of the present application can obtain the loss function for model training based on the preset language to which each of the first training audios belongs and the predicted language (that is, obtained by processing the to-be-processed phoneme sequence by the initial acoustic model). This loss function is used to indicate whether the predicted language is the same as the preset language. For example, this loss function can be a 0-1 loss function, a cross-entropy loss function, etc., and the present application does not limit this; when the value of the loss function decreases to a stable state as the training progresses, it represents the end of the training phase.
[0149] For example, Figure 5 is a schematic diagram of the training process of an initial language classification model provided by an embodiment of the present application. As Figure 5 shown, each obtained to-be-processed phoneme sequence can be used as input data to be input into the initial language classification model to obtain the predicted language T2 to which each of the first training audios belongs; then the loss function is used to calculate the difference between the predicted language T2 to which each of the first training audios belongs and the preset language T1, and the calculation result is the loss value; finally, this loss value is fed back to the initial language classification model to train the parameters in the model.
[0150] In some possible implementation manners, after the present application obtains the trained target acoustic model and the trained target language classification model according to the above process, it is also possible to further perform fine-tuning training on these two models to improve the accuracy of language classification. Among them, the fine-tuning training process specifically includes: obtaining a second training audio set, where the second training audio set includes at least one second training audio of at least one language; obtaining the spectral features of each second training audio and the preset language to which it belongs; sequentially calling the target acoustic model and the target language classification model to process the spectral features of each second training audio to obtain the predicted language to which each second training audio belongs; and updating the target acoustic model and the target language classification model based on the predicted language and the preset language to which each second training audio belongs.
[0151] Optionally, the preset language to which each second training audio belongs is included in the preset languages to which the first training audio in the first training audio set belongs. Among them, the preset language to which the second training audio in the second training audio set belongs is related to the specific language classification scenario. For example, if the preset languages to which the first training audio in the first training audio set belong include Mandarin in Chinese, Cantonese in Chinese, English, Japanese, and Korean, the target acoustic model and the target language classification model obtained by training with the first training audio set can classify the audio of these five languages (the classification principle is the same as the corresponding description in the above Figure 2 corresponding embodiment and will not be elaborated here). When only some of these five languages (for example, Mandarin, Cantonese, and English) need to be classified in a specific application scenario, the target acoustic model and the target language classification model obtained above can be stacked together as a whole model for fine-tuning and updating through this possible implementation manner. The updated target acoustic model and target language classification model have better classification effects on Mandarin, Cantonese, and English.
[0152] Optionally, the number of second training audios included in the second training audio set may also be less than or equal to the number of first training audios included in the first training audio set. This optional method can reduce the processing workload of model fine-tuning and improve the processing efficiency. Moreover, the method of training the target acoustic model and the target language classification model again can also improve the accuracy of language classification.
[0153] In some embodiments, when the present application is applied to a specific language classification scenario and the language classification scenario includes a new language that was not involved in the training during training, a training audio set including the new language can be re-obtained to train the entire model of the above-obtained target language classification model again.
[0154] From Figure 2 the corresponding embodiment, it can be seen that when the phoneme sequence of an audio (including various training audios and the audio to be classified) is input into the above-trained target language classification model, first, the feature extraction sub-model in the target language classification model extracts a phoneme feature vector composed of multiple phoneme features from the phoneme sequence, and then the language determination sub-model maps the values of the phoneme feature vector to the probability values that the audio to be classified belongs to each language and determines the language to which the audio belongs based on the probability values. Thus, it can be seen that the phoneme sequence is determined by the target acoustic model based on the phonemes in the phoneme dictionary, and the phoneme feature vector determined by the phoneme sequence is not affected by the subsequent mapping to determine the language to which it belongs.
[0155] Therefore, in some other possible implementation manners, if it is required that the target language classification model also has the language classification function for new languages not participating in the training, the present application may also only update and train the language determination sub-model in the above-mentioned trained target language classification model, and no longer update and train the feature extraction sub-model. Compared with retraining the entire target language classification model, this method can effectively reduce the training workload and improve the training efficiency. The training process specifically includes: obtaining a third training audio set, where the third training audio set includes at least one third training audio of at least one language; the preset language to which the third training audio belongs is different from the preset language to which the first training audio belongs; obtaining the preset language to which each third training audio belongs; sequentially calling the target acoustic model and the target language classification model to process each third training audio to obtain the predicted language of each third training audio; and updating the language determination sub-model in the target language classification model based on the preset language and the predicted language to which each third training audio belongs.
[0156] Specifically, the third training audio is a new language audio (different from the language to which the first training audio belongs), and the at least one language here includes two possible cases: one new language and multiple new languages. In the process of updating the feature extraction sub-model, it is necessary to sequentially call the target acoustic model and the target language classification model to process each third training audio to obtain the predicted language of each third training audio. This process is the same as the process in the corresponding embodiment above and will not be elaborated here; then, the loss function value is determined from the predicted language and the preset language of the third training audio (that is, the language in the real situation), and the language determination sub-model is updated based on the value of the loss function. The updated language determination sub-model can map the phoneme feature vector to the probability distribution of languages including the new language. Figure 2 Correspondingly, the updated language determination sub-model can map the phoneme feature vector to the probability distribution of languages including the new language.
[0157] Exemplarily, Figure 6 is a schematic diagram before and after the update of a language determination sub-model provided by an embodiment of the present application. As Figure 6 shown, when the language determination sub-model before the update receives the phoneme feature vector (a1, 2, 3,..., N ), through the matrix operation inside the model, the phoneme feature vector is mapped to the probabilities that the audio to be classified belongs to language A and language B are 0.7 and 0.3 respectively. Based on this probability, it can be determined that the audio to be classified belongs to language A. When the updated language determination sub-model receives the above phoneme feature vector, through the matrix operation inside the model, the phoneme feature vector is mapped to the probabilities that the audio to be classified belongs to language A, language B, language C, and language D (language C and language D are new languages) are 0.3, 0.1, 0.5, and 0.1 respectively. Based on this probability, it can be determined that the audio to be classified belongs to language C.
[0158] It should be noted that when this application is applied to specific products and technologies, the first training audio set, the second training audio set, the third training audio set in this embodiment, and the relevant data of these training audio sets need to obtain the permission or consent of the object before they can be obtained, and the collection, use, and processing of these data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0159] In Figure 3 the corresponding embodiment, this application can first train the initial acoustic model to obtain the trained target acoustic model, and then train the initial language classification model based on the trained target acoustic model to obtain the trained target language classification model. The two trained models can be applied to Figure 2 the corresponding embodiment.
[0160] See Figure 7 , which is a schematic structural diagram of a language classification device provided by an embodiment of this application. The language classification device includes an acquisition unit 701 and a processing unit 702. Among them:
[0161] The acquisition unit 701 is used to acquire the spectral features of the audio to be classified that has been trained.
[0162] The processing unit 702 is used to call the trained target acoustic model to process the spectral features to obtain the phoneme sequence of the audio to be classified; the trained target acoustic model is a neural network model trained based on a phoneme dictionary, and the phoneme dictionary is used to indicate the correspondence between characters and phonemes of different languages.
[0163] The processing unit 702 is further used to call the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs; the trained target language classification model is trained by the phoneme sequences of multiple training audios, and each training audio has a labeled preset language label. The trained target language classification model records the correspondence between the phoneme sequence of the audio and the language to which the audio belongs.
[0164] In a possible implementation manner, the trained target language classification model includes a feature extraction sub-model and a language determination sub-model. When the processing unit 702 is used to call the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs, it specifically includes:
[0165] Call the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector, which is composed of multiple phoneme features of the phoneme sequence, and the phoneme features have a correspondence with the language.
[0166] Call the language determination sub-model to process the phoneme feature vector, obtain the probabilities of the audio to be classified belonging to each language, and determine the language to which the audio to be classified belongs according to the probabilities of the audio to be classified belonging to each language.
[0167] In a possible implementation manner, the trained feature extraction sub-model includes an embedding layer, a self-attention layer, and a batch normalization layer;
[0168] When the processing unit 702 is used to call the feature extraction sub-model to process the phoneme sequence and obtain a phoneme feature vector, it specifically includes:
[0169] Call the embedding layer to perform vector encoding on the phoneme sequence to obtain a phoneme embedding vector;
[0170] Call the self-attention layer to process the phoneme embedding vector based on the correlation between each vector component in the phoneme embedding vector to obtain an initial feature vector;
[0171] Call the batch normalization layer to normalize the initial feature vector to obtain a phoneme feature vector.
[0172] In a possible implementation manner, when the processing unit 702 is used to call the language determination sub-model to process the phoneme feature vector and obtain the probabilities of the audio to be classified belonging to each language, it specifically includes:
[0173] Call the mapping rule in the language determination sub-model to map the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language.
[0174] In a possible implementation manner, the obtaining unit 701 is further used for:
[0175] Obtain a first training audio set, where the first training audio set includes at least one first training audio of at least one language; obtain the spectral features of each first training audio and the first character sequence corresponding to the first training audio;
[0176] The processing unit 702 is further used for:
[0177] Call the initial acoustic model to process the spectral features of each first training audio to obtain the predicted phoneme sequence of each first training audio;
[0178] According to the phoneme dictionary and the predicted phoneme sequence of each first training audio, obtain the second character sequence of each first training audio;
[0179] Training the parameters in the initial acoustic model based on the first character sequence and the second character sequence to obtain the trained target acoustic model.
[0180] In a possible implementation manner, the obtaining unit 701 is further configured to:
[0181] Obtain the phoneme sequence to be processed of each first training audio, where the phoneme sequence to be processed is obtained by processing the spectral features of each first training audio by invoking the trained target acoustic model;
[0182] Obtain the preset language to which each first training audio belongs;
[0183] The processing unit 702 is further configured to:
[0184] Invoke the initial language classification model to process the phoneme sequence to be processed of each first training audio to obtain the predicted language to which each first training audio belongs;
[0185] Training the parameters in the initial language classification model based on the preset language and the predicted language to which each first training audio belongs to obtain the trained target language classification model.
[0186] In a possible implementation manner, the obtaining unit 701 is further configured to:
[0187] Obtain a second training audio set, where the second training audio set includes at least one second training audio of at least one language;
[0188] Obtain the spectral features and the preset language to which each second training audio in the second training audio set belongs;
[0189] The processing unit 702 is further configured to:
[0190] Successively invoke the target acoustic model and the target language classification model to process the spectral features of each second training audio to obtain the predicted language to which each second training audio belongs;
[0191] Updating the target acoustic model and the target language classification model based on the predicted language and the preset language to which each second training audio belongs.
[0192] In a possible implementation manner, the preset language to which each second training audio belongs is included in the preset languages to which the first training audios in the first training audio set belong; the number of second training audios included in the second training audio set is less than or equal to the number of first training audios included in the first training audio set.
[0193] In a possible implementation, the obtaining unit 701 is further configured to:
[0194] Obtain a third training audio set, where the third training audio set includes at least one third training audio of at least one language; the preset language to which the third training audio belongs is different from the preset language to which the first training audio belongs;
[0195] Obtain the preset language to which each third training audio in the third training audio set belongs;
[0196] The processing unit 702 is further configured to:
[0197] Call the target acoustic model and the target language classification model in sequence to process each third training audio, and obtain the predicted language of each third training audio; based on the preset language and the predicted language to which each third training audio belongs, update the target language classification model.
[0198] It should be noted that the functions of the various unit modules of the language classification device in the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here.
[0199] See Figure 8 , which is a schematic structural diagram of a terminal device provided in an embodiment of the present application. As Figure 8 shown, the terminal device in the embodiments of the present application may include: one or more processors 801, a memory 802, and a network interface 803. The above-mentioned processors 801, memory 802, and network interface 803 are connected through a bus 804. The memory 802 is used to store a computer program, and the computer program includes program instructions. The processor 801 and the network interface 803 are used to execute the program instructions stored in the memory 802 and perform the following operations:
[0200] Obtain the spectral features of the audio to be classified;
[0201] Call the trained target acoustic model to process the spectral features, and obtain the phoneme sequence of the audio to be classified; the trained target acoustic model is a neural network model trained based on a phoneme dictionary, and the phoneme dictionary is used to indicate the correspondence between characters and phonemes of different languages;
[0202] Call the trained target language classification model to process the phoneme sequence, and obtain the language to which the audio to be classified belongs; the trained target language classification model is trained by the phoneme sequences of multiple training audios, and each of the training audios has a labeled preset language label. The trained target language classification model records the correspondence between the phoneme sequence of the audio and the language to which the audio belongs.
[0203] It should be understood that in some feasible embodiments, the above-mentioned processor 801 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc. The above-mentioned memory 802 may include a read-only memory (ROM) and a random access memory (RAM), and provide instructions and data to the processor 801. A part of the memory 802 may also include a non-volatile random access memory. For example, the memory 802 may also store information about the device type.
[0204] In specific implementation, the above-mentioned terminal device may execute the implementation manners provided in each of the above Figures 2 to 3 steps. For the specific implementation manners, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated herein.
[0205] The embodiment of the present application also provides a computer-readable storage medium, and computer-readable instructions executed by the text processing device mentioned above are stored in the computer-readable storage medium. The computer-readable instructions include program instructions. When the processor executes the above program instructions, it can execute the Figures 2 to 3 method in the corresponding embodiment mentioned above. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application. As an example, the program instructions may be deployed on a computer device, or executed on multiple computer devices located at one place, or alternatively, executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network may form a blockchain system.
[0206] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can execute the method in the corresponding embodiment described above. Therefore, it will not be elaborated herein. Figures 2 to 3 The method in the corresponding embodiment is thus not elaborated herein any further.
[0207] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above various methods.
[0208] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A language classification method, characterized in that, The method includes: Obtaining the spectral features of the audio to be classified; Invoking the trained target acoustic model to process the spectral features to obtain the phoneme sequence of the audio to be classified; the target acoustic model is a neural network model trained based on a phoneme dictionary, and the phoneme dictionary is used to indicate the correspondence between characters of different languages and phonemes; Invoking the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs; the target language classification model is trained by the phoneme sequences of multiple training audios, each of the training audios has a labeled preset language label, and the trained target language classification model records the correspondence between the phoneme sequence of the audio and the language to which the audio belongs; The method further includes: Obtaining a first training audio set, the first training audio set including at least one first training audio of at least one language; obtaining the spectral features of each first training audio and the first character sequence corresponding to the first training audio; Invoking the initial acoustic model to process the spectral features of each first training audio to obtain the predicted phoneme sequence of each first training audio; According to the phoneme dictionary and the predicted phoneme sequence of each first training audio, obtaining the second character sequence of each first training audio; Training the parameters in the initial acoustic model based on the first character sequence and the second character sequence to obtain the trained target acoustic model.
2. The method according to claim 1, characterized in that, The target language classification model includes a feature extraction sub-model and a language determination sub-model. Invoking the trained target language classification model to process the phoneme sequence to obtain the language to which the audio to be classified belongs includes: Invoking the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector, the phoneme feature vector being composed of multiple phoneme features of the phoneme sequence, and the phoneme features having a correspondence with the language; Invoking the language determination sub-model to process the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language, and determining the language to which the audio to be classified belongs according to the probabilities of the audio to be classified belonging to each language.
3. The method according to claim 2, characterized in that, The feature extraction sub-model includes an embedding layer, a self-attention layer, and a batch normalization layer. Invoking the feature extraction sub-model to process the phoneme sequence to obtain a phoneme feature vector includes: Invoking the embedding layer to perform vector encoding on the phoneme sequence to obtain a phoneme embedding vector; Invoking the self-attention layer to process the phoneme embedding vector based on the correlation between each vector component in the phoneme embedding vector to obtain an initial feature vector; Invoking the batch normalization layer to normalize the initial feature vector to obtain a phoneme feature vector.
4. The method according to claim 2, wherein Invoking the language determination sub-model to process the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language includes: Invoking the mapping rule in the language determination sub-model to map the phoneme feature vector to obtain the probabilities of the audio to be classified belonging to each language.
5. The method according to claim 1, characterized in that, The method further includes: Obtain the to-be-processed phoneme sequence of each first training audio, where the to-be-processed phoneme sequence is obtained by processing the spectral features of each first training audio using the target acoustic model; Obtain the preset language to which each first training audio belongs; Call the initial language classification model to process the to-be-processed phoneme sequence of each first training audio, and obtain the predicted language to which each first training audio belongs; Based on the preset language and the predicted language to which each first training audio belongs, train the parameters in the initial language classification model to obtain the trained target language classification model.
6. The method according to claim 5, characterized in that, The method further includes: Obtain a second training audio set, where the second training audio set includes at least one second training audio of at least one language; obtain the spectral features and the preset language to which each second training audio in the second training audio set belongs; Successively call the target acoustic model and the target language classification model to process the spectral features of each second training audio, and obtain the predicted language to which each second training audio belongs; Based on the predicted language and the preset language to which each second training audio belongs, update the target acoustic model and the target language classification model.
7. The method according to claim 6, characterized in that, The preset language to which each second training audio belongs is included in the preset languages to which the first training audios in the first training audio set belong; the number of second training audios included in the second training audio set is less than or equal to the number of first training audios included in the first training audio set.
8. The method according to claim 5, wherein The method further includes: Obtain a third training audio set, where the third training audio set includes at least one third training audio of at least one language; the preset language to which the third training audio belongs is different from the preset language to which the first training audio belongs; Obtain the preset language to which each third training audio in the third training audio set belongs; Successively call the target acoustic model and the target language classification model to process each third training audio, and obtain the predicted language of each third training audio; based on the preset language and the predicted language to which each third training audio belongs, update the target language classification model.
9. A terminal device, characterized in that, The terminal device includes a processor, a memory, and a network interface, and the processor is connected to the memory and the network interface; the network interface is used to provide network communication functions, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to call the program code to implement the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Language recognition method and device
CN113744717A