Language recognition method and device, electronic equipment and storage medium

By conducting joint training of self-supervised phoneme segmentation and language recognition tasks on neural network models, the generated language recognition model effectively solves the problem of poor speech recognition performance in traditional algorithms without seeing, and improves the accuracy and generalization ability of language recognition.

CN120356456APending Publication Date: 2025-07-22BEIJING YUANJIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510630973.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, traditional language recognition algorithms have poor performance for unseen speech language recognition, poor generalization of algorithms, and deep learning-based language recognition algorithms have little language distinction and limited performance of phrase speech language recognition.

Method used

By conducting joint training on the neural network model of self-supervised phoneme segmentation task and language recognition task, a language recognition model is generated, and audio features are processed using the shared convolutional neural network layer, the Transformer encoder layer and the linear projection layer to achieve language recognition at the audio band level and sentence level.

Benefits of technology

It significantly improves the accuracy of language recognition, reduces the performance bottleneck of scenarios with small language distinction and phrase speech recognition, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356456A_ABST
    Figure CN120356456A_ABST
Patent Text Reader

Abstract

The invention provides a language recognition method and device, electronic equipment and a storage medium. The language recognition method comprises the following steps: acquiring a to-be-recognized audio; the method comprises the following steps: inputting to-be-recognized audio into a language recognition model to carry out audio feature extraction, carrying out language coding processing and phoneme coding processing on the audio features to generate a phoneme embedding vector sequence at an audio segment level, carrying out feature coding processing, sentence level statistical processing and linear projection processing on the phoneme embedding vector sequence, and carrying out sentence level statistical processing and linear projection processing on the phoneme embedding vector sequence. Outputting the language type of the to-be-recognized audio; wherein the language recognition model is obtained by performing joint training of a self-supervised phoneme segmentation task and a language recognition task on a neural network model. The language recognition model is obtained through joint training of the phoneme segmentation task and the language recognition task, and the accuracy of language recognition of the audio is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio recognition, and in particular, to a method, apparatus, electronic device, and storage medium for language identification. Background Art

[0002] Language identification (Spoken Language Identification, LID) refers to the technology of analyzing and processing speech segments to determine the language to which the speech belongs. Similar to speaker identification, language identification is also divided into two tasks: language discrimination and language confirmation. In the discrimination task, given a speech segment, the system has to select one of several possible languages as the language to which the speech segment belongs; in the confirmation task, given a speech segment, the system needs to determine whether the speech segment belongs to a certain language. Currently, traditional language identification algorithms have poor performance in identifying unseen speech languages, and the algorithm generalization ability is poor. The language identification algorithm based on deep learning effectively improves the performance of language identification. However, currently, the language identification algorithm based on deep learning has a significant decline in performance in scenarios where the language discrimination is not large (such as dialects), and at the same time, the performance of short speech language identification is limited. Therefore, how to improve the accuracy of language identification has become a technical problem that cannot be underestimated. Summary of the Invention

[0003] In view of this, the purpose of the present application is to provide a method, apparatus, electronic device, and storage medium for language identification. The language identification model obtained through the joint training of the phoneme segmentation task and the language identification task effectively improves the accuracy of audio language identification.

[0004] The embodiment of the present application provides a method for language identification, and the method for language identification includes:

[0005] Obtain the audio to be recognized;

[0006] Input the audio to be recognized into the language identification model to extract audio features, perform language encoding processing and phoneme encoding processing on the audio features to generate a sequence of phoneme embedding vectors at the audio segment level, perform feature encoding processing, sentence-level statistical processing, and linear projection processing on the sequence of phoneme embedding vectors, and output the language category of the audio to be recognized;

[0007] Wherein, the language identification model is obtained by jointly training a neural network model for a self-supervised phoneme segmentation task and a language identification task.

[0008] In a possible implementation manner, the performing language encoding processing and phoneme encoding processing on the audio features to generate a sequence of phoneme embedding vectors at the audio segment level includes:

[0009] Input the audio features into the shared convolutional neural network layer of the language recognition model for language encoding processing and phoneme encoding, and output the encoded audio features;

[0010] Input the encoded audio features into the audio segment-level statistical pooling layer of the language recognition model for mean processing and standard deviation processing, and concatenate the obtained mean vector and standard deviation vector, and output the concatenated feature vector;

[0011] Input the concatenated feature vector into the linear projection layer of the language recognition model for encoding processing, and output the phoneme embedding vector sequence at the audio segment level.

[0012] In a possible implementation manner, the feature encoding processing, sentence-level statistical processing, and linear projection processing of the phoneme embedding vector sequence to output the language category of the audio to be recognized include:

[0013] Input the phoneme embedding vector sequence into the Transformer encoder layer of the language recognition model, and perform feature encoding processing on the phoneme embedding vector sequence based on the self-attention mechanism and the multi-head attention mechanism, and output the feature-encoded phoneme embedding vector sequence;

[0014] Input the feature-encoded phoneme embedding vector sequence into the sentence-level statistical pooling layer of the language recognition model, perform mean processing and standard deviation processing and then aggregation processing, and output the language embedding vector at the sentence level;

[0015] Input the language embedding vector into the linear projection layer of the language recognition model to calculate the language category score to generate a language score matrix, and determine the language category corresponding to the index of the maximum value of the language score matrix in the language dictionary as the language category of the audio to be recognized.

[0016] In a possible implementation manner, the language recognition model is determined through the following steps:

[0017] Input the sample audio into the acoustic feature extraction layer of the neural network model for feature extraction processing, and output the sample audio features;

[0018] Input the sample audio features into the shared convolutional neural network layer of the neural network model for self-supervised phoneme segmentation tasks and language recognition tasks, and output the reduced-dimensional phoneme features and language features of the sample audio features;

[0019] Input the phoneme features and the language features after dimensionality reduction into the audio segment-level statistical pooling layer, linear projection layer, Transformer encoder layer, and sentence-level statistical pooling layer of the neural network model for processing, and output the predicted language category of the sample audio;

[0020] Based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the phoneme features after dimensionality reduction of each frame, determine the loss value of the neural network model;

[0021] If the loss value is greater than the preset loss value, perform iterative training on the neural network model. If the loss value is less than or equal to the preset loss value, stop the iterative training of the neural network model to generate the language recognition model.

[0022] In a possible implementation manner, the inputting the sample audio features into the shared convolutional neural network layer of the neural network model for self-supervised phoneme segmentation tasks and language recognition tasks, and outputting the phoneme features and language features after dimensionality reduction of the sample audio features includes:

[0023] Input the sample audio features into the shared convolutional neural network layer for phoneme encoding processing to output phoneme features, and perform dimensionality reduction processing on the phoneme features based on the linear projection layer to output the phoneme features after dimensionality reduction of each frame;

[0024] Input the sample audio features into the shared convolutional neural network layer of the neural network model for language encoding processing to output language features.

[0025] In a possible implementation manner, the determining the loss value of the neural network model based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the phoneme features after dimensionality reduction of each frame includes:

[0026] Calculate the phoneme features after dimensionality reduction of each frame based on the noise contrast estimation loss function to determine the noise contrast estimation loss value;

[0027] Calculate the predicted language category and the actual language category of the sample audio based on the cross-entropy loss function to determine the cross-entropy loss value;

[0028] Based on the noise contrast estimation loss value and the cross-entropy loss value, determine the loss value of the neural network model.

[0029] The embodiments of the present application also provide a language recognition device, and the language recognition device includes:

[0030] An acquisition module for acquiring the audio to be recognized;

[0031] An identification module for inputting the audio to be recognized into a language identification model to extract audio features, performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, and performing feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence to output the language category of the audio to be recognized; wherein, the language identification model is obtained by jointly training a neural network model for a self-supervised phoneme segmentation task and a language identification task.

[0032] In a possible implementation manner, when the identification module is used for performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, the identification module specifically is used for:

[0033] Inputting the audio features into a shared convolutional neural network layer of the language identification model for language encoding processing and phoneme encoding, and outputting the encoded audio features;

[0034] Inputting the encoded audio features into an audio segment-level statistical pooling layer of the language identification model for mean processing and standard deviation processing, and concatenating the obtained mean vector and standard deviation vector to output a concatenated representation vector;

[0035] Inputting the concatenated representation vector into a linear projection layer of the language identification model for encoding processing, and outputting a phoneme embedding vector sequence at the audio segment level.

[0036] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the language identification method as described above are executed.

[0037] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the language identification method as described above are executed.

[0038] A language recognition method, device, electronic device, and storage medium provided by an embodiment of the present application. The language recognition method includes: obtaining an audio to be recognized; inputting the audio to be recognized into a language recognition model to extract audio features, performing language encoding processing and phoneme encoding processing on the audio features to generate a sequence of phoneme embedding vectors at the audio segment level, and performing feature encoding processing, sentence-level statistical processing, and linear projection processing on the sequence of phoneme embedding vectors to output the language category of the audio to be recognized; wherein, the language recognition model is obtained by jointly training a neural network model for a self-supervised phoneme segmentation task and a language recognition task. The language recognition model obtained by the joint training of the phoneme segmentation task and the language recognition task effectively improves the accuracy of audio language recognition.

[0039] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides detailed descriptions as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0041] Figure 1 A flowchart of a language recognition method provided by an embodiment of the present application;

[0042] Figure 2 A schematic diagram of a language recognition method provided by an embodiment of the present application;

[0043] Figure 3 A schematic diagram of the structure of a language recognition device provided by an embodiment of the present application;

[0044] Figure 4 A schematic diagram of the structure of a language recognition device provided by an embodiment of the present application;

[0045] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts falls within the scope of protection of the present application.

[0047] First, the applicable application scenarios of the present application are introduced. The present application can be applied to the field of audio recognition technology.

[0048] Through research, it is found that spoken language identification (LID) refers to the technology of discriminating the language to which a speech segment belongs by analyzing and processing the speech segment. Similar to speaker identification, spoken language identification is also divided into two tasks: language discrimination and language confirmation. In the discrimination task, given a speech segment, the system has to select one of several possible languages as the language to which the speech segment belongs; in the confirmation task, given a speech segment, the system needs to determine whether the speech segment belongs to a certain language. Currently, traditional spoken language identification algorithms have poor performance in identifying unseen speech languages, and the algorithm generalization ability is poor. The spoken language identification algorithm based on deep learning effectively improves the performance of spoken language identification. However, currently, the spoken language identification algorithm based on deep learning has a significant decline in performance in scenarios where the language discrimination is not large (such as dialects), and at the same time, the performance of identifying short speech languages is limited. Therefore, how to improve the accuracy of spoken language identification has become a technical problem that cannot be underestimated.

[0049] Based on this, the embodiments of the present application provide a spoken language identification method. The spoken language identification model obtained through the joint training of the phoneme segmentation task and the spoken language identification task effectively improves the accuracy of audio language identification.

[0050] Please refer to Figure 1 , Figure 1 which is a flowchart of a spoken language identification method provided by the embodiments of the present application. As shown in Figure 1 , the spoken language identification method provided by the embodiments of the present application includes:

[0051] S101: Obtain the audio to be identified.

[0052] Here, the audio to be identified may include at least one speaker.

[0053] S102: Input the audio to be recognized into the language recognition model for audio feature extraction, perform language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, perform feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence, and output the language category of the audio to be recognized.

[0054] In this step, input the audio to be recognized into the language recognition model for audio feature extraction, perform language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, perform feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence, and output the language category of the audio to be recognized.

[0055] Among them, the language recognition model is obtained by jointly training a neural network model for self-supervised phoneme segmentation tasks and language recognition tasks.

[0056] Here, the process of inputting the audio to be recognized into the language recognition model for audio feature extraction is as follows: Use the pre-trained Wav2vec2-large-xlsr-53 model to extract high-dimensional representations from the audio to be recognized. The output of the 16th encoder layer of the Wav2vec2-large-xlsr-53 model for the audio to be recognized is used as the extracted acoustic representation x0 ∈ R T×D , where T is the number of frames of the acoustic representation, and D is the dimension of the acoustic representation, both set to 1024. The acoustic representation x0 is then reshaped into a three-dimensional vector x1 ∈ R T‘×20×D , where T‘ represents the number of segments of the reshaped acoustic representation, 20 represents that each acoustic representation has 20 frames, and D represents the dimension of the acoustic representation, all set to 1024. The specific reshaping rule is: First, truncate the representation to the maximum frame length divisible by 20, and then reshape it. Considering that the average duration of common phonemes is 50ms to 200ms, this solution defines the number of frames of each audio representation segment so that each phoneme unit in the model can cover at least two phonemes. The reshaped acoustic features are fed into a shared convolutional neural network layer, which is jointly optimized by the main task of language recognition branch and the auxiliary task of self-supervised phoneme segmentation task.

[0057] In a possible implementation manner, the performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level includes:

[0058] A: Input the audio features into the shared convolutional neural network layer of the language recognition model for language encoding processing and phoneme encoding, and output the encoded audio features.

[0059] Here, the audio features are input into the shared convolutional neural network layer of the language recognition model for language encoding processing and phoneme encoding, and the encoded audio features are output, where the encoded audio features contain phoneme information and language information.

[0060] B: The encoded audio features are input into the audio segment-level statistical pooling layer of the language recognition model for mean processing and standard deviation processing, and the obtained mean vector and standard deviation vector are concatenated to output the concatenated representation vector.

[0061] Here, in the audio segment-level statistical pooling layer, the encoded audio features are subjected to mean processing and standard deviation processing to obtain a mean vector and a standard deviation vector respectively, and the mean vector and the standard deviation vector are concatenated to obtain the concatenated representation vector.

[0062] C: The concatenated representation vector is input into the linear projection layer of the language recognition model for encoding processing to output a sequence of phoneme embedding vectors at the audio segment level.

[0063] Here, the concatenated representation vector is input into the linear projection layer for encoding processing to convert the concatenated representation vector into a sequence of phoneme embedding vectors at a higher dimension at the audio segment level.

[0064] In a possible implementation manner, the feature encoding processing, sentence-level statistical processing, and linear projection processing are performed on the sequence of phoneme embedding vectors to output the language category of the audio to be recognized, including:

[0065] a: The sequence of phoneme embedding vectors is input into the Transformer encoder layer of the language recognition model, and the sequence of phoneme embedding vectors is subjected to feature encoding processing based on the self-attention mechanism and the multi-head attention mechanism to output the sequence of phoneme embedding vectors after feature encoding.

[0066] Here, the sequence of phoneme embedding vectors is input into the Transformer encoder layer to model the global dependencies of phoneme information, thereby modeling the language information. The Transformer encoder layer calculates the similarity between each phoneme embedding vector and all other phoneme embedding vectors through the self-attention mechanism. This enables the model to capture the long-range dependencies between phoneme information. For example, certain phonemes may appear at different positions in the speech segment, but there is a certain pattern or correlation between them, and the Transformer can discover and utilize this correlation. To analyze the relationships between phoneme information from multiple perspectives, the Transformer typically uses the multi-head attention mechanism. This means that the model will calculate multiple attention weight matrices simultaneously to more comprehensively understand the input data. After the attention calculation, the Transformer further processes the data through a feed-forward neural network and combines residual connections to ensure that information is not lost. After being processed by the Transformer encoder layer, the originally isolated phoneme embedding vectors are integrated into an information representation with global semantics (the sequence of phoneme embedding vectors after feature encoding). This representation contains the complex relationships between phonemes in the speech segment, and these relationships are the key to distinguishing different languages.

[0067] b: Input the sequence of phoneme embedding vectors after feature encoding into the sentence-level statistical pooling layer of the language recognition model, perform mean processing and standard deviation processing, and then perform aggregation processing to output a sentence-level language embedding vector.

[0068] Here, the sequence of phoneme embedding vectors after feature encoding is input into the sentence-level statistical pooling layer, perform mean processing and standard deviation processing, and then perform aggregation processing to output a sentence-level language embedding vector.

[0069] c: Input the language embedding vector into the linear projection layer of the language recognition model to calculate the language category scores and generate a language score matrix, and determine the language category corresponding to the index corresponding to the maximum value of the language score matrix in the language dictionary as the language category of the audio to be recognized.

[0070] Here, input the language embedding vector into the linear projection layer to output a language score matrix, and predict the language as the language corresponding to the index corresponding to the maximum value in the score matrix in the language dictionary.

[0071] In a possible implementation manner, the language recognition model is determined through the following steps:

[0072] (1): Input the sample audio into the acoustic feature extraction layer of the neural network model for feature extraction processing, and output the sample audio features.

[0073] Here, the process of obtaining the sample audio features is consistent with the process of obtaining the above-mentioned audio features, and this part will not be elaborated further.

[0074] (2): Input the sample audio features into the shared convolutional neural network layer of the neural network model to perform self-supervised phoneme segmentation tasks and language identification tasks, and output the reduced-dimensional phoneme features and language features of the sample audio features.

[0075] Here, input the sample audio features into the shared convolutional neural network layer to perform self-supervised phoneme segmentation tasks and language identification tasks, and output the reduced-dimensional phoneme features and language features of the sample audio features.

[0076] In a possible implementation manner, the step of inputting the sample audio features into the shared convolutional neural network layer of the neural network model to perform self-supervised phoneme segmentation tasks and language identification tasks, and outputting the reduced-dimensional phoneme features and language features of the sample audio features includes:

[0077] I: Input the sample audio features into the shared convolutional neural network layer to perform phoneme encoding processing to output phoneme features, and perform dimensionality reduction processing on the phoneme features based on a linear projection layer, and output the reduced-dimensional phoneme features of each frame.

[0078] Here, a shared convolutional neural network layer is used for phoneme encoding. This network layer is a one-dimensional convolution with three layers having regularization techniques, ReLU, and batch normalization. Considering that the minimum phoneme duration is about 50ms, in this application, the convolution kernel size of the shared convolutional neural network layer is 1, so that the receptive field of the output unit of this module is shorter than the minimum phoneme duration. The output of the shared convolutional neural network layer is sent to a linear projection layer for dimensionality reduction to obtain the reduced-dimensional phoneme features.

[0079] II: Input the sample audio features into the shared convolutional neural network layer of the neural network model to perform language encoding processing, and output language features.

[0080] (3): Input the reduced-dimensional phoneme features and language features into the audio segment-level statistical pooling layer, linear projection layer, Transformer encoder layer, and sentence-level statistical pooling layer of the neural network model for processing, and output the predicted language category of the sample audio.

[0081] Here, the process of processing the reduced-dimensional phoneme features and language features is consistent with the process of processing the encoded audio features mentioned above, and this part will not be elaborated further.

[0082] (4): Determine the loss value of the neural network model based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the dimensionality-reduced phoneme features of each frame.

[0083] Here, use the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the dimensionality-reduced phoneme features of each frame to determine the loss value of the neural network model.

[0084] In a possible implementation manner, the determining the loss value of the neural network model based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the dimensionality-reduced phoneme features of each frame includes:

[0085] i: Calculate the dimensionality-reduced phoneme features of each frame based on the noise contrast estimation loss function to determine the noise contrast estimation loss value.

[0086] Here, determine the noise contrast estimation loss value through the following formula:

[0087]

[0088] Among them, L NCE is the noise contrast estimation loss value of the neural network model, K is the audio feature x of each audio segment t contains a total of K frames, T is the number of audio segments, L(z i ) is the frame-level noise contrast loss value, sim is the cosine similarity function, and the calculation process is to first use the current frame as the anchor point, the next frame as the positive sample, and non-adjacent frames as negative samples, calculate the cosine similarity between each frame and its positive sample and randomly sampled negative samples. If adjacent frames belong to the same phoneme, their similarity should be relatively high. Next, splice the calculated positive and negative sample similarities, represents the negative sample set of M z i These negative samples are randomly selected from all non-connected frames. Given that the similarity between adjacent frames can reveal phoneme boundaries, after optimization, the phoneme segmentation branch can spontaneously encode phoneme information. Since phoneme information represents the combination of phonemes in short segments of speech, the learned phoneme information can enable the model to fuse phoneme information into segment-level features.

[0089] ii: Calculate the predicted language category of the sample audio and the actual language category of the sample audio based on the cross-entropy loss function to determine the cross-entropy loss value;

[0090] iii: Determine the loss value of the neural network model based on the noise contrast estimation loss value and the cross-entropy loss value.

[0091] Here, the loss value of the neural network model is determined according to the added value of the noise contrast estimation loss value and the cross-entropy loss value.

[0092] (5): If the loss value is greater than the preset loss value, iterative training of the neural network model is performed. If the loss value is less than or equal to the preset loss value, iterative training of the neural network model is stopped to generate the language recognition model.

[0093] In this application, a language recognition model that hierarchically fuses phoneme and phoneme information is proposed. This model does not require any phoneme transcription or annotation during training. The model includes a self-supervised phoneme segmentation branch and a language recognition branch. The two branches share a convolutional neural network module, which simultaneously encodes language and sequential phoneme information from the input audio representation to generate a sequence of phoneme embedding vectors. This sequence is then encoded with language features using the Transformer encoder layer. This solution combines different language information in an end-to-end manner. Compared with the method of fusing language information using multiple subsystems, the complexity is lower. At the same time, the self-supervised training method effectively reduces the cost and difficulty of manually annotating phoneme sequences, and significantly improves the performance of language recognition.

[0094] Further, please refer to Figure 2 , Figure 2 , which is a schematic diagram of a language recognition method provided by an embodiment of this application. As shown in Figure 2 , the audio features are input into the shared convolutional neural network layer for language encoding and phoneme encoding, and the encoded audio features are output. The encoded audio features are input into the audio segment-level statistical pooling layer to output the concatenated representation vector. The concatenated representation vector is input into the linear projection layer to output a sequence of phoneme embedding vectors. The sequence of phoneme embedding vectors is input into the Transformer encoder layer to output the sequence of phoneme embedding vectors after feature encoding. The sequence of phoneme embedding vectors after feature encoding is input into the sentence-level statistical pooling layer to output the language embedding vector at the sentence level. The language embedding vector is input into the linear projection layer to output the language category of the audio to be recognized.

[0095] A language recognition method provided by an embodiment of the present application, the language recognition method includes: obtaining an audio to be recognized; inputting the audio to be recognized into a language recognition model for audio feature extraction, performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, performing feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence, and outputting the language category of the audio to be recognized; wherein, the language recognition model is obtained by jointly training a neural network model for a self-supervised phoneme segmentation task and a language recognition task. The language recognition model obtained by jointly training the phoneme segmentation task and the language recognition task effectively improves the accuracy of audio language recognition.

[0096] Please refer to Figure 3 、 Figure 4 , Figure 3 which is one of the structural schematic diagrams of a language recognition device provided by an embodiment of the present application; Figure 4 This is the second structural schematic diagram of a language recognition device provided by an embodiment of the present application. As Figure 3 shown in

[0097] an acquisition module 310, configured to acquire an audio to be recognized;

[0098] a recognition module 320, configured to input the audio to be recognized into a language recognition model for audio feature extraction, perform language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, perform feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence, and output the language category of the audio to be recognized; wherein, the language recognition model is obtained by jointly training a neural network model for a self-supervised phoneme segmentation task and a language recognition task.

[0099] Further, when the recognition module 320 is used for performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, the recognition module 320 is specifically configured to:

[0100] input the audio features into a shared convolutional neural network layer of the language recognition model for language encoding processing and phoneme encoding, and output the encoded audio features;

[0101] input the encoded audio features into an audio segment-level statistical pooling layer of the language recognition model for mean processing and standard deviation processing, and splice the obtained mean vector and standard deviation vector, and output the spliced representation vector;

[0102] Input the spliced representation vectors into the linear projection layer of the language recognition model for encoding processing, and output a sequence of phoneme embedding vectors at the audio segment level.

[0103] Further, when the recognition module 320 is used to perform feature encoding processing, sentence-level statistical processing, and linear projection processing on the sequence of phoneme embedding vectors and output the language category of the audio to be recognized, the recognition module 320 is specifically used for:

[0104] Input the sequence of phoneme embedding vectors into the Transformer encoder layer of the language recognition model, and perform feature encoding processing on the sequence of phoneme embedding vectors based on the self-attention mechanism and the multi-head attention mechanism, and output the sequence of phoneme embedding vectors after feature encoding;

[0105] Input the sequence of phoneme embedding vectors after feature encoding into the sentence-level statistical pooling layer of the language recognition model, perform mean processing and standard deviation processing and then aggregation processing, and output the language embedding vector at the sentence level;

[0106] Input the language embedding vector into the linear projection layer of the language recognition model to calculate the language category score and generate a language score matrix, and determine the language category corresponding to the index with the maximum value in the language score matrix in the language dictionary as the language category of the audio to be recognized.

[0107] Further, as Figure 4 shown, the language recognition device 300 further includes a model training module 330, and the model training module 330 determines the language recognition model through the following steps:

[0108] Input the sample audio into the acoustic feature extraction layer of the neural network model for feature extraction processing, and output the sample audio features;

[0109] Input the sample audio features into the shared convolutional neural network layer of the neural network model to perform self-supervised phoneme segmentation tasks and language recognition tasks, and output the reduced-dimensional phoneme features and language features of the sample audio features;

[0110] Input the reduced-dimensional phoneme features and the language features into the audio segment-level statistical pooling layer, linear projection layer, Transformer encoder layer, and sentence-level statistical pooling layer of the neural network model for processing, and output the predicted language category of the sample audio;

[0111] Based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the reduced-dimensional phoneme features of each frame, determine the loss value of the neural network model;

[0112] If the loss value is greater than a preset loss value, iterative training of the neural network model is performed. If the loss value is less than or equal to the preset loss value, iterative training of the neural network model is stopped to generate the language recognition model.

[0113] Further, when the model training module 330 is used to input the sample audio features into the shared convolutional neural network layer of the neural network model to perform self-supervised phoneme segmentation tasks and language recognition tasks, and output the reduced-dimensional phoneme features and language features of the sample audio features, the model training module 330 is specifically used for:

[0114] Input the sample audio features into the shared convolutional neural network layer for phoneme encoding processing to output phoneme features, and perform dimensionality reduction processing on the phoneme features based on the linear projection layer to output the reduced-dimensional phoneme features of each frame;

[0115] Input the sample audio features into the shared convolutional neural network layer of the neural network model for language encoding processing to output language features.

[0116] Further, when the model training module 330 is used to determine the loss value of the neural network model based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the reduced-dimensional phoneme features of each frame, the model training module 330 is specifically used for:

[0117] Calculate the reduced-dimensional phoneme features of each frame based on the noise contrast estimation loss function to determine the noise contrast estimation loss value;

[0118] Calculate the predicted language category and the actual language category of the sample audio based on the cross-entropy loss function to determine the cross-entropy loss value;

[0119] Determine the loss value of the neural network model based on the noise contrast estimation loss value and the cross-entropy loss value.

[0120] A language recognition device provided by an embodiment of the present application, the language recognition device includes: an acquisition module, configured to acquire an audio to be recognized; a recognition module, configured to input the audio to be recognized into a language recognition model to extract audio features, perform language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, and perform feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence, and output the language category of the audio to be recognized; wherein, the language recognition model is obtained by jointly training a neural network model on a self-supervised phoneme segmentation task and a language recognition task. The language recognition model obtained by jointly training the phoneme segmentation task and the language recognition task effectively improves the accuracy of audio language recognition.

[0121] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown in the figure, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.

[0122] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 runs, the processor 510 communicates with the memory 520 through the bus 530. When the machine-readable instructions are executed by the processor 510, they can execute the steps of the language recognition method in the method embodiments as described above Figure 1 and Figure 2 shown. For the specific implementation manner, reference can be made to the method embodiments, which will not be elaborated here.

[0123] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it can execute the steps of the language recognition method in the method embodiments as described above Figure 1 and Figure 2 shown. For the specific implementation manner, reference can be made to the method embodiments, which will not be elaborated here.

[0124] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.

[0125] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0126] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0127] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0128] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0129] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A language identification method, characterized in that, The language identification method includes: Obtaining the audio to be identified; Inputting the audio to be identified into a language identification model for audio feature extraction, performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, and performing feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence to output the language category of the audio to be identified; Among them, the language identification model is obtained by jointly training a neural network model for self-supervised phoneme segmentation tasks and language identification tasks.

2. The language recognition method according to claim 1, wherein The performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level includes: Inputting the audio features into the shared convolutional neural network layer of the language identification model for language encoding processing and phoneme encoding, and outputting the encoded audio features; Inputting the encoded audio features into the audio segment-level statistical pooling layer of the language identification model for mean processing and standard deviation processing, and concatenating the obtained mean vector and standard deviation vector to output the concatenated representation vector; Inputting the concatenated representation vector into the linear projection layer of the language identification model for encoding processing to output a phoneme embedding vector sequence at the audio segment level.

3. The language recognition method according to claim 1, wherein The performing feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence to output the language category of the audio to be identified includes: Inputting the phoneme embedding vector sequence into the Transformer encoder layer of the language identification model, and performing feature encoding processing on the phoneme embedding vector sequence based on the self-attention mechanism and the multi-head attention mechanism to output the phoneme embedding vector sequence after feature encoding; Inputting the phoneme embedding vector sequence after feature encoding into the sentence-level statistical pooling layer of the language identification model, performing mean processing and standard deviation processing and then aggregation processing to output a language embedding vector at the sentence level; Inputting the language embedding vector into the linear projection layer of the language identification model to calculate the language category score to generate a language score matrix, and determining the language category corresponding to the index of the maximum value of the language score matrix in the language dictionary as the language category of the audio to be identified.

4. The language identification method according to claim 1, wherein The language identification model is determined through the following steps: Inputting the sample audio into the acoustic feature extraction layer of the neural network model for feature extraction processing to output sample audio features; Inputting the sample audio features into the shared convolutional neural network layer of the neural network model for self-supervised phoneme segmentation tasks and language identification tasks to output the reduced-dimensional phoneme features and language features of the sample audio features; Inputting the reduced-dimensional phoneme features and the language features into the audio segment-level statistical pooling layer, linear projection layer, Transformer encoder layer, and sentence-level statistical pooling layer of the neural network model for processing to output the predicted language category of the sample audio. Determine the loss value of the neural network model based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the reduced-dimensional phoneme features of each frame; If the loss value is greater than the preset loss value, perform iterative training on the neural network model. If the loss value is less than or equal to the preset loss value, stop the iterative training of the neural network model to generate the language recognition model.

5. The language identification method according to claim 4, wherein The step of inputting the sample audio features into the shared convolutional neural network layer of the neural network model to perform self-supervised phoneme segmentation tasks and language recognition tasks, and outputting the reduced-dimensional phoneme features and language features of the sample audio features includes: Input the sample audio features into the shared convolutional neural network layer for phoneme encoding processing to output phoneme features, and perform dimensionality reduction processing on the phoneme features based on the linear projection layer to output the reduced-dimensional phoneme features of each frame; Input the sample audio features into the shared convolutional neural network layer of the neural network model for language encoding processing to output language features.

6. The language identification method according to claim 4, wherein The step of determining the loss value of the neural network model based on the loss function, the predicted language category of the sample audio, the actual language category of the sample audio, and the reduced-dimensional phoneme features of each frame includes: Calculate the reduced-dimensional phoneme features of each frame based on the noise contrast estimation loss function to determine the noise contrast estimation loss value; Calculate the predicted language category and the actual language category of the sample audio based on the cross-entropy loss function to determine the cross-entropy loss value; Determine the loss value of the neural network model based on the noise contrast estimation loss value and the cross-entropy loss value.

7. A language recognition device, characterized in that, The language recognition device includes: An acquisition module for acquiring the audio to be recognized; A recognition module for inputting the audio to be recognized into the language recognition model to extract audio features, performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, and performing feature encoding processing, sentence-level statistical processing, and linear projection processing on the phoneme embedding vector sequence to output the language category of the audio to be recognized; wherein, the language recognition model is obtained by jointly training a neural network model for self-supervised phoneme segmentation tasks and language recognition tasks.

8. The language recognition device according to claim 7, characterized in that When the recognition module is used for performing language encoding processing and phoneme encoding processing on the audio features to generate a phoneme embedding vector sequence at the audio segment level, the recognition module specifically is used for: Input the audio features into the shared convolutional neural network layer of the language recognition model for language encoding processing and phoneme encoding, and output the encoded audio features; Input the encoded audio features into the audio segment-level statistical pooling layer of the language recognition model for mean processing and standard deviation processing, and splice the obtained mean vector and standard deviation vector to output a spliced representation vector; Input the spliced representation vector into the linear projection layer of the language recognition model for encoding processing to output a phoneme embedding vector sequence at the audio segment level.

9. An electronic device, characterized in that, including: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device operates, the processor communicates with the memory via the bus, and when the machine-readable instructions are run by the processor, the steps of the language identification method according to any one of claims 1 to 6 are executed.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the language identification method according to any one of claims 1 to 6 are executed.