Language classification model training methods, language identification methods, devices and intelligent equipment

By training an N-layer encoder and fusion module to create a language classification model, global and local features of audio segments are captured, solving the problems of insufficient accuracy and robustness in language recognition in existing technologies, and achieving efficient language recognition in multilingual and complex scenarios.

CN119785773BActive Publication Date: 2026-01-30NIO TECH ANHUI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411902205.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-01-30
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing language recognition technologies have poor accuracy and robustness in multilingual and complex scenarios, making it difficult to meet the needs of in-vehicle voice interaction systems.

Method used

By acquiring the acoustic features of audio segments, a language classification model is trained using an N-layer encoder and a fusion module to capture global and local features of the audio segments. Combined with multi-head attention and attention statistical pooling layers, the accuracy and robustness of language recognition are improved.

Benefits of technology

It improves the accuracy and robustness of language recognition, making it suitable for in-vehicle voice interaction systems in multilingual and complex scenarios, and meeting the requirements for real-time and multilingual recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785773B_ABST
    Figure CN119785773B_ABST
Patent Text Reader

Abstract

This application relates to the field of intelligent speech processing technology, and provides a language classification model training method, a language recognition method, a device, and an intelligent device. The language classification model training method includes: acquiring a first acoustic feature, where the first acoustic feature refers to the acoustic feature of an audio segment; sequentially inputting the acoustic feature of the audio segment into an N-layer encoder of an initial language classification model to obtain a first encoded feature output by each of the N layers of encoders, where N is an integer greater than 1; inputting each of the first encoded features into a fusion module of the initial language classification model to obtain a first fused feature; and training the initial language classification model based on the first fused feature and the language recognition module of the initial language classification model to obtain a target language classification model, which is used to identify the language category of audio data. This application can improve the accuracy and robustness of language recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of intelligent speech processing technology, and in particular relates to a language classification model training method, a language recognition method, a device, and an intelligent equipment. Background Technology

[0002] In recent years, with the rapid development of intelligent vehicle technology, in-vehicle voice interaction systems have become one of the core functions for enhancing user experience. In a globalized context, automakers expanding into overseas markets need to address the demands of multilingual and multi-scenario voice recognition. Reliable and robust Spoken Language Identification (LID) is crucial for achieving high performance and accuracy in these systems. Existing language recognition technologies perform poorly in multilingual and complex scenarios, reducing the accuracy and robustness of language recognition. Summary of the Invention

[0003] This application provides a language classification model training method, a language identification method, a device, and an intelligent device, which can improve the accuracy and robustness of language identification.

[0004] In a first aspect, embodiments of this application provide a language classification model training method, including:

[0005] Obtain the first acoustic feature, which refers to the acoustic feature of an audio segment, wherein the audio segment includes at least one character of the language category to which it belongs;

[0006] The acoustic features of the audio segment are sequentially input into the N-layer encoder of the initial language classification model to obtain the first encoded features output by each of the N layers of encoders; N is an integer greater than 1.

[0007] Each of the first encoded features is input into the fusion module of the initial language classification model to obtain the first fused feature;

[0008] Based on the first fusion feature and the initial language classification model, the language identification module trains the initial language classification model to obtain the target language classification model, which is used to identify the language category of the audio data.

[0009] In this embodiment, by acquiring the acoustic features of an audio segment and sequentially inputting these acoustic features into the N-layer encoder of the initial language classification model, the first encoded features output by each of the N layers of encoders can be obtained. These first encoded features are then input into the fusion module of the initial language classification model to obtain the first fused features. Based on the first fused features and the language recognition module of the initial classification, the initial language classification model can be trained, thereby obtaining the target language classification model. This scheme fuses the encoded features output by each of the N layers of encoders, capturing both global and local features of the audio segment. Based on this, the initial language classification model is trained using the fused features (i.e., the first fused features) and the language recognition module, which improves the accuracy and robustness of the target language classification model's language recognition.

[0010] Secondly, embodiments of this application provide a language identification method, including:

[0011] Obtain the audio data to be recognized;

[0012] Obtain the second acoustic feature, which refers to the acoustic feature of the audio data to be identified;

[0013] Based on the second acoustic feature and the target language classification model, the language category of the audio data to be identified is determined; the target language classification model is obtained by training the initial language classification model based on the method described in the first aspect above.

[0014] The target language classification model includes an N-layer encoder, a fusion module, and a language identification module, where N is an integer greater than 1.

[0015] The encoder in layer N is used to: determine the corresponding second coding feature based on the second acoustic feature;

[0016] The fusion module is used to: fuse each of the second encoded features to obtain the target fused feature;

[0017] The language identification module is used to: determine the language category of the audio data to be identified based on the target fusion features.

[0018] In this embodiment, by acquiring the acoustic features (i.e., the second acoustic features) of the audio data to be identified and inputting these acoustic features into a target language classification model, the language category to which the audio data to be identified belongs can be obtained. The target language classification model includes an N-layer encoder, a fusion module, and a language identification module. The N-layer encoder is used to determine the corresponding second coding features based on the acoustic features of the audio data to be identified. The fusion module is used to fuse the second coding features to obtain a target fused feature. The language identification module is used to determine the language category to which the audio data to be identified belongs based on the target fused feature. This solution, when performing language identification on the audio data to be identified using the target language classification model, can capture both global and local features of the audio data to be identified, thereby improving the accuracy and robustness of language identification.

[0019] Thirdly, embodiments of this application provide a language classification model training device, which can be applied to a smart device or can function as a smart device. The language classification model training device includes:

[0020] The feature acquisition module is used to acquire a first acoustic feature, which refers to the acoustic feature of an audio segment, wherein the audio segment includes at least one character of the language category to which it belongs;

[0021] The feature encoding module is used to sequentially input the first acoustic features into the N-layer encoder of the initial language classification model to obtain the first encoded features output by each of the N layers of encoders, where N is an integer greater than 1;

[0022] The feature input module is used to input each of the first encoded features into the fusion module of the initial language classification model to obtain the first fused feature;

[0023] The model training module is used to train the initial language classification model based on the first fusion feature and the language identification module of the initial language classification model to obtain the target language classification model, which is used to identify the language category of the audio data.

[0024] Fourthly, embodiments of this application provide a language recognition device, which can be applied to a smart device or can function as a smart device. The language recognition device includes:

[0025] The data acquisition module is used to acquire the audio data to be recognized;

[0026] An acoustic acquisition module is used to acquire a second acoustic feature, which refers to the acoustic feature of the audio data to be identified.

[0027] The language identification module is used to determine the language category of the audio data to be identified based on the audio data to be identified and the target language classification model; the target language classification model is obtained by training an initial language classification model based on the method described in the first aspect above;

[0028] The target language classification model includes an N-layer encoder, a fusion module, and a language identification module, where N is an integer greater than 1.

[0029] The encoder in layer N is used to: determine the corresponding second coding features based on the second acoustic features;

[0030] The fusion module is used to: fuse each of the second encoded features to obtain the target fused feature;

[0031] The language identification module is used to: determine the language category of the audio data to be identified based on the target fusion features.

[0032] Fifthly, embodiments of this application provide an intelligent device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the language classification model training method as described in the first aspect above, or implements the language recognition method as described in the second aspect above.

[0033] Sixthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a computer, implements the language classification model training method as described in the first aspect above, or the language recognition method as described in the second aspect above.

[0034] In a seventh aspect, embodiments of this application provide a computer program product, including a computer program that, when run, causes the language classification model training method as described in the first aspect above to be executed, or causes the language recognition method as described in the second aspect above to be executed.

[0035] It is understood that the beneficial effects of the third to seventh aspects mentioned above can be found in the relevant descriptions in the first and second aspects mentioned above, and will not be repeated here. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart illustrating the language classification model training method provided in the embodiments of this application;

[0038] Figure 2 This is a flowchart illustrating the language recognition method provided in the embodiments of this application;

[0039] Figure 3 This is a schematic diagram of the language classification model training device provided in the embodiments of this application;

[0040] Figure 4 This is a schematic diagram of the language recognition device provided in the embodiments of this application;

[0041] Figure 5 This is a schematic diagram of the structure of the smart device provided in the embodiments of this application. Detailed Implementation

[0042] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0043] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0044] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0045] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0046] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0047] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0048] Please see Figure 1 , Figure 1 The illustration shows a flowchart of a language classification model training method provided in an embodiment of this application. This is an example and not a limitation, and the method can be applied to smart devices (e.g., smart devices may include driving equipment, smart cars, robots, etc.) or electronic devices. For example, a smart device could be a smart car. Electronic devices could be servers, mobile phones, tablets, computers, or other related electronic devices with model training capabilities. For example, a server could also be a cloud server that communicates with the smart device.

[0049] Specifically, Figure 1 The method provided in this application is described using an example of its application to smart devices. Specifically, the method includes the following steps:

[0050] Step 101: Obtain the first acoustic feature, which refers to the acoustic feature of the audio segment.

[0051] The audio segment includes at least one character representing the language category, which enables the audio segment to contain language information, thus allowing it to be used to train a language classification model.

[0052] Optionally, in order to achieve language recognition for short audio clips, the audio length of the aforementioned audio segment can be limited to a preset length range, the maximum value of which can be the length threshold described below. For example, the preset length range is 640 milliseconds to 2 seconds.

[0053] To improve the accuracy and robustness of the language classification model, a larger number of audio segments can be used to train the model, i.e., multiple audio segments. Of course, it is understood that the number of audio segments could be only one, and this application does not impose this limitation.

[0054] To better distinguish between untrained and trained language classification models, the untrained model can be referred to as the initial language classification model, while the trained model can be referred to as the target language classification model.

[0055] The first acoustic feature described above and the second acoustic feature described below are of the same type, and this application does not limit the specific type of the acoustic feature. For example, the acoustic feature can be a filter bank (FBank) feature, Mel-Frequency Cepstrum Coefficients (MFCC), Linear Predictive Cepstral Coefficient (LPCC), or Perceptual Linear Predictive (PLP) coefficients.

[0056] By way of example and not limitation, the first and second acoustic features in this application are FBank features, which are able to process audio signals in a manner similar to the human ear, thereby improving the performance of speech recognition.

[0057] In one embodiment, the smart device can acquire at least one audio training sample. It can directly use each audio training sample as an audio segment, or, in order to increase the diversity of audio segments, randomly cut each audio training sample to obtain an audio segment. This application does not limit this.

[0058] The aforementioned audio training samples constitute the training dataset for the initial language classification model, used to train the initial language classification model. When training the initial language classification model based on the training dataset, audio segments can be randomly cut from each audio training sample in the training dataset, and then the language classification model can be trained based on the audio segments. In other words, audio segments need to be randomly cut from the audio training samples in each training session, which ensures the diversity of audio segments.

[0059] For the clipped audio segments, data augmentation techniques such as temporal perturbation, noise addition, frequency domain transformation, speed adjustment, and temporal and frequency domain masking can be applied to enhance the data of the audio segments and improve the robustness of the target language classification model.

[0060] In one embodiment, the smart device can collect relevant language audio data from the network and user voice data actually collected in the vehicle. This data is used as raw audio data. The raw audio data can be used directly as audio training samples, or the raw audio data can be preprocessed by the Voice Activity Detection (VAD) module to remove silent segments from the raw audio data and crop out audio training samples containing effective language information, thereby eliminating the influence of non-human voice data on the training process of the initial language classification model.

[0061] The process of obtaining audio segments by cropping audio training samples by intelligent devices may include: obtaining the timestamp of each character in each audio training sample; and for any audio training sample, randomly cropping the audio training sample based on the timestamp of each character in the audio training sample during each training session to obtain multiple audio segments.

[0062] The timestamp of a character records the time information of that character in the corresponding audio training sample.

[0063] Optionally, the audio training samples can be processed using the Whisper model to obtain the timestamp of each character in the audio training samples. This timestamp is a character-level timestamp, which allows for more precise random cropping of the audio training samples, ensuring that the resulting audio segments include at least one character. For audio training samples without language category labels, or those labeled with multiple language categories, the Whisper model can be used to identify the language category, thus obtaining the true language value (i.e., the expression label). Based on this true language value, an initial language classification model can be trained.

[0064] This embodiment uses random cropping based on the timestamps of each character to ensure that each audio segment includes at least one character, thus containing language information that can be used to train the initial language classification model. Furthermore, random cropping of the audio training samples simulates the randomness and diversity of audio length in real-world scenarios (such as in-vehicle scenarios), thereby improving the training of the initial language classification model.

[0065] From the start of training the initial language classification model to its convergence, it usually requires many training rounds. During each training, each audio training sample is randomly cropped. Because it is random cropping, the audio segments obtained from the same audio training sample may be different in different cropping times. This makes the training data (i.e., audio segments) more consistent with the real in-vehicle scenario and improves the robustness of the target language classification model.

[0066] In one embodiment, a length threshold can be preset to obtain audio segments suitable for short audio recognition from the audio training samples. That is, the audio segments are relatively short. By training the initial language classification model with these audio segments, the accuracy and robustness of the target language classification model for short audio language recognition can be improved.

[0067] One approach to obtaining suitable audio segments for short audio recognition from audio training samples using a length threshold is as follows: if the audio length of the audio training sample exceeds the length threshold, the audio training sample is randomly cropped based on the timestamp of each character in the audio training sample during each training session; if the audio duration of the audio training sample does not exceed the length threshold, the audio training sample is determined to be an audio segment.

[0068] Optionally, a length threshold can be set according to the needs of the scenario. The specific value of the length threshold in this application is not limited. As an example and not a limitation, the length threshold is 2 seconds.

[0069] In real-world in-vehicle scenarios, user voice commands typically appear as short audio clips (e.g., audio length of 1 to 2 seconds). In streaming recognition scenarios, it's crucial to identify the language category of the audio within a few hundred milliseconds to facilitate subsequent speech recognition and Natural Language Understanding (NLU) decision-making logic. This embodiment utilizes a short audio enhancement method combining VAD pruning and random pruning. This method makes the audio clips used to train the initial language classification model more consistent with real-world in-vehicle scenarios and more accurately distinguishes features of highly similar languages. This improves the accuracy and robustness of the target language classification model in identifying the language of short audio clips, enabling efficient recognition of multilingual short audio clips. It is applicable to both real-world in-vehicle scenarios and streaming recognition scenarios.

[0070] As an example, and not a limitation, in a real-world in-vehicle scenario, intelligent devices collect audio data in relevant languages ​​from the internet and real-vehicle data. First, the VAD (Visual Audio Decoding) module is used to trim this audio data (which may include audio data in one language or multiple languages) to obtain audio training samples that match the real-world scenario. Then, the Whisper model is used to further identify the language of the audio training samples that are not labeled with a language category or whose language category accuracy is low, and the timestamp of each character in the audio training sample is output, resulting in a high-quality multilingual training dataset. To further improve the short audio classification capability, the audio training samples can be trimmed to a random length based on the timestamp to obtain audio segments containing at least one character of the corresponding language category. Furthermore, during random cropping, audio training samples with an audio length greater than 2 seconds are cropped, and audio segments containing real word pronunciations with a length of 640 milliseconds to 2 seconds are extracted. Data augmentation techniques such as speed variation, temporal perturbation, noise addition, and frequency variation are applied to the cropped audio segments. Then, the FBank feature extraction method is used to extract 80-dimensional FBank features from the augmented audio segments, and after performing temporal and frequency domain masking, they are used as input for subsequent model training to enhance the model's robustness to noise and interference.

[0071] Step 102: Input the first acoustic features into the N-layer encoder of the initial language classification model in sequence to obtain the first encoded features output by each of the N-layer encoders.

[0072] Where N is an integer greater than 1.

[0073] Optionally, the aforementioned N-layer encoder can be a portion of the encoders in the speech recognition model, or it can be all of the encoders; this application does not limit this.

[0074] Each encoder layer may sequentially include: a first position-wise feedforward network (FFN), a multi-head self-attention (MHA) layer module, a convolutional module, and a second FFN.

[0075] As an example, and not a limitation, let's take the i-th layer encoder as an example. The i-th layer encoder is any encoder in the N-layer encoder except for the first layer encoder. The encoded features output by the i-th layer encoder can be represented as follows:

[0076] encoder_out i =FFN i (Conv i (MHA i (FFN_Macaron i (encoder_outi-1 ))))

[0077] Among them, encoder_out i FFN_Macaron represents the encoded feature output by the i-th layer encoder. i Let FFN represent the first FFN in the i-th layer encoder. i Denotes the second FFN in the i-th layer encoder, Conv i MHA represents the convolutional module in the i-th layer encoder. i Represents the MHA in the i-th layer encoder, encoder_out i-1 This represents the encoded features output by the (i-1)th layer encoder.

[0078] Optionally, the above-mentioned speech recognition model can be any type of speech recognition model, and this application does not limit it.

[0079] As an example and not a limitation, the above speech recognition model is the Wav2Vec2.0-BERT model that includes a 12-layer encoder, and the N-layer encoder mentioned above can refer to these 12 layers.

[0080] Step 103: Input each first coding feature into the fusion module of the initial language classification model to obtain the first fused feature.

[0081] Optionally, after inputting each first coding feature into the fusion module, the fusion module can splice the first coding features to integrate more semantic information in the audio segment, thereby capturing the global features and fine-grained differences of the language through the synergistic effect of hierarchical information, solving the problem of unstable single-layer features, and is suitable for language classification tasks in multilingual and complex scenarios.

[0082] Since this embodiment selects N layers of encoders in the speech recognition model for feature extraction, rather than the entire speech recognition model, the above fusion module can also be called a lightweight layer fusion module. This avoids the computational complexity problem caused by directly splicing all layers, while also taking into account the real-time requirements of voice interaction scenarios (such as in-vehicle voice interaction scenarios).

[0083] The lightweight layer fusion module in this embodiment is isolated from the language recognition module, which accelerates the training process of the initial language classification model, increases the diversity of language recognition, and allows for targeted optimization of language recognition. This not only improves the accuracy and robustness of the target language classification model but also meets the performance requirements of in-vehicle real-time voice interaction.

[0084] Step 104: Based on the first fusion feature and the language identification module of the initial language classification model, train the initial language classification model to obtain the target language classification model.

[0085] The aforementioned target language classification model is used to identify the language category of audio data.

[0086] In one possible implementation, after inputting the first fused feature into the language recognition module, the N-layer encoder, fusion module, and language recognition module can be trained based on the output of the language recognition module. Alternatively, in step 102, the trained speech recognition model (i.e., the N-layer encoder has been pre-trained) can be used for feature encoding, and in step 103, the trained fusion module can be used for feature fusion. Based on this, the language recognition module can be trained based on the output of the language recognition module. This approach can avoid additional training costs.

[0087] In one possible implementation, the language identification module sequentially includes a first multi-head attention layer, a first attention statistical pooling layer, a second multi-head attention layer, and a second attention statistical pooling layer; the language identification module, based on the first fused features and the initial language classification model, trains the initial language classification model to obtain the target language classification model, including:

[0088] The first fused feature is input into the first multi-head attention layer to perform attention weighting on the time dimension of the first fused feature, thereby obtaining the second fused feature;

[0089] The second fusion feature is input into the first attention statistical pooling layer to fuse the features at different time steps in the second fusion feature to obtain the third fusion feature;

[0090] The third fusion feature is input into the second multi-head attention layer to perform attention weighting on the encoder dimension of the third fusion feature, thus obtaining the fourth fusion feature;

[0091] The fourth fusion feature is input into the second attention statistical pooling layer to fuse the encoded features output by the N layers of encoders contained in the fourth fusion feature, thus obtaining the fifth fusion feature.

[0092] Based on the fifth fusion feature, an initial language classification model is trained to obtain a target language classification model.

[0093] The first multi-head attention layer performs attention weighting on the temporal dimension of the first fused feature, enabling it to extract the relationships between features at different time steps within the first fused feature. The second multi-head attention layer performs attention weighting on the encoder dimension of the third fused feature, enabling it to extract the relationships between the encoded features output by the N encoder layers contained in the third fused feature.

[0094] The first multi-head attention layer focuses on the feature relationships at different time steps within an audio segment, capturing temporal correlations and dynamic change patterns. The second multi-head attention layer focuses on the information interaction and fusion between the outputs of different encoders, extracting rich contextual semantics from multiple levels. By combining the first and second multi-head attention layers, the language classification model can achieve a comprehensive expression of features across both time and layer dimensions, providing more discriminative feature representations for language classification tasks.

[0095] The first attention statistical pooling layer dynamically adjusts the attention distribution based on the actual features of the audio segment, effectively focusing on key language-related information while ignoring noise or redundancy. The second attention statistical pooling layer highlights the contributions of different encoders through weight allocation, achieving dynamic aggregation of layer features. By combining the first and second attention statistical pooling layers, not only is the adaptability of the language classification model to complex audio data improved, but the robustness and accuracy of language classification are also significantly enhanced. After fusing the first encoding features, the resulting first fused feature has a high dimensionality. To avoid redundant information interfering with downstream classification tasks and to improve computational efficiency, the initial language recognition module can also include a downsampling layer. Before inputting the first fused feature into the first multi-head attention layer, the first fused feature is first input into the downsampling layer to perform dimensionality reduction. This dimensionality reduction effectively reduces computational complexity while retaining important information in the first fused feature.

[0096] In one possible implementation, based on training an initial language classification model using a pre-trained speech recognition model and fusion module (i.e., the N-layer encoder and fusion module of the initial language classification model are pre-trained), the above-mentioned initial language classification model is trained based on the fifth fusion feature, including:

[0097] With the parameters of the N-layer encoder being the first parameter and the parameters of the fusion module being the second parameter, the parameters of the language recognition module are updated to the third parameter based on the fifth fusion feature.

[0098] The initial language classification model obtained by setting the parameters of the N-layer encoder as the first parameter, the parameters of the fusion module as the second parameter, and the parameters of the language recognition module as the third parameter is used as the target language classification model.

[0099] Since the N-layer encoder of the initial language classification model is pre-trained, the parameters of the aforementioned N-layer encoder are fixed, and the first parameter mentioned above refers to the parameters of the pre-trained N-layer encoder. Similarly, since the fusion module of the initial language classification model is also pre-trained, the parameters of the aforementioned fusion module are also fixed, and the second parameter mentioned above refers to the parameters of the pre-trained fusion module.

[0100] After the language recognition module is trained based on the fifth fusion feature, the parameters of the language recognition module are also fixed. Therefore, the third parameter mentioned above is the parameter of the trained language recognition module.

[0101] In this embodiment, when the parameters of the N-layer encoder are the first parameters and the parameters of the fusion module are the second parameters, updating the parameters of the language recognition module means that while freezing the parameters of the N-layer encoder and the fusion module, only the parameters of the language recognition module are updated. That is, in the training of the language classification model, the N-layer encoder and the fusion module only generate features during the forward inference process and do not participate in parameter updates. This method can make full use of the powerful feature extraction capability of the encoder and the feature fusion capability of the fusion module, while avoiding additional training costs, thereby improving the efficiency and stability of model training.

[0102] This embodiment maintains feature stability by freezing the parameters of the N-layer encoder and the fusion module. By combining multi-head attention layers and attention statistical pooling layers in both time and layer dimensions, it can fully extract information in both time and layer dimensions.

[0103] Optionally, the language recognition module may also include a fully connected layer and an activation function, that is, the fifth fusion feature is sequentially input into the fully connected layer and the activation function to obtain the language prediction value of the corresponding audio segment. The language prediction value and the true language value of the audio segment are used to update the parameters of the downsampling layer, the first multi-head attention layer, the first attention statistical pooling layer, the second multi-head attention layer, the second attention statistical pooling layer and the fully connected layer in the language recognition module.

[0104] During training, the language classification model can be optimized using the standard cross-entropy loss function to enable efficient multilingual recognition of short audio clips. After training, the target language classification model is used to determine the language.

[0105] After obtaining the target language classification model, a threshold mechanism can be constructed based on positive and negative sample data. As an example, and not a limitation, a large number of positive samples (target language audio data) and negative samples (non-target language audio data) from in-vehicle scenarios are analyzed to calculate the classification score distribution of the target language classification model for different languages. Based on the score distribution of positive and negative samples, a score threshold for language discrimination is set. After setting the score threshold, it can be applied to language recognition scenarios. For example, if the target language score is greater than the score threshold, the audio data is considered to belong to that language; otherwise, the language classification result is considered unclear, and a general multilingual resource package is recorded. After the target language classification score exceeds the score threshold, the target language pronunciation dictionary and Weighted Finite-State Transducer (WFST) resource packages are loaded to optimize the accuracy of the speech decoding process. A hot word list related to the target language is called to ensure that important words in specific domains are prioritized in speech recognition. The corresponding NLU module for the target language is called to further improve the ability to interpret the speech input intent. Language identification is performed using a target language classification model.

[0106] This application embodiment fuses the coding features output by each of the N-layer encoders, which can capture the global and local features of audio segments. Based on this, an initial language classification model is trained using the fused features and the language recognition module, which can improve the accuracy and robustness of the target language classification model in language recognition, and provide reliable technical support for intelligent voice interaction and multilingual support.

[0107] Please see Figure 2 , Figure 2 The diagram illustrates a language recognition method provided in an embodiment of this application. This is an example and not a limitation, and the method can be applied to smart devices (e.g., smart devices may include driving equipment, smart cars, robots, etc.) or electronic devices such as mobile phones. For example, a smart car is a smart device. Specifically... Figure 2 The following description uses the application of the method to a smart device as an example. Specifically, the smart device has a target language classification model, and the method includes the following steps:

[0108] Step 201: Obtain the audio data to be recognized.

[0109] The aforementioned audio data to be identified may refer to audio data for which language identification is required.

[0110] As an example and not a limitation, the audio data to be recognized can be audio data collected by the voice acquisition unit in the smart device, or audio data to be recognized input by the user to the smart device through the interactive interface of the smart device.

[0111] Taking a smart device as an example, when a user is riding in the vehicle, the voice acquisition unit in the vehicle can collect the language of one or more users communicating in real time, and then use the collected language as audio data to be recognized.

[0112] After acquiring the audio data to be recognized, the VAD module can be used to preprocess the audio data to remove silent segments and cut out audio segments containing valid language information. Language recognition is then performed based on these audio segments, thereby eliminating the influence of non-human voice data on the language recognition of the audio data to be recognized.

[0113] Step 202: Obtain the second acoustic feature, which refers to the acoustic feature of the audio data to be identified.

[0114] By preprocessing the audio data to be identified using the VAD module, a second acoustic feature can be extracted from the audio segment to be identified.

[0115] Step 203: Based on the second acoustic features and the target language classification model, determine the language category to which the audio data to be identified belongs.

[0116] The target language classification model includes an N-layer encoder, a fusion module, and a language identification module, where N is an integer greater than 1.

[0117] The N-layer encoder is used to: determine the corresponding second coding feature based on the second acoustic feature;

[0118] The fusion module is used to fuse the second coding features to obtain the target fused feature;

[0119] The language identification module is used to determine the language category of the audio data to be identified based on the target fusion features.

[0120] As an example, a target language classification model can be generated by a smart device or server through... Figure 1 The method shown is used for training, or the data is obtained by a smart device or server from device A, where device A is for execution. Figure 1 The device that trains the target language classification model using the method shown.

[0121] Based on the second acoustic features and the target language classification model, the language category of the audio data to be identified is determined, including:

[0122] In one scenario, if a target language classification model is deployed in the smart device or server, the second acoustic feature is input into the target language classification model to obtain the language category to which the audio data to be identified belongs.

[0123] In another scenario, if the smart device or server does not have a target language classification model deployed, the smart device will send the second acoustic feature or the audio data to be identified to a device that has a target language classification model deployed. Then, the device with the target language classification model will determine the language category of the audio data to be identified based on the second acoustic feature or the audio data to be identified, and send it to the smart device or server.

[0124] The language classification model in this application combines short audio enhancement and model layer fusion technology. When performing language identification on the audio data to be identified, the language classification model can improve the accuracy and robustness of language identification, and is especially suitable for multilingual speech recognition systems and related application scenarios.

[0125] As an example, and not a limitation, in practical applications, a target language classification model can be used to perform real-time language recognition on voice commands acquired in real time. The language recognition results can then be used to optimize subsequent speech recognition processes (such as selecting language-specific speech recognition language models and corresponding language dictionaries, and determining the language to be switched in the downstream NLU). By combining real-time requirements with the characteristics of in-vehicle scenarios, a seamless integration of multilingual recognition and voice interaction can be achieved.

[0126] In this application embodiment, when performing language identification on the audio data to be identified using the target language classification model, the global and local features of the audio data to be identified can be captured, thereby improving the accuracy and robustness of language identification.

[0127] It is worth noting that the above Figure 1 The illustrated embodiments and Figure 2 When the illustrated embodiments are executed by the same smart device (server), the smart device (server) can execute first. Figure 1 The illustrated embodiment, and then, in the case of an audio category recognition requirement, is executed again. Figure 2 The example shown.

[0128] It should be noted that any solutions not detailed in this embodiment can be found in the descriptions of the relevant solutions in the foregoing embodiments, and will not be repeated here.

[0129] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0130] It is worth noting that the intelligent device or server that performs the language classification model training method and the intelligent device or server that performs the language recognition method can be the same device or different devices.

[0131] For example, the server executes the language classification model training method, while the smart device performs language recognition based on the target language classification model obtained by the server through the language classification model training method.

[0132] For example, after the server trains the target language classification model in the above manner, it sends the target language classification model to the smart device, or the smart device obtains the target language classification model from the server, and then the smart device executes the above language recognition method based on the target language classification model.

[0133] Alternatively, the devices that perform the language classification model training method and the language recognition method are different servers. For example, server 1 performs the language classification model training method and then sends the target language classification model to server 2, which then performs the language recognition method based on the target language classification model.

[0134] Alternatively, the devices that perform the language classification model training method and the devices that perform the language recognition method may be different smart devices. For example, smart device 1 performs the language classification model training method and then sends the target language classification model to smart device 2. Smart device 2 then performs the language recognition method based on the target language classification model.

[0135] Alternatively, after obtaining the target language classification model by executing the above-mentioned language classification model training method, server 1 or smart device 1 stores the target language classification model in server 1 or smart device 1. When server 2 or smart device 2 needs to perform language identification on the audio data to be identified, server 2 or smart device 2 sends the audio data to be identified to server 1 or smart device 1. Then, server 1 or smart device 1 performs language category identification on the audio data to be identified according to the target language classification model, and then sends the identification result to server 2 or smart device 2.

[0136] Corresponding to the language classification model training method described in the above embodiments, Figure 3 A schematic diagram of the structure of a language classification model training device provided in an embodiment of this application is shown. This language classification model training device can be applied to a smart device or can function as a smart device. For ease of explanation, only the parts relevant to the embodiments of this application are shown.

[0137] Reference Figure 3 The device includes:

[0138] The feature acquisition module 301 is used to acquire a first acoustic feature, which refers to the acoustic feature of an audio segment, the audio segment including at least one character;

[0139] The feature encoding module 302 is used to sequentially input the first acoustic features into the N-layer encoder of the initial language classification model to obtain the first encoded features output by each of the N layers of encoders, where N is an integer greater than 1;

[0140] The feature input module 303 is used to input each of the first encoded features into the fusion module of the initial language classification model to obtain the first fused feature;

[0141] The model training module 304 is used to train the initial language classification model based on the first fusion feature and the language recognition module of the initial language classification model to obtain the target language classification model, which is used to identify the language category of the audio data.

[0142] Optionally, the above-mentioned device further includes:

[0143] The sample acquisition module is used to acquire at least one audio training sample.

[0144] The timestamp acquisition module is used to acquire the timestamp of each character in each of the audio training samples;

[0145] The random cropping module is used to randomly crop any audio training sample based on the timestamp of each character in the audio training sample during each training session to obtain multiple audio segments.

[0146] Optionally, the above-mentioned random cropping module is specifically used for:

[0147] If the audio length of the audio training sample exceeds the length threshold, the audio training sample will be randomly cropped based on the timestamp of each character in the audio training sample during each training session.

[0148] Optionally, the language recognition module sequentially includes a first multi-head attention layer, a first attention statistical pooling layer, a second multi-head attention layer, and a second attention statistical pooling layer; the model training module 304 includes:

[0149] The first input unit is used to input the first fusion feature into the first multi-head attention layer to perform attention weighting on the time dimension of the first fusion feature to obtain the second fusion feature;

[0150] The second input unit is used to input the second fusion feature into the first attention statistical pooling layer to fuse the features at different time steps in the second fusion feature to obtain the third fusion feature;

[0151] The third input unit is used to input the third fusion feature into the second multi-head attention layer to perform attention weighting on the encoder dimension of the third fusion feature to obtain the fourth fusion feature;

[0152] The fourth input unit is used to input the fourth fusion feature into the second attention statistical pooling layer to fuse the encoded features output by the N layers of the encoder contained in the fourth fusion feature to obtain the fifth fusion feature;

[0153] The training unit is used to train the initial language classification model based on the fifth fusion feature to obtain the target language classification model.

[0154] Optionally, the above training unit is specifically used for:

[0155] When the parameters of the encoder in layer N are the first parameter and the parameters of the fusion module are the second parameter, the parameters of the language recognition module are updated to the third parameter based on the fifth fusion feature.

[0156] The initial language classification model obtained by setting the parameters of the N-layer encoder as the first parameter, the parameters of the fusion module as the second parameter, and the parameters of the language recognition module as the third parameter is used as the target language classification model.

[0157] Corresponding to the language identification method described in the above embodiments, Figure 4 A schematic diagram of the language recognition device provided in an embodiment of this application is shown. This language recognition device can be applied to a smart device or can function as a smart device. For ease of explanation, only the parts relevant to the embodiments of this application are shown.

[0158] Reference Figure 4 The device includes:

[0159] Data acquisition module 401 is used to acquire audio data to be recognized;

[0160] The acoustic acquisition module 402 is used to acquire a second acoustic feature, wherein the second acoustic feature refers to the acoustic feature of the audio data to be identified;

[0161] The language identification module 403 is used to obtain the language category of the audio data to be identified based on the second acoustic features and the target language classification model; the target language classification model is obtained by training the initial language classification model based on the above-mentioned language classification model training method.

[0162] The language classification model includes an N-layer encoder, a fusion module, and a language identification module, where N is an integer greater than 1.

[0163] The encoder in layer N is used to: determine the corresponding second coding feature based on the second acoustic feature;

[0164] The fusion module is used to: fuse each of the second encoded features to obtain the target fused feature;

[0165] The language identification module is used to: determine the language category of the audio data to be identified based on the target fusion features.

[0166] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0167] Figure 5 This is a schematic diagram of the structure of a smart device provided in an embodiment of this application. Figure 5 As shown, the smart device 5 in this embodiment includes: at least one processor 50 ( Figure 5 (Only one is shown in the diagram), memory 51, and computer program 52 stored in said memory 51 and executable on said at least one processor 50, wherein said processor 50 executes said computer program 52 to implement the steps in any of the above method embodiments.

[0168] The smart device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will understand that... Figure 5 This is merely an example of smart device 5 and does not constitute a limitation on smart device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0169] The processor 50 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0170] In some embodiments, the memory 51 may be an internal storage unit of the smart device 5, such as a hard drive or memory of the smart device 5. In other embodiments, the memory 51 may be an external storage device of the smart device 5, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the smart device 5. Furthermore, the memory 51 may include both internal and external storage units of the smart device 5. The memory 51 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 51 can also be used to temporarily store data that has been output or will be output.

[0171] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0172] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / smart device, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0173] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0174] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0175] In the embodiments provided in this application, it should be understood that the disclosed devices / smart devices and methods can be implemented in other ways. For example, the device / smart device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0176] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0177] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for training a language classification model, characterized in that, The method comprises: obtaining a first acoustic feature, the first acoustic feature being an acoustic feature of an audio segment, the audio segment comprising at least one character of a language category; inputting the first acoustic feature into an N-layer encoder of an initial language classification model in sequence to obtain first encoding features output by the N-layer encoder respectively, N being an integer greater than 1; inputting each of the first encoding features into a fusion module of the initial language classification model to obtain a first fusion feature; training the initial language classification model based on the first fusion feature and a language recognition module of the initial language classification model to obtain a target language classification model, the target language classification model being used to identify a language category of audio data; the language recognition module comprises a first multi-head attention layer, a first attention statistical pooling layer, a second multi-head attention layer and a second attention statistical pooling layer in sequence; and the training of the initial language classification model based on the first fusion feature and the language recognition module of the initial language classification model to obtain the target language classification model comprises: inputting the first fusion feature into the first multi-head attention layer to perform attention weighting on a time dimension of the first fusion feature to obtain a second fusion feature; inputting the second fusion feature into the first attention statistical pooling layer to fuse features at different time steps in the second fusion feature to obtain a third fusion feature; inputting the third fusion feature into the second multi-head attention layer to perform attention weighting on an encoder dimension of the third fusion feature to obtain a fourth fusion feature; inputting the fourth fusion feature into the second attention statistical pooling layer to fuse encoding features output by the N-layer encoder included in the fourth fusion feature to obtain a fifth fusion feature; in a case where parameters of the N-layer encoder are first parameters and parameters of the fusion module are second parameters, updating parameters of the language recognition module to be third parameters based on the fifth fusion feature, the first parameters being parameters of the N-layer encoder trained, and the second parameters being parameters of the fusion module trained; using the initial language classification model obtained in a case where the parameters of the N-layer encoder are the first parameters, the parameters of the fusion module are the second parameters, and the parameters of the language recognition module are the third parameters as the target language classification model. 2.The language classification model training method of claim 1, wherein, Before the obtaining of the first acoustic feature, the method further comprises: obtaining at least one audio training sample; obtaining a timestamp of each of the characters in each of the audio training samples; for any of the audio training samples, performing random clipping on the audio training sample based on the timestamps of the characters in the audio training sample at each time of training to obtain a plurality of audio segments. 3.The language classification model training method of claim 2, wherein, The performing of the random clipping on the audio training sample based on the timestamps of the characters in the audio training sample at each time of training comprises: if an audio length of the audio training sample exceeds a length threshold, performing random clipping on the audio training sample based on the timestamps of the characters in the audio training sample at each time of training.

4. A language identification method characterized by, The method comprises: Obtaining to-be-recognized audio data; Obtaining a second acoustic feature, the second acoustic feature refers to an acoustic feature of the to-be-recognized audio data; Based on the second acoustic feature and the target language classification model, determine the language category to which the to-be-recognized audio data belongs; the target language classification model is obtained by training the initial language classification model based on the method of any one of claims 1 to 3; wherein the target language classification model comprises N layers of encoder, fusion module and language recognition module, N is an integer greater than 1; The N-layer encoder is used to determine the corresponding second encoding feature based on the second acoustic feature; The fusion module is used to fuse each second encoding feature to obtain a target fusion feature; The language recognition module is used to determine the language category to which the to-be-recognized audio data belongs based on the target fusion feature. 5.A language classification model training apparatus, characterized by comprising: Comprising: The feature acquisition module is used for acquiring a first acoustic feature, the first acoustic feature refers to the acoustic feature of the audio segment, and the audio segment includes at least one character of the language category; The feature encoding module is used for inputting the first acoustic feature into the N-layer encoder of the initial language classification model in turn to obtain the first encoding feature output by each of the N-layer encoder, N is an integer greater than 1; The feature input module is used for inputting each first encoding feature into the fusion module of the initial language classification model to obtain a first fusion feature; The model training module is used for training the initial language classification model based on the first fusion feature and the language recognition module of the initial language classification model to obtain a target language classification model, and the target language classification model is used for identifying the language category of audio data; The language recognition module comprises a first multi-head attention layer, a first attention statistical pooling layer, a second multi-head attention layer and a second attention statistical pooling layer in turn; The model training module comprises: The first input unit is used for inputting the first fusion feature into the first multi-head attention layer to perform attention weighting on the time dimension of the first fusion feature to obtain a second fusion feature; The second input unit is used for inputting the second fusion feature into the first attention statistical pooling layer to fuse the features at different time steps in the second fusion feature to obtain a third fusion feature; The third input unit is used for inputting the third fusion feature into the second multi-head attention layer to perform attention weighting on the encoder dimension of the third fusion feature to obtain a fourth fusion feature; The fourth input unit is used for inputting the fourth fusion feature into the second attention statistical pooling layer to fuse the encoding features output by the N-layer encoder contained in the fourth fusion feature to obtain a fifth fusion feature; The fifth input unit is configured to update the parameters of the language recognition module to third parameters based on the fifth fusion feature in a case where the parameters of the N-layer encoder are first parameters and the parameters of the fusion module are second parameters, the first parameters being the trained parameters of the N-layer encoder, and the second parameters being the trained parameters of the fusion module; and obtain the initial language classification model in a case where the parameters of the N-layer encoder are the first parameters, the parameters of the fusion module are the second parameters, and the parameters of the language recognition module are the third parameters, as the target language classification model.

6. A language identification apparatus characterized by comprising: Comprising: a data acquisition module configured to acquire to-be-recognized audio data; an acoustic acquisition module configured to acquire second acoustic features, the second acoustic features being acoustic features of the to-be-recognized audio data; a language recognition module configured to determine a language category to which the to-be-recognized audio data belongs based on the second acoustic features and a target language classification model, the target language classification model being obtained by training an initial language classification model based on the method of any one of claims 1 to 3; wherein the target language classification model comprises an N-layer encoder, a fusion module, and a language recognition module, N being an integer greater than 1; the N-layer encoder is configured to determine corresponding second encoding features based on the second acoustic features; the fusion module is configured to fuse each of the second encoding features to obtain target fusion features; the language recognition module is configured to determine the language category to which the to-be-recognized audio data belongs based on the target fusion features.

7. An intelligent device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program, so that the intelligent device implements the language classification model training method of any one of claims 1 to 3, or implements the language recognition method of claim 4.

8. A computer program product, characterised in that, The computer program is executed, so that the language classification model training method of any one of claims 1 to 3 is executed, or so that the language recognition method of claim 4 is executed.

Citation Information

Patent Citations

  • Multilingual identification method and system based on ASR information

    CN115064151A

  • Method and device for training language recognition model and language recognition

    CN115565522A