Language recognition method and related device, electronic device and storage medium

By extracting feature and projecting parameter processing of the spectrum graph, gender information interference in language recognition is eliminated, language recognition accuracy is improved, and the problems of weak language information and gender interference in language recognition are solved.

CN114283785BActive Publication Date: 2025-07-11ANHUI IFLYTEK INTELLIGENT SYST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111506374.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-07-11
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

There is a problem of weak language information in language recognition and is susceptible to gender information, resulting in low recognition accuracy.

Method used

By extracting the feature of the spectrogram, using projection parameters to eliminate interference information, especially gender information, and adjusting network parameters in combination with the training process of the language recognition network, gradually narrowing the differences between the predicted language and sample language, and improving the accuracy of language characteristics.

Benefits of technology

Effectively reduce interference information in language recognition, improve the accuracy of language recognition, ensure that language characteristics contain more language-related information, and reduce gender information interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283785B_ABST
    Figure CN114283785B_ABST
Patent Text Reader

Abstract

The present application discloses a language identification method, related devices, electronic devices, and storage media. The language identification method includes: extracting features from the spectrogram of the speech to be identified to obtain language features; projecting the language features using projection parameters to obtain projection features, where the projection parameters are used to reduce interference information in the language features, and the interference information at least includes gender information; and making a prediction based on the projection features to obtain the target language of the speech to be identified. The above solution can reduce the interference of interference information such as gender information on language identification as much as possible and improve the accuracy of language identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and particularly to a language identification method, related devices, electronic devices, and storage media. Background Art

[0002] With the deepening of international exchanges and cooperation, how to communicate without barriers across the country or even the world has become an issue that cannot be ignored. Facing such a huge language system, the languages that an individual can master are very limited. Therefore, the importance and status of automatic language identification technology have become increasingly prominent.

[0003] However, language information belongs to the weak information in speech signals, which cannot be directly and explicitly reflected like content information, and there are also interference information such as gender information, making language identification extremely difficult. In view of this, how to improve the accuracy of language identification has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide a language identification method, related devices, electronic devices, and storage media, which can improve the accuracy of language identification.

[0005] To solve the above technical problem, a first aspect of the present application provides a language identification method, including: extracting features from the spectrogram of the speech to be identified to obtain language features; projecting the language features using projection parameters to obtain projection features; wherein the projection parameters are used to reduce the interference information in the language features, and the interference information at least includes gender information; predicting based on the projection features to obtain the target language of the speech to be identified; wherein the target language is identified by a language identification network, and during the training process of the language identification network, the network parameters of the language identification network are adjusted based on the difference between the predicted language of the sample speech obtained based on the sample projection features and the sample language labeled for the sample speech, the sample projection features are obtained by projecting the sample language features of the sample speech using sample projection parameters, the sample projection parameters are obtained based on the intra-class feature difference of the sample speech with the same sample language regarding the sample language features, and the projection parameters are the sample projection parameters obtained when the language identification network converges after several rounds of training.

[0006] To solve the above technical problems, a language identification device is provided in the second aspect of the present application, including: a feature extraction module, a feature projection module, and a language prediction module. The feature extraction module is configured to extract features from the spectrogram of the speech to be identified to obtain language features; the feature projection module is configured to project the language features using projection parameters to obtain projection features; wherein the projection parameters are used to eliminate interference information in the language features, and the interference information at least includes gender information; the language prediction module is configured to make a prediction based on the projection features to obtain the target language of the speech to be identified; wherein the target language is identified by a language identification network. During the training process of the language identification network, the network parameters of the language identification network are adjusted based on the difference between the predicted language of the sample speech obtained based on the sample projection features and the sample language annotated for the sample speech. The sample projection features are obtained by projecting the sample language features of the sample speech using sample projection parameters, and the sample projection parameters are obtained based on the within-class feature difference of the sample speech with the same sample language regarding the sample language features, and the projection parameters are the sample projection parameters obtained when the language identification network converges after several rounds of training.

[0007] To solve the above technical problems, an electronic device is provided in the third aspect of the present application, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the language identification method in the first aspect above.

[0008] To solve the above technical problems, a computer-readable storage medium is provided in the fourth aspect of the present application, storing program instructions that can be run by a processor, and the program instructions are used to implement the language identification method in the first aspect above.

[0009] In the above solution, feature extraction is performed on the spectrogram of the speech to be recognized to obtain language features, and the language features are projected using projection parameters to obtain projection features. The projection parameters are used to eliminate interference information in the language features, and the interference information at least includes gender information. On this basis, prediction is performed based on the projection features to obtain the target language of the speech to be recognized, and the target language is recognized by a language recognition network. During the training process of the language recognition network, the network parameters of the language recognition network are adjusted based on the difference between the predicted language of the sample speech obtained based on the sample projection features and the sample language annotated for the sample speech. The sample projection features are obtained from the sample language features of the sample speech through sample projection parameters, and the sample projection parameters are obtained based on the intra-class feature difference of the sample speech with the same sample language regarding the sample language features. The projection parameters are the sample projection parameters obtained when the sample language recognition network converges after several rounds of training. Therefore, during the training process of the language recognition network, by gradually reducing the difference between the constrained predicted language and the sample language, the interference information in the sample projection features is also reduced as much as possible. Moreover, the sample projection features themselves are obtained from the sample language features through sample projection parameters, and the sample projection parameters are obtained based on the intra-class feature difference of the sample speech with the same sample language regarding the sample language features. Therefore, it can help to continuously reduce the difference between the sample language features of the sample speech with the same sample language, enabling the sample language features of the sample speech with the same sample language to contain as much language-related feature information as possible and as little interference information such as gender information as possible. Thus, it can improve the ability of the projection parameters obtained based on this to eliminate interference information, and further reduce the interference information such as gender information contained in the projection features as much as possible, which is beneficial to highlighting the language-related feature information in the projection features and further improving the accuracy of language recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a flowchart of an embodiment of the language recognition method of the present application;

[0011] Figure 2 is a framework diagram of an embodiment of the language recognition network;

[0012] Figure 3 is a framework diagram of an embodiment of the language recognition device of the present application;

[0013] Figure 4 is a framework diagram of an embodiment of the electronic device of the present application;

[0014] Figure 5 is a framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The following will combine the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of the present application.

[0016] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0017] In this article, the terms "system" and "network" are often used interchangeably. The term "and / or" in this article merely describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.

[0018] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the language identification method of the present application.

[0019] Specifically, it may include the following steps:

[0020] Step S11: Extract features from the spectrogram of the speech to be recognized to obtain language features.

[0021] In an implementation scenario, the specific language of the speech to be recognized may not be limited. For example, the specific language of the speech to be recognized may include but is not limited to: Chinese, English, French, etc., which are not limited here. In addition, the gender of the speaker of the speech to be recognized may also not be limited. That is, the speaker of the speech to be recognized can be male or female, which is not limited here.

[0022] In an implementation scenario, the speech to be recognized can be windowed to obtain a number of speech frames, and Fourier transforms are performed on each speech frame respectively to obtain the acoustic features of each speech frame, and then the acoustic features of each speech frame are stitched frame by frame to obtain a spectrogram. It should be noted that the above acoustic features may include but are not limited to: Filterbank, MFCC (Mel Frequency Cepstral Coefficient), etc., which are not limited here. In addition, the dimension of the acoustic features can be denoted as d, and its specific value is not limited here.

[0023] In an implementation scenario, in order to improve the language identification efficiency, a language identification network can be pre-trained, and the language identification network can include a feature extraction sub-network for extracting the first features related to the language of the spectrogram as language features.

[0024] In a specific implementation scenario, spectrograms can be feature-extracted to obtain a number of first feature maps. Statistical pooling can be performed on each of the first feature maps respectively to obtain first statistical features corresponding to each of the first feature maps. Then, based on the first statistical features corresponding to each of the first feature maps respectively, a first feature is obtained. In the above manner, by performing statistical pooling on each of the first feature maps, the feature dimension can be effectively reduced. Moreover, by combining the first statistical features corresponding to each of the first feature maps respectively to obtain the first feature, it is beneficial to improve the accuracy of the first feature.

[0025] In a specific implementation scenario, please refer to Figure 2 , Figure 2 which is a schematic framework diagram of an embodiment of a language identification network. As Figure 2 shown, the feature extraction sub-network may include a feature extraction layer, which may be, for example, a deep residual network (i.e., Deep Residual Network), etc., and is not limited herein.

[0026] In a specific implementation scenario, please continue to refer to Figure 2 , the feature extraction sub-network may further include a statistical pooling layer (Statistics Pooling, SP). Taking the resolution of the first feature map as M*N as an example, statistical pooling can be performed on each column of the first feature map to obtain the mean and variance of each column. Then, for each first feature map, after statistical pooling, a first statistical feature with a size of 2*N can be obtained. For C first feature maps, C first statistical features with a size of 2*N can be obtained. Other cases can be analogized accordingly and will not be exemplified one by one here.

[0027] In a specific implementation scenario, please continue to refer to Figure 2 , the feature extraction sub-network may further include a dimensionality reduction layer, which may include, but is not limited to, a full connection layer (Full Connection, FC), etc., and is not limited herein. Still taking the above-mentioned first statistical feature with a size of 2*N as an example, the first statistical features corresponding to each of the first feature maps can be stretched and spliced to form a high-dimensional feature of 2*N*C. On this basis, the dimensionality reduction layer can be used to compress it into a low-dimensional vector as the first feature. For ease of description, the first feature can be denoted as w0.

[0028] In an implementation scenario, as described above, in order to improve the language identification efficiency, a language identification network can be pre-trained, and the language identification network may include a gender suppression sub-network for extracting second features that are independent of gender from spectrograms as language features.

[0029] In a specific implementation scenario, the spectrogram can be divided into several sub-spectrograms from the frequency domain dimension, and feature extraction can be performed on the several sub-spectrograms respectively to obtain second feature maps corresponding to the respective sub-spectrograms, and feature dimensionality reduction can be performed on each of the second feature maps respectively to obtain dimensionality-reduced features corresponding to the respective second feature maps, and weights corresponding to the respective second feature maps can be obtained. On this basis, the dimensionality-reduced features corresponding to the respective second feature maps can be weighted respectively using the weights corresponding to the respective second feature maps to obtain second features. By dividing the spectrogram from the frequency domain dimension in the above manner, it is possible to suppress gender information for sub-spectrograms in different frequency bands respectively through a gender suppression network, further weight the dimensionality-reduced features after suppressing gender information in each frequency band, obtain second features, and thus be able to fully suppress gender information and improve the accuracy of the second features.

[0030] In a specific implementation scenario, from the frequency domain dimension, the spectrogram can be equally divided into several sub-spectrograms. Exemplarily, as Figure 2 shown, the spectrogram can be equally divided into three sub-spectrograms from the frequency domain dimension. For ease of description, the number of sub-spectrograms can be denoted as G. By equally dividing and cutting the spectrogram in the frequency domain dimension in the above manner, it is beneficial to suppress gender information in each frequency band subsequently and to refer to different frequency bands with emphasis respectively.

[0031] In a specific implementation scenario, please continue to refer to Figure 2 , the gender suppression sub-network can include several feature extraction layers for respectively extracting second feature maps corresponding to the respective sub-spectrograms. Exemplarily, taking G = 3 as an example, the gender suppression sub-network can include 3 feature extraction layers for respectively extracting second feature maps corresponding to 3 sub-spectrograms. Other cases can be inferred by analogy and will not be elaborated one by one here. In addition, the feature extraction layer can include, but is not limited to, a Time Delay Neural Network (TDNN), etc., and is not limited here.

[0032] In a specific implementation scenario, statistical pooling can be performed on each of the second feature maps respectively to obtain second statistical features corresponding to the respective second feature maps, and based on the second statistical features corresponding to the respective second feature maps, dimensionality-reduced features corresponding to the respective second feature maps can be obtained. Please refer to Figure 2, similarly to the foregoing description, the gender suppression network may also include a statistical pooling layer for performing statistical pooling on each of the second feature maps respectively. For the specific process of statistical pooling, reference may be made to the foregoing relevant description and will not be elaborated herein. In addition, similarly to the foregoing description, the gender suppression network may also include a dimensionality reduction layer for reducing the dimension of the second statistical features corresponding to each of the second feature maps respectively. For the specific process of dimensionality reduction, reference may also be made to the foregoing relevant description and will not be elaborated herein. By performing statistical pooling on each of the second feature maps respectively in the above manner, the second statistical features corresponding to each of the second feature maps are obtained, and based on the second statistical features corresponding to each of the second feature maps respectively, the dimensionality-reduced features corresponding to each of the second feature maps are obtained, which can achieve feature dimensionality reduction through a series of operations such as statistical pooling, and is beneficial to reducing the complexity of feature dimensionality reduction. In addition, for the sake of convenience of description, the dimensionality-reduced feature corresponding to the i-th second feature map may be denoted as w gen-i .

[0033] In a specific implementation scenario, please continue to refer to Figure 2 , and each of the second feature maps can be respectively subjected to global average pooling (Global Average Pooling, GAP) to obtain the initial weights corresponding to each of the second feature maps respectively. For the sake of convenience of description, the initial weight corresponding to the i-th second feature map may be denoted as β gen-i . Further, as Figure 2 shown, the gender suppression sub-network may also include a normalization layer (such as, softmax, etc.), then after global average pooling, the initial weights corresponding to each of the second feature maps can be further normalized to obtain the weights corresponding to each of the second features respectively:

[0034]

[0035] In the above formula (1), α gen-i represents the weight corresponding to the i-th second feature map. Further, the weight α gen-i corresponding to the i-th second feature map can be used to perform weighted summation on the dimensionality-reduced feature w gen-i corresponding to the i-th second feature map to obtain the second feature w2:

[0036]

[0037] In one implementation scenario, as described above, in order to improve the efficiency of language recognition, a language recognition network can be pre-trained, and the language recognition network can include a feature extraction sub-network and a gender suppression sub-network. The feature extraction sub-network is used to extract the first feature related to the language of the spectrogram, while the gender suppression sub-network is used to extract the second feature unrelated to the gender of the spectrogram, and the language feature is obtained by fusing the first feature and the second feature. In the above manner, by fusing the first feature and the second feature, the accuracy of the language feature can be improved.

[0038] In a specific implementation scenario, weights can be set for the first feature and the second feature respectively in advance, and the first feature and the second feature are weighted using the pre-set weights to obtain the language feature. For ease of description, the weight corresponding to the first feature can be denoted as the first weight γ1, and the weight corresponding to the second feature can be denoted as the second weight γ2. It should be noted that the specific magnitudes of the first weight γ1 and the second weight γ2 can be set according to actual application requirements. For example, in the case of emphasizing the feature extraction sub-network, the first weight γ1 can be set to be greater than the second weight γ2, such as the first weight γ1 can be set to 0.7 and the second weight γ2 can be set to 0.3; or, in the case of emphasizing the gender suppression sub-network, the first weight γ1 can be set to be less than the second weight γ2, such as the first weight γ1 can be set to 0.3 and the second weight γ2 can be set to 0.7; or, in the case of equally emphasizing the feature extraction sub-network and the gender suppression sub-network, the first weight γ1 can be set to be equal to the second weight γ2, which is not limited herein.

[0039] In a specific implementation scenario, before fusion, the second feature can also be mapped to the same feature space as the first feature, and after mapping, the first feature and the mapped second feature are weighted and summed to obtain the language feature. Specifically, the second feature can be dimension-reduced through a fully connected layer to be mapped to the same feature space as the first feature.

[0040] Step S12: Project the language feature using the projection parameter to obtain the projection feature.

[0041] In the embodiments of the present disclosure, the projection parameter is used to eliminate the interference information in the language feature, and the interference information at least includes gender information. Specifically, the projection parameter can be represented in matrix form, and after projection, the original feature dimension of the language feature can remain unchanged. For ease of description, the projection parameter is denoted as P, then the projection feature w final can be calculated by multiplying the projection parameter by the language feature (i.e., Pw), where w represents the language feature.

[0042] In the embodiments of the present disclosure, the target language of the speech to be recognized can be recognized by the aforementioned language recognition network. During the training process of the language recognition network, the network parameters of the language recognition network can be adjusted based on the difference between the predicted language of the sample speech predicted based on the sample projection features and the sample language annotated for the sample speech. The sample projection features are obtained by projecting the sample language features of the sample speech through sample projection parameters. The sample projection parameters are obtained based on the within-class feature difference of the sample speech with the same sample language with respect to the sample language features, and the projection parameters are the sample projection parameters obtained when the language recognition network converges after several rounds of training. Taking the language recognition network converging after N rounds of training as an example, the sample projection parameters obtained during the Nth round of training can be used as the projection parameters adopted in the subsequent application stage. Other situations can be deduced by analogy and will not be exemplified one by one here.

[0043] In one implementation scenario, during the training process of the language recognition network, the size of the batch size (i.e., batch size) can be denoted as M, that is, a batch of training samples can contain M sample speeches, and each sample speech can be annotated with its sample language. Further, the sample languages annotated for these M sample speeches can all be the same or not all the same. Exemplarily, among them, there can be M - Q sample speeches annotated with the sample language "Chinese", and the other Q sample speeches are annotated with the sample language "English".

[0044] In one implementation scenario, the within-class feature difference of the sample speech with the same sample language with respect to the sample language features can be obtained, and the within-class feature differences corresponding to various sample languages are fused to obtain a fused feature difference, and based on the feature decomposition result of the fused feature difference, sample projection parameters are obtained. By the above method, by fusing the within-class feature differences corresponding to various sample languages to obtain a fused feature difference, and through feature decomposition to obtain sample projection parameters, it is beneficial to improve the accuracy of the sample projection parameters.

[0045] In a specific implementation scenario, the within-class feature difference is the variance of the sample speech with the same sample language with respect to the sample language features. Exemplarily, the variance corresponding to the l-th sample language can be denoted as W l , then the variance W l corresponding to the l-th sample language can be expressed as:

[0046]

[0047] In the above formula (3), represents the sample language features extracted from the j-th sample speech of the l-th sample language, denotes the average of the sample language features extracted from each sample voice of the l-th sample language, and I represents the total number of sample voices of the l-th sample language. It should be noted that for the specific process of extracting the sample language features from the sample voices, reference can be made to the aforementioned process of extracting language features, which will not be elaborated here. The above method calculates the variance of the sample voices with the same sample language with respect to the sample language features to obtain the within-class feature difference, which is beneficial to reducing the complexity of calculating the within-class feature difference.

[0048] In a specific implementation scenario, the fused feature difference is obtained by averaging the within-class feature differences corresponding to various sample languages. As mentioned above, a batch of training samples can contain M sample voices, and the total number of sample languages involved can be denoted as L. Then, the fused feature difference W can be expressed as:

[0049]

[0050] The above method obtains the fused feature difference by averaging the within-class feature differences corresponding to various sample languages, which is beneficial to reducing the complexity of calculating the fused feature difference.

[0051] In a specific implementation scenario, the feature decomposition result can include several eigenvalues and the eigenvectors corresponding to each eigenvalue respectively. Then, based on the numerical magnitudes of the eigenvalues, at least one eigenvector can be selected, and based on the selected eigenvectors, a perturbation matrix can be obtained. The perturbation matrix contains the aforementioned interference information, and based on the identity matrix and the perturbation matrix, the sample projection parameters can be obtained. Specifically, since the larger the eigenvalue, the greater the interference information contained in the corresponding eigenvector, the eigenvalues can be sorted in descending order, and the eigenvectors corresponding to the eigenvalues located before the preset order position can be selected. Exemplarily, the eigenvectors corresponding to the first r eigenvalues can be selected, so as to reduce the complexity of selecting eigenvectors; further, the selected eigenvectors can be directly combined to obtain the perturbation matrix R, so as to reduce the complexity of obtaining the perturbation matrix; on this basis, the eigenmatrix RR after multiplying the perturbation matrix R by its transpose matrix R T can be obtained T , and the identity matrix E is subtracted from the eigenmatrix RR TThe difference matrix is used as the sample projection parameter P. Therefore, through simple operations such as matrix multiplication and subtraction, the interference information in the sample projection parameter can be removed as much as possible. In the above manner, the eigen-decomposition result includes several eigenvalues and the eigenvectors corresponding to each eigenvalue respectively. On this basis, based on the numerical magnitudes of the eigenvalues, at least one eigenvector is selected, and based on the selected eigenvector, a perturbation matrix is obtained, and the perturbation matrix contains interference information. And based on the identity matrix and the perturbation matrix, the sample projection parameter is obtained. Therefore, through simple operations such as eigen-decomposition and feature selection, the sample projection parameter can be obtained, which is beneficial to reducing the complexity of calculating the sample projection parameter.

[0052] In an implementation scenario, similar to the language feature, the sample language feature can also be obtained by fusing the first sample feature and the second sample feature. The first sample feature can be extracted from the sample spectrogram by the feature extraction sub-network, and the second sample feature can be extracted from the sample spectrogram by the gender suppression sub-network. The specific process can refer to the extraction processes of the aforementioned first feature and second feature respectively, which will not be elaborated here. Based on this, gender prediction can be performed based on the second sample feature to obtain the predicted probability values of the sample speaker of the sample speech being of various genders, and based on the difference between the predicted language of the sample speech obtained based on the sample projection feature and the sample language annotated for the sample speech, a first loss is obtained, and based on the difference between the predicted probability values of the sample speaker of the sample speech being of various genders and the target probability values of various genders, a second loss is obtained, and the target probability values of various genders are the same. On this basis, based on the first loss and the second loss, the network parameters of the language recognition network are adjusted. In the above manner, on the one hand, during the training process of the language recognition network, gender prediction is used to assist language prediction, which is beneficial to improving the network performance of the language recognition network. On the other hand, the second loss is measured by the difference between the predicted probability values of the sample speaker of the sample speech being of various genders and the target probability values of various genders, and the target probability values of various genders are the same, which can enable the gender suppression sub-network to suppress the interference of gender information as much as possible and further promote the recognition accuracy of the language recognition network.

[0053] In a specific implementation scenario, the specific process of obtaining the sample spectrogram of the sample speech can refer to the process of obtaining the spectrogram of the speech to be recognized, which will not be elaborated here.

[0054] In a specific implementation scenario, the language recognition network may further include a language prediction subnetwork. After obtaining the sample projection features, the sample projection features may be input into the language prediction subnetwork to obtain the predicted probability values ​​of the sample speech being in several preset languages, thereby being able to calculate the first loss based on the sample language marked by the sample speech and the predicted probability values ​​of the sample speech being in several preset languages ​​using a loss function such as cross entropy. For the specific calculation process, please refer to the technical details of loss functions such as cross entropy, which will not be described here. In addition, the language prediction subnetwork may specifically include a fully connected layer, a softmax layer, etc., which are not limited here.

[0055] In a specific implementation scenario, as mentioned above, the target probability values ​​of various genders can be set to be the same. Taking various genders including "male" and "female" as an example, the target probability value of "female" can be set to 0.5, and the target probability value of "male" can also be set to 0.5, that is, the second sample feature extracted by the gender suppression sub-network cannot distinguish whether the sample speaker of the sample speech is "male" or "female", so through the above optimization, the gender suppression sub-network can reduce gender information as much as possible during the feature extraction process, and only retain feature information related to the language.

[0056] In a specific implementation scenario, the first loss and the second loss may be weighted to obtain a weighted loss, on which basis the network parameters of the language recognition network may be adjusted by an optimization method such as gradient descent. The specific parameter adjustment process may refer to the technical details of optimization methods such as gradient descent, which will not be described in detail here.

[0057] In a specific implementation scenario, as described above, the above training process can be performed for several rounds until the language recognition network training converges. During this process, the intra-class feature differences of sample speech with the same sample language with respect to the sample language features become smaller and smaller, and the sample projection parameters gradually become stable. When the language recognition network training converges, it can be considered that the sample projection parameters at this time have been able to reduce interference information such as gender information as much as possible.

[0058] Step S13: Predicting based on the projection features to obtain the target language of the speech to be recognized.

[0059] Specifically, as mentioned above, the language recognition network can further include a language prediction subnetwork. After obtaining the projection features, the projection features can be input into the language prediction subnetwork to obtain the predicted probability values ​​of the speech to be recognized being several preset languages. On this basis, the preset language corresponding to the maximum prediction probability value can be selected as the target language of the speech to be recognized.

[0060] In the above solution, the spectrogram of the speech to be recognized is subjected to feature extraction to obtain language features, and the language features are projected using projection parameters to obtain projection features. The projection parameters are used to reduce the interference information in the language features, and the interference information at least includes gender information. On this basis, prediction is performed based on the projection features to obtain the target language of the speech to be recognized, and the target language is recognized by a language recognition network. During the training process of the language recognition network, the network parameters of the language recognition network are adjusted based on the difference between the predicted language of the sample speech obtained by predicting based on the sample projection features and the sample language annotated for the sample speech. The sample projection features are obtained from the sample language features of the sample speech through sample projection parameters, and the sample projection parameters are obtained based on the intra-class feature differences of the sample speech with the same sample language regarding the sample language features. The projection parameters are the sample projection parameters obtained when the sample language recognition network converges after several rounds of training. Therefore, during the training process of the language recognition network, by gradually reducing the difference between the constrained predicted language and the sample language, the interference information in the sample projection features can also be reduced as much as possible. Moreover, the sample projection features themselves are obtained from the sample language features through the sample projection parameters, and the sample projection parameters are obtained based on the intra-class feature differences of the sample speech with the same sample language regarding the sample language features. Therefore, it can help to continuously reduce the differences between the sample language features of the sample speech with the same sample language, enabling the sample language features of the sample speech with the same sample language to contain as much language-related feature information as possible and as little interference information such as gender information as possible. Thus, it can enhance the ability of the obtained projection parameters to reduce interference information, and further can reduce the interference information such as gender information contained in the projection features as much as possible, which is beneficial to highlighting the language-related feature information in the projection features, and further can be beneficial to improving the accuracy of language recognition.

[0061] Please refer to Figure 3 , Figure 3It is a schematic framework diagram of an embodiment of the language identification device 30 of the present application. The language identification device 30 includes: a feature extraction module 31, a feature projection module 32, and a language prediction module 33. The feature extraction module 31 is configured to extract features from the spectrogram of the speech to be identified to obtain language features. The feature projection module 32 is configured to project the language features using projection parameters to obtain projection features. Among them, the projection parameters are used to eliminate interference information in the language features, and the interference information at least includes gender information. The language prediction module 33 is configured to make a prediction based on the projection features to obtain the target language of the speech to be identified. Among them, the target language is identified by a language identification network. During the training process of the language identification network, based on the difference between the predicted language of the sample speech obtained based on the sample projection features and the sample language labeled for the sample speech, the network parameters of the language identification network are adjusted. The sample projection features are obtained by projecting the sample language features of the sample speech through sample projection parameters. The sample projection parameters are obtained based on the intra-class feature differences of the sample speech with the same sample language regarding the sample language features, and the projection parameters are the sample projection parameters obtained when the language identification network converges after several rounds of training.

[0062] In the above solution, during the training process of the language identification network, by gradually reducing the difference between the constrained predicted language and the sample language, the interference information in the sample projection features can also be reduced as much as possible. And the sample projection features themselves are obtained by projecting the sample language features through the sample projection parameters, and the sample projection parameters are obtained based on the intra-class feature differences of the sample speech with the same sample language regarding the sample language features. Therefore, it can help to continuously reduce the differences between the sample language features of the sample speech with the same sample language, so that the sample language features of the sample speech with the same sample language contain as much language-related feature information as possible and as little interference information such as gender information as possible. Thus, it can enhance the ability of the obtained projection parameters to eliminate interference information, and further can reduce the interference information such as gender information contained in the projection features as much as possible, which is beneficial to highlighting the language-related feature information in the projection features, and further can be beneficial to improving the accuracy of language identification.

[0063] In some publicly disclosed embodiments, the language identification device 30 further includes an intra-class difference statistics module, configured to obtain the intra-class feature differences of the sample speech with the same sample language regarding the sample language features. The language identification device 30 further includes an intra-class difference fusion module, configured to fuse the intra-class feature differences corresponding to various sample languages to obtain a fused feature difference. The language identification device 30 further includes a projection parameter acquisition module, configured to obtain the sample projection parameters based on the feature decomposition result of the fused feature difference.

[0064] Therefore, by fusing the intra-class feature differences corresponding to various sample languages, the fused feature differences are obtained, and through feature decomposition, the sample projection parameters are obtained, which is beneficial to improving the accuracy of the sample projection parameters.

[0065] In some disclosed embodiments, the feature decomposition result includes a plurality of eigenvalues and eigenvectors respectively corresponding to each eigenvalue. The projection parameter acquisition module includes a vector selection sub-module for selecting at least one eigenvector based on the numerical magnitudes of the respective eigenvalues; the projection parameter acquisition module includes a vector combination sub-module for obtaining a perturbation matrix based on the selected eigenvectors; wherein the perturbation matrix contains interference information; the projection parameter acquisition module includes a matrix operation sub-module for obtaining the sample projection parameters based on the identity matrix and the perturbation matrix.

[0066] Therefore, the feature decomposition result includes a plurality of eigenvalues and eigenvectors respectively corresponding to each eigenvalue. On this basis, at least one eigenvector is selected based on the numerical magnitudes of the respective eigenvalues, and a perturbation matrix is obtained based on the selected eigenvectors, and the perturbation matrix contains interference information, and the sample projection parameters are obtained based on the identity matrix and the perturbation matrix. Therefore, through simple operations such as feature decomposition and feature selection, the sample projection parameters can be obtained, which is beneficial to reducing the complexity of calculating the sample projection parameters.

[0067] In some disclosed embodiments, the vector selection sub-module is specifically configured to sort the eigenvalues in descending order and select the eigenvectors respectively corresponding to the eigenvalues located before a preset order position; the vector combination sub-module is specifically configured to combine the selected eigenvectors to obtain a perturbation matrix; the matrix operation sub-module is specifically configured to obtain the eigenmatrix after multiplying the perturbation matrix by its transpose matrix, and use the difference matrix obtained by subtracting the eigenmatrix from the identity matrix as the sample projection parameters.

[0068] Therefore, selecting the eigenvectors respectively corresponding to the eigenvalues located in the front preset order position can reduce the complexity of selecting eigenvectors; further, directly combining the selected eigenvectors to obtain a perturbation matrix can reduce the complexity of obtaining the perturbation matrix; on this basis, obtaining the eigenmatrix after multiplying the perturbation matrix by its transpose matrix, and using the difference matrix obtained by subtracting the eigenmatrix from the identity matrix as the sample projection parameters, so the interference information in the sample projection parameters can be removed as much as possible through simple operations such as matrix multiplication and subtraction.

[0069] In some disclosed embodiments, the intra-class feature difference is the variance of the sample speech with the same sample language with respect to the sample language feature; and / or, the fused feature difference is obtained by averaging the intra-class feature differences corresponding to various sample languages.

[0070] Therefore, by calculating the variance of the sample speech with the same sample language about the sample language features, the intra-class feature differences are obtained, which is conducive to reducing the complexity of calculating the intra-class feature differences; and by averaging the intra-class feature differences corresponding to various sample languages ​​to obtain the fused feature differences, it is conducive to reducing the complexity of calculating the fused feature differences.

[0071] In some disclosed embodiments, the language identification network includes a feature extraction subnetwork and a gender suppression subnetwork; wherein the feature extraction subnetwork is used to extract a first feature of the spectrogram related to the language, and the gender suppression subnetwork is used to extract a second feature of the spectrogram that is independent of gender, and the language feature is obtained based on the fusion of the first feature and the second feature.

[0072] Therefore, by fusing the first feature and the second feature, the accuracy of the language feature can be improved.

[0073] In some disclosed embodiments, the sample language feature is obtained by fusing the first sample feature and the second sample feature, and the first sample feature is extracted from the sample spectrogram of the sample speech by the feature extraction subnetwork, and the second sample feature is extracted from the sample spectrogram by the gender suppression subnetwork; the language recognition device 30 also includes a gender prediction module, which is used to predict the gender based on the second sample feature, and obtain the predicted probability values ​​of the sample speakers of the sample speech for each gender; the language recognition device 30 also includes a first loss module, which is used to obtain a first loss based on the difference between the predicted language of the sample speech and the sample language annotated by the sample speech; the language recognition device 30 also includes a second loss module, which is used to obtain a second loss based on the difference between the predicted probability values ​​of the sample speakers for each gender and the target probability values ​​of each gender; wherein the target probability values ​​of each gender are the same; the language recognition device 30 also includes a parameter adjustment module, which is used to adjust the network parameters of the language recognition network based on the first loss and the second loss.

[0074] Therefore, on the one hand, during the training process of the language recognition network, gender prediction is used to assist language prediction, which is beneficial to improving the network performance of the language recognition network. On the other hand, the second loss is measured by the difference between the predicted probability values ​​of the sample speakers for each gender and the target probability values ​​for each gender. The target probability values ​​for each gender are the same, which can enable the gender suppression sub-network to suppress the interference of gender information as much as possible, and can further promote the recognition accuracy of the language recognition network.

[0075] In some disclosed embodiments, the feature extraction module 31 includes a first extraction sub-module for extracting first features. The first extraction sub-module includes a first feature map extraction unit for extracting features from the spectrogram to obtain a plurality of first feature maps; the first extraction sub-module includes a first statistical pooling unit for performing statistical pooling on the plurality of first feature maps respectively to obtain first statistical features corresponding to each of the first feature maps; the first extraction sub-module includes a first feature acquisition unit for obtaining first features based on the first statistical features corresponding to each of the first feature maps.

[0076] Therefore, by performing statistical pooling on each of the first feature maps, the feature dimension can be effectively reduced, and by combining the first statistical features corresponding to each of the first feature maps to obtain the first features, it is beneficial to improve the accuracy of the first features.

[0077] In some disclosed embodiments, the feature extraction module 31 includes a second extraction sub-module for extracting second features. The second extraction sub-module includes a spectrogram partitioning unit for partitioning the spectrogram into a plurality of sub-spectrograms from the frequency domain dimension; the second extraction sub-module includes a second feature map extraction unit for extracting features from each of the sub-spectrograms respectively to obtain second feature maps corresponding to each of the sub-spectrograms; the second extraction sub-module includes a feature dimensionality reduction unit for performing feature dimensionality reduction on each of the second feature maps respectively to obtain reduced-dimensional features corresponding to each of the second feature maps; the second extraction sub-module includes a weight acquisition unit for acquiring weights corresponding to each of the second feature maps; the second extraction sub-module includes a second feature acquisition unit for weighting the reduced-dimensional features corresponding to each of the second feature maps respectively using the weights corresponding to each of the second feature maps to obtain second features.

[0078] Therefore, by partitioning the spectrogram from the frequency domain dimension, the gender suppression network can respectively suppress gender information for sub-spectrograms of different frequency bands, further weight the reduced-dimensional features after suppressing gender information for each frequency band to obtain second features, and thus can fully suppress gender information and improve the accuracy of the second features.

[0079] In some disclosed embodiments, the feature dimensionality reduction unit is specifically configured to perform statistical pooling on each of the second feature maps respectively to obtain second statistical features corresponding to each of the second feature maps, and based on the second statistical features corresponding to each of the second feature maps, obtain reduced-dimensional features corresponding to each of the second feature maps.

[0080] Therefore, by performing statistical pooling on each of the second feature maps respectively to obtain second statistical features corresponding to each of the second feature maps, and based on the second statistical features corresponding to each of the second feature maps, obtaining reduced-dimensional features corresponding to each of the second feature maps, feature dimensionality reduction can be achieved through a series of operations such as statistical pooling, which is beneficial to reducing the complexity of feature dimensionality reduction.

[0081] In some disclosed embodiments, the spectrogram division unit is specifically configured to equally divide a spectrogram into a plurality of sub-spectrograms in the frequency domain dimension.

[0082] Therefore, by equally dividing and cutting the spectrogram in the frequency domain dimension, it is beneficial to subsequently suppress gender information in each frequency band respectively and to refer to different frequency bands with emphasis.

[0083] Please refer to Figure 4 , Figure 4 which is a schematic framework diagram of an embodiment of the electronic device 40 of the present application. The electronic device 40 includes a mutually coupled memory 41 and a processor 42. Program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above-described embodiment of the language recognition method. Specifically, the electronic device 40 may include, but is not limited to: a desktop computer, a laptop computer, a server, a mobile phone, a tablet computer, etc., which are not limited herein.

[0084] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above-described embodiment of the language recognition method. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip having signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.

[0085] In the above solution, during the training process of the language recognition network, by gradually reducing the difference between the constrained predicted language and the sample language, the interference information in the sample projection features can also be reduced as much as possible. The sample projection features themselves are obtained from the sample language features through sample projection parameters, and the sample projection parameters are obtained based on the within-class feature differences of the sample languages of the sample voices with the same sample language regarding the sample language features. Therefore, it can help to continuously reduce the differences between the sample language features of the sample voices with the same sample language, enabling the sample language features of the sample voices with the same sample language to contain as much language-related feature information as possible and as little interference information such as gender information as possible. Thus, it can enhance the ability of the projection parameters obtained based on this to eliminate interference information, and further can reduce as much as possible the interference information such as gender information contained in the projection features, which is beneficial to highlighting the language-related feature information in the projection features, and further can be beneficial to improving the accuracy of language recognition.

[0086] Please refer to Figure 5 , Figure 5 FIG. is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above-described embodiments of the language recognition method.

[0087] In the above solution, during the training process of the language recognition network, by gradually reducing the difference between the constrained predicted language and the sample language, the interference information in the sample projection features can also be reduced as much as possible. The sample projection features themselves are obtained from the sample language features through sample projection parameters, and the sample projection parameters are obtained based on the within-class feature differences of the sample languages of the sample voices with the same sample language regarding the sample language features. Therefore, it can help to continuously reduce the differences between the sample language features of the sample voices with the same sample language, enabling the sample language features of the sample voices with the same sample language to contain as much language-related feature information as possible and as little interference information such as gender information as possible. Thus, it can enhance the ability of the projection parameters obtained based on this to eliminate interference information, and further can reduce as much as possible the interference information such as gender information contained in the projection features, which is beneficial to highlighting the language-related feature information in the projection features, and further can be beneficial to improving the accuracy of language recognition.

[0088] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments, and their specific implementations can refer to the descriptions of the above method embodiments. For the sake of brevity, they will not be elaborated here.

[0089] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and their similarities can be referred to each other. For the sake of brevity, they will not be elaborated in this article.

[0090] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the apparatuses or units can be in electrical, mechanical or other forms.

[0091] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0092] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0093] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks or optical discs that can store program codes.

Claims

1. A language identification method, characterized in that, Including: Performing feature extraction on the spectrogram of the speech to be recognized to obtain language features; Projecting the language features by using projection parameters to obtain projection features; wherein, the projection parameters are used to reduce interference information in the language features, and the interference information at least includes gender information; Performing prediction based on the projection features to obtain the target language of the speech to be recognized; Wherein, the target language is recognized by a language recognition network. During the training process of the language recognition network, based on the difference between the predicted language of the sample speech obtained by predicting based on the sample projection features and the sample language labeled for the sample speech, the network parameters of the language recognition network are adjusted. The sample projection features are obtained by projecting the sample language features of the sample speech through sample projection parameters. The sample projection parameters are obtained based on the intra-class feature difference of the sample speech with the same sample language regarding the sample language features. During the training process of the language recognition network, the sample projection parameters gradually tend to be stable, and the projection parameters are the sample projection parameters obtained when the language recognition network converges after several rounds of training.

2. The method according to claim 1, wherein The obtaining steps of the sample projection parameters include: Obtaining the intra-class feature difference of the sample speech with the same sample language regarding the sample language features; Fusing the intra-class feature differences corresponding to various sample languages respectively to obtain a fused feature difference; Based on the eigen-decomposition result of the fused feature difference, obtaining the sample projection parameters.

3. The method according to claim 2, wherein The eigen-decomposition result includes several eigenvalues and eigenvectors respectively corresponding to each eigenvalue. Based on the eigen-decomposition result of the fused feature difference, obtaining the sample projection parameters includes: Based on the numerical magnitudes of the eigenvalues, selecting at least one eigenvector; Based on the selected eigenvectors, obtaining a perturbation matrix; wherein, the perturbation matrix contains the interference information; Based on the identity matrix and the perturbation matrix, obtaining the sample projection parameters.

4. The method according to claim 3, wherein The based on the numerical magnitudes of the eigenvalues, selecting at least one eigenvector includes: Sorting the eigenvalues in descending order and selecting the eigenvectors corresponding to the eigenvalues located before a preset order position; And / or, the based on the selected eigenvectors, obtaining a perturbation matrix includes: Combining the selected eigenvectors to obtain the perturbation matrix; And / or, the based on the identity matrix and the perturbation matrix, obtaining the sample projection parameters includes: Obtaining an eigenmatrix after multiplying the perturbation matrix by its transpose matrix, and using the difference matrix obtained by subtracting the eigenmatrix from the identity matrix as the sample projection parameters.

5. The method according to claim 2, wherein The intra-class feature difference is the variance of the sample speech with the same sample language regarding the sample language features; And / or, the fused feature difference is obtained by averaging the intra-class feature differences corresponding to various sample languages respectively.

6. The method according to claim 1, characterized in that, The language recognition network includes a feature extraction sub-network and a gender suppression sub-network; Among them, the feature extraction sub-network is used to extract the first feature related to the language of the spectrogram, the gender suppression sub-network is used to extract the second feature independent of gender of the spectrogram, and the language feature is obtained by fusing the first feature and the second feature.

7. The method according to claim 6, wherein The sample language feature is obtained by fusing the first sample feature and the second sample feature, and the first sample feature is extracted by the feature extraction sub-network from the sample spectrogram of the sample speech, and the second sample feature is extracted by the gender suppression sub-network from the sample spectrogram; The training steps of the language recognition network include: Performing gender prediction based on the second sample feature to obtain the predicted probability values of the sample speakers of the sample speech for each gender; Obtaining a first loss based on the difference between the predicted language of the sample speech and the sample language annotated for the sample speech, and obtaining a second loss based on the difference between the predicted probability values of the sample speakers for each gender and the target probability values for each gender; wherein, the target probability values for each gender are the same; Adjusting the network parameters of the language recognition network based on the first loss and the second loss.

8. The method according to claim 6, characterized in that, The extraction steps of the first feature include: Performing feature extraction on the spectrogram to obtain a number of first feature maps; Performing statistical pooling on each of the number of first feature maps respectively to obtain the first statistical feature corresponding to each of the first feature maps; Obtaining the first feature based on the first statistical feature corresponding to each of the first feature maps.

9. The method according to claim 6, characterized in that, The extraction steps of the second feature include: Dividing the spectrogram into a number of sub-spectrograms from the frequency domain dimension; Performing feature extraction on each of the number of sub-spectrograms respectively to obtain the second feature map corresponding to each of the sub-spectrograms; Performing feature dimensionality reduction on each of the second feature maps respectively to obtain the dimensionality-reduced feature corresponding to each of the second feature maps, and obtaining the weight corresponding to each of the second feature maps; Using the weight corresponding to each of the second feature maps to weight the dimensionality-reduced feature corresponding to each of the second feature maps respectively to obtain the second feature.

10. The method according to claim 9, characterized in that, The performing feature dimensionality reduction on each of the second feature maps respectively to obtain the dimensionality-reduced feature corresponding to each of the second feature maps includes: Performing statistical pooling on each of the second feature maps respectively to obtain the second statistical feature corresponding to each of the second feature maps; Obtaining the dimensionality-reduced feature corresponding to each of the second feature maps based on the second statistical feature corresponding to each of the second feature maps.

11. The method according to claim 9, wherein The dividing the spectrogram into a number of sub-spectrograms from the frequency domain dimension includes: Dividing the spectrogram equally into the number of sub-spectrograms from the frequency domain dimension.

12. A language identification device, characterized in that, Includes: A feature extraction module for performing feature extraction on the spectrogram of the speech to be recognized to obtain a language feature; A feature projection module for projecting the language feature using projection parameters to obtain a projection feature; wherein, the projection parameters are used to eliminate the interference information in the language feature, and the interference information at least includes gender information; A language prediction module, configured to perform prediction based on the projection features to obtain the target language of the speech to be recognized; Wherein, the target language is recognized by a language recognition network. During the training process of the language recognition network, the network parameters of the language recognition network are adjusted based on the difference between the predicted language of the sample speech predicted based on the sample projection features and the sample language annotated for the sample speech. The sample projection features are obtained by projecting the sample language features of the sample speech through sample projection parameters, and the sample projection parameters are obtained based on the within-class feature difference of the sample speech with the same sample language with respect to the sample language features. During the training process of the language recognition network, the sample projection parameters gradually tend to be stable, and the projection parameters are the sample projection parameters obtained when the language recognition network converges after several rounds of training.

13. An electronic device, characterized in that, It includes a memory and a processor that are coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the language recognition method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, It stores program instructions that can be run by a processor, and the program instructions are used to implement the language recognition method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Language vector obtaining and language identification methods and related devices

    CN110164417A

  • Language recognition method for field of civil aviation air-ground call

    CN112216272A