Training Method for Speech Intention Recognition Model, Speech Intention Recognition Method and Device

The end-to-end speech intention recognition model is constructed through a two-stage training method, and the semantic extraction network is pre-trained with speech-text paired data, and the intent recognition network is trained with a small amount of manual annotation data, which solves the problems of low accuracy and high cost in traditional methods, and achieves efficient speech intention recognition.

CN114974224BActive Publication Date: 2025-07-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210767379.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-08
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Traditional speech intention recognition methods have low accuracy when processing declarative content and rely on a large number of manually labeled speech-intention pair training data, resulting in high costs.

Method used

Using a two-stage training method, firstly, the semantic extraction network is pre-trained using speech-text paired data, and then a small amount of manually labeled speech-intention paired data is used to train the intent recognition network to build an end-to-end speech intent recognition model.

Benefits of technology

It improves the accuracy of speech intention recognition, reduces the consumption of computing resources, and reduces the need for manual labeling of data, and reduces the training cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974224B_ABST
    Figure CN114974224B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method, a speech intent recognition method, and a device for a speech intent recognition model. The training method includes: obtaining a text sample and a first speech sample carrying a semantic label, wherein the first speech sample corresponds to the content of the text sample, and the semantic label is the text semantic feature of the text sample; using the first speech sample to pre-train a semantic extraction network in the speech intent recognition model to be trained, so as to obtain a pre-trained speech intent recognition model, wherein the pre-trained speech intent recognition model includes a pre-trained semantic extraction network and an intent recognition network to be trained; obtaining a second speech sample carrying an intent label; and using the second speech sample to train the pre-trained speech intent recognition model to obtain a trained speech intent recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech recognition technology, and in particular, to a method for training a speech intent recognition model, a speech intent recognition method, and a device. Background Art

[0002] Traditional speech intent recognition is mostly used in the field of human-computer interaction, such as smart home and mobile phone voice assistants. In these scenarios, the speech often contains simple directive content, and the speech content itself is basically equivalent to the intent. In this scenario, traditional speech intent recognition methods often apply a two-stage model. In the first step, it first passes through an ASR (Automatic Speech Recognition) model to convert the speech into text (ASR result). In the second step, the ASR result is then input into an NLU (Natural Language Understanding) model to output a predefined intent category.

[0003] With the expansion of applications, there is an increasing need to recognize the intent of a large amount of declarative speech content, but the declarative content itself is often not equivalent to the intent. For example, in an e-commerce live broadcast, the anchor may need to create an atmosphere by chatting about daily life and giving out red envelopes, and the content of the speech itself is not equivalent to the intent of creating an atmosphere. At this time, if the traditional two-stage model is applied, there will be a problem that more accurate speech recognition may not necessarily make the accuracy of the overall intent recognition better, so that the optimization goals of the two models are not necessarily consistent, and it is difficult to ensure the accuracy of speech intent recognition.

[0004] To adapt to such new scenarios, there is another method in related technologies, which is to construct an end-to-end E2E-SLU (End to End Spoken Language Understanding) model for speech content understanding. The speech signal is input, and the intent category is output. Compared with the traditional method, this method has a consistent global optimization goal and better accuracy. However, the intent recognition accuracy of this method highly depends on the data volume of the manually labeled speech-intent paired training data. Summary of the Invention

[0005] The present disclosure provides a method for training a speech intent recognition model, a speech intent recognition method, and a device, so as to at least solve the problem of highly depending on the data volume of manually labeled training data in related technologies, or may not solve any of the above problems.

[0006] According to a first aspect of the present disclosure, there is provided a method for training a speech intent recognition model, the training method comprising: obtaining a text sample and a first speech sample carrying a semantic label, wherein the first speech sample corresponds to the content of the text sample, and the semantic label is the text semantic feature of the text sample; using the first speech sample to pre-train a semantic extraction network in a speech intent recognition model to be trained, obtaining a pre-trained speech intent recognition model, wherein the pre-trained speech intent recognition model includes a pre-trained semantic extraction network and an intent recognition network to be trained; obtaining a second speech sample carrying an intent label; and using the second speech sample to train the pre-trained speech intent recognition model to obtain a trained speech intent recognition model.

[0007] Optionally, the using the first speech sample to pre-train a semantic extraction network in a speech intent recognition model to be trained includes: inputting the speech feature of the first speech sample into the semantic extraction network in the speech intent recognition model to be trained to obtain a first speech semantic feature of the first speech sample; determining a semantic similarity between the first speech semantic feature and the text semantic feature; and based on the semantic similarity, adjusting parameters of the semantic extraction network in the speech intent recognition model to be trained to pre-train the semantic extraction network in the speech intent recognition model to be trained.

[0008] Optionally, the determining a semantic similarity between the first speech semantic feature and the text semantic feature includes: respectively determining a speech representation vector corresponding to the first speech semantic feature and a text representation vector corresponding to the text semantic feature; and determining a semantic similarity between the speech representation vector and the text representation vector as the semantic similarity between the first speech semantic feature and the text semantic feature.

[0009] Optionally, the determining a speech representation vector corresponding to the speech semantic feature includes: performing pooling processing on the first speech semantic feature in a time dimension to obtain the speech representation vector.

[0010] Optionally, determining a text representation vector corresponding to the text semantic feature includes: using the first character of the text semantic feature as the text representation vector; or performing pooling processing on the text semantic feature in a time dimension to obtain the text representation vector.

[0011] Optionally, training the pre-trained speech intent recognition model using the second speech sample includes: inputting the speech features of the second speech sample into the pre-trained semantic extraction network to obtain the second speech semantic features of the second speech sample; inputting the second speech semantic features into the intent recognition network to be trained to obtain a predicted intent; determining a loss value based on the predicted intent and the intent label; and adjusting the parameters of the intent recognition network to be trained or adjusting the parameters of the pre-trained semantic extraction network and the intent recognition network to be trained based on the loss value to train the pre-trained speech intent recognition model.

[0012] According to a second aspect of the present disclosure, there is provided a speech intent recognition method, the speech intent recognition method including: obtaining a speech to be recognized; inputting the speech features of the speech to be recognized into a speech intent recognition model to obtain a predicted intent of the speech to be recognized, where the speech intent recognition model is trained using the above training method.

[0013] According to a third aspect of the present disclosure, there is provided a training apparatus for a speech intent recognition model, the training apparatus including: an acquisition unit configured to: acquire a text sample and a first speech sample carrying a semantic label, where the first speech sample corresponds to the content of the text sample, and the semantic label is the text semantic feature of the text sample; a first training unit configured to: pre-train a semantic extraction network in a speech intent recognition model to be trained using the first speech sample to obtain a pre-trained speech intent recognition model, where the pre-trained speech intent recognition model includes a pre-trained semantic extraction network and an intent recognition network to be trained; the acquisition unit is further configured to: acquire a second speech sample carrying an intent label; and a second training unit configured to: train the pre-trained speech intent recognition model using the second speech sample to obtain a trained speech intent recognition model.

[0014] Optionally, the first training unit is further configured to: input the speech features of the first speech sample into the semantic extraction network in the speech intent recognition model to be trained to obtain the first speech semantic features of the first speech sample; determine the semantic similarity between the first speech semantic features and the text semantic features; and adjust the parameters of the semantic extraction network in the speech intent recognition model to be trained based on the semantic similarity to pre-train the semantic extraction network in the speech intent recognition model to be trained.

[0015] Optionally, the first training unit is further configured to: respectively determine a speech representation vector corresponding to the first speech semantic feature and a text representation vector corresponding to the text semantic feature; determine a semantic similarity between the speech representation vector and the text representation vector as the semantic similarity between the first speech semantic feature and the text semantic feature.

[0016] Optionally, the first training unit is further configured to: perform pooling processing on the first speech semantic feature in the time dimension to obtain the speech representation vector.

[0017] Optionally, the first training unit is further configured to: use the first character of the sentence of the text semantic feature as the text representation vector; or perform pooling processing on the text semantic feature in the time dimension to obtain the text representation vector.

[0018] Optionally, the second training unit is further configured to: input the speech feature of the second speech sample into the pre-trained semantic extraction network to obtain a second speech semantic feature of the second speech sample; input the second speech semantic feature into the to-be-trained intent recognition network to obtain a predicted intent; determine a loss value according to the predicted intent and the intent label; based on the loss value, adjust parameters of the to-be-trained intent recognition network, or adjust parameters of the pre-trained semantic extraction network and the to-be-trained intent recognition network to train the pre-trained speech intent recognition model.

[0019] According to a fourth aspect of the present disclosure, there is provided a speech intent recognition device, including: an acquisition unit configured to acquire a speech to be recognized; a recognition unit configured to input the speech feature of the speech to be recognized into a speech intent recognition model to obtain a predicted intent of the speech to be recognized, where the speech intent recognition model is trained by using the above training method.

[0020] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, where when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method or the speech intent recognition method of the speech intent recognition model according to the present disclosure.

[0021] According to a sixth aspect of the present disclosure, there is provided a computer-readable storage medium, where when instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the training method or the speech intent recognition method of the speech intent recognition model according to the present disclosure.

[0022] According to a seventh aspect of the present disclosure, there is provided a computer program product including computer instructions which, when executed by at least one processor, implement a method for training a voice intent recognition model or a method for voice intent recognition according to the present disclosure.

[0023] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0024] For the method for training a voice intent recognition model, the method for voice intent recognition, and the device according to the embodiments of the present disclosure, the voice intent recognition model is an end-to-end model with a consistent global optimization objective, which is easy to ensure the recognition accuracy. Moreover, only one model is required, and the system architecture is simple, which can reduce the consumption of computing resources. In addition, by adopting a two-stage model training method, in the first stage, a large amount of easily obtainable voice-text paired data, that is, the first voice samples and text samples, are used to pre-train the semantic extraction network, which can improve the accuracy of the pre-trained semantic extraction network. In the second stage, by using a small amount of manually labeled voice-intent paired data, that is, the second voice samples carrying intent labels, the training of the voice intent recognition model can be achieved. Further, the cost of manually labeling training data is relatively high, so it is difficult to obtain a large amount of manually labeled paired data. According to the embodiments of the present disclosure, the preparation cost of training data can be reduced while ensuring the model accuracy, and the feasibility of model training can be improved.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0027] Figure 1 is a flowchart showing a method for training a voice intent recognition model according to an exemplary embodiment of the present disclosure.

[0028] Figure 2 is a schematic flowchart showing the first-stage training of a voice intent recognition model according to an exemplary embodiment of the present disclosure.

[0029] Figure 3 is a schematic flowchart showing the second-stage training of a voice intent recognition model according to an exemplary embodiment of the present disclosure.

[0030] Figure 4 is a flowchart showing a method for voice intent recognition according to an exemplary embodiment of the present disclosure.

[0031] Figure 5It is a block diagram showing a training device for a voice intent recognition model according to an exemplary embodiment of the present disclosure.

[0032] Figure 6 It is a block diagram showing a voice intent recognition device according to an exemplary embodiment of the present disclosure.

[0033] Figure 7 It is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0034] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order different from those illustrated or described here. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0036] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations including "any one of the several items", "any combination of several items of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0037] Traditional voice intent recognition is mostly used in the field of human-computer interaction, such as smart home and mobile phone voice assistants. In these scenarios, the voice often contains simple directive content, and the voice content itself is basically equivalent to the intent, and the changes in the specific expression methods are relatively few. For example, if the intent is "open the curtain", the corresponding voice content can be "open the curtain", "open the curtain", "open the curtain"; another example is that if the intent is "play music", the specific expression can be "play music", "play a song", "play some music".

[0038] In this scenario, traditional speech intent recognition methods often use a two-stage model. In the first step, the speech is converted into text (ASR result) through an ASR model. In the second step, the ASR result is input into an NLU model to output predefined intent categories.

[0039] With the rise of e-commerce live streaming, more and more merchants have started live streaming with goods while operating offline or online stores. During the e-commerce live streaming process, the live streamer mainly promotes the transaction process by explaining and demonstrating products. Among them, the live streamer's explanation can be further divided into intentions such as creating an atmosphere, introducing products, and answering viewers' questions. Different intentions often indicate different links in the process of selling products. Different live streaming platforms have different traffic distribution strategies for e-commerce live streaming. Therefore, it is a very important task to identify the explanation intentions of live streamers in e-commerce live streaming.

[0040] Different from the traditional scenario, for a new scenario such as e-commerce live streaming, the content of the speech to be recognized (such as the live streamer's explanation speech) is often long and declarative. The explanation content itself is not equal to the intention, and the same intention can have a large number of different specific expressions. At this time, if the traditional two-stage model is applied, there will be a problem that more accurate speech recognition does not necessarily make the overall intent recognition accuracy better, so that the optimization objectives of the two models are not necessarily the same, and it is difficult to guarantee the accuracy of speech intent recognition. At the same time, the system architectures of the two models require more computing resources.

[0041] To adapt to such new scenarios, there is another method in related technologies, which is to construct an end-to-end E2E-SLU model for speech content understanding. Input the speech signal and output the intent category. Compared with the traditional method, this method has two advantages: 1. The global optimization objective is the same and the accuracy is better; 2. Only one model is needed, the system architecture is simpler, and it usually consumes less computing resources. However, the intent recognition accuracy of this method depends heavily on the amount of manually labeled speech-intent paired training data.

[0042] The speech intent recognition model according to an exemplary embodiment of the present disclosure is also an end-to-end model, but includes a semantic extraction network and an intent recognition network. Among them, the semantic extraction network is used to extract the speech semantic features of the speech as the intermediate data of the model and is not output. The intent recognition network is used to obtain the estimated intent based on the speech semantic features. Therefore, the speech intent recognition model has all the advantages of an end-to-end model. In addition, the training method of the speech intent recognition model according to an exemplary embodiment of the present disclosure adopts a two-stage training method. In the first stage, a large amount of easily obtainable speech-text paired data is used to pre-train the semantic extraction network, which can improve the accuracy of the pre-trained semantic extraction network. In the second stage, the training of the speech intent recognition model can be achieved by using a small amount of manually labeled speech-intent paired data. Further, the cost of manually labeling training data is relatively high, so it is difficult to obtain a large amount of manually labeled paired data. According to the exemplary embodiment of the present disclosure, the preparation cost of training data can be reduced while ensuring the accuracy of the model, and the feasibility of model training can be improved.

[0043] Next, the training method of the speech intent recognition model, the training device of the speech intent recognition model, the speech intent recognition method, and the speech intent recognition device according to the exemplary embodiment of the present disclosure will be specifically described with reference to Figures 1 to 5 FIG.

[0044] Figure 1 FIG. is a flowchart showing the training method of the speech intent recognition model according to an exemplary embodiment of the present disclosure. It should be understood that the training method of the speech intent recognition model according to an exemplary embodiment of the present disclosure can be implemented in a terminal device such as a smart phone, a tablet computer, or a personal computer (PC), or can be implemented in a device such as a server.

[0045] Referring to Figure 1 , in step 101, a text sample and a first speech sample carrying a semantic label are obtained.

[0046] Among them, the first speech sample corresponds to the content of the text sample, constituting speech-text paired data. This kind of data is the training data of the ASR model, and speech recognition is a relatively mature technology, so this kind of paired data is relatively easy to obtain.

[0047] The semantic label is the text semantic feature of the text sample. Among them, to obtain the text semantic feature of the text sample, it is necessary to first extract the text feature of the text sample. Text feature extraction is a mature technology. As an example, the text feature can be extracted by looking up a table according to the text, and the present disclosure does not limit this. After obtaining the text feature, a text semantic extraction network can be used to extract the text semantic feature. The text semantic extraction network can adopt a masked language model pre-trained with a large amount of unsupervised text data, and there are already many open-source versions. As an example, the text semantic extraction network can be BERT (Bidirectional Encoder Representation from Transformers, a deep bidirectional self-attention network), or it can be replaced with a comparable RoBERTa (Robustly Optimized BERT Pretraining Approach, a robust deep bidirectional self-attention network) or MacBERT (MLM as correction BERT, a masked corrected deep bidirectional self-attention network), and the present disclosure does not limit this.

[0048] In step 102, using the first speech sample, pre-train the semantic extraction network in the speech intent recognition model to be trained to obtain a pre-trained speech intent recognition model. Among them, the pre-trained speech intent recognition model includes a pre-trained semantic extraction network and an intent recognition network to be trained. In other words, this involves the first-stage training of the speech intent recognition model, specifically the pre-training of the semantic extraction network. By using the text semantic feature of the text sample corresponding to the first speech sample as the semantic label of the first speech sample at this stage, it is possible to use the extraction of speech semantic features by imitating text semantic features as the extraction target of the semantic extraction network, realize the pre-training of the semantic extraction network, enable the semantic extraction network to learn rich semantic information, improve the extraction accuracy of the pre-trained semantic extraction network, and ensure the recognition accuracy of the subsequent intent recognition network. As an example, the semantic extraction network can adopt a deep self-attention network with convolution (Conformer). Conformer is a neural network with good effects in the speech field and has not been used in the speech intent recognition field.

[0049] Figure 2 It is a schematic flowchart showing the first-stage training of the speech intent recognition model according to an exemplary embodiment of the present disclosure.

[0050] Refer to Figure 2 , optionally, step 102 further includes but is not limited to the following three steps:

[0051] In the first step, input the speech features of the first speech sample into the semantic extraction network in the speech intent recognition model to be trained, and obtain the first speech semantic features of the first speech sample. Among them, speech feature extraction is a mature technology and will not be elaborated here.

[0052] In the second step, determine the semantic similarity between the first speech semantic features and the text semantic features. By determining the semantic similarity, a clear reference standard can be provided for training.

[0053] The second step may include, for example: respectively determining the speech representation vector corresponding to the first speech semantic features and the text representation vector corresponding to the text semantic features; determining the semantic similarity between the speech representation vector and the text representation vector as the semantic similarity between the first speech semantic features and the text semantic features. By converting the first speech semantic features and the text semantic features into sentence-level speech representation vectors and text representation vectors respectively, the similarity of the two vectors can be directly calculated, which can not only reduce the amount of calculation but also simplify the similarity calculation. As an example, cosine similarity can be used to calculate the semantic similarity. It should be understood that the sentence level is mentioned here because speech data and text data usually go through preprocessing and are segmented according to sentences, so a speech sample or a text sample often contains the content of a single sentence, and the speech representation vector of the speech sample or the text representation vector of the text sample is a sentence-level representation vector.

[0054] Optionally, the first speech semantic features output by the semantic extraction network may be in matrix form, with a feature vector extracted for each time unit (such as an audio frame), and the feature vectors of all time units of a first speech sample are pooled together to form a feature matrix. When determining the speech representation vector, pooling processing can be performed on the first speech semantic features in the time dimension, such as average pooling processing or max pooling processing, to convert the feature matrix into a feature vector and obtain the speech representation vector, thereby condensing the amount of information in the time dimension and helping to reduce the amount of feature data while retaining the basic information of the speech. Of course, other feasible conversion methods can also be adopted, and the present disclosure does not limit this.

[0055] It should be understood that when the text semantic extraction network is BERT, RoBERTa, or MacBERT, since such networks will add a character at the beginning of the sentence when outputting text semantic features to represent the text representation of this sentence, the first character of the text semantic features can be directly taken as the text representation vector at the sentence level, which helps to simplify the extraction of the text representation vector. For networks that do not output this first character, the aforementioned pooling process or other transformation processes can be similarly performed on the text semantic features, which can also condense the information volume in the time dimension and help reduce the feature data volume while retaining the basic information of the text. It should be understood that taking Chinese text as an example, a common way to process text is to first divide a sentence of text into multiple words through word segmentation, and then extract features for each word. At this time, a time unit of the text is a word, and the time dimension of the text can be regarded as the dimension of words. The text semantic feature vectors of all words in a text sample are pooled together to form a feature matrix as the text semantic features of this text sample.

[0056] In the third step, based on the semantic similarity, adjust the parameters of the semantic extraction network in the speech intent recognition model to be trained to pre-train the semantic extraction network in the speech intent recognition model to be trained. In this process, the parameters of the text semantic extraction network can be kept fixed without update, and with the goal of maximizing the semantic similarity, the parameters of the semantic extraction network are updated. After continuous update and iteration, the pre-training of the semantic extraction network can be completed, and a semantic extraction network that learns rich semantic information from the text semantic extraction network can be obtained. As an example, the parameter update method can be the commonly used backpropagation method in deep neural networks. The condition for the end of training can be that the calculation result converges, that is, the improvement rate of the semantic similarity is less than the threshold, that is, the change of the semantic similarity tends to be stable, or it can also be that the number of iteration steps reaches the set number of steps. The present disclosure does not limit this.

[0057] It should be understood that for the convenience of corresponding to reflect the extraction processes of the first speech semantic features and the text semantic features, the two are drawn in a parallel form in Figure 2 This is not a limitation on the execution time of the two. During actual training, the first speech semantic features and the text speech features can be extracted synchronously or asynchronously, which are all implementation manners of the present disclosure and fall within the protection scope of the present disclosure.

[0058] Returning to Figure 1 In step 103, a second speech sample carrying an intent label is obtained. What is obtained in this step is the manually labeled speech-intent paired data. Since the pre-training of the semantic extraction network is completed in steps 101 and 102, only a small amount of manually labeled data can be obtained in this step, reducing the data preparation cost.

[0059] In step 104, the pre-trained speech intent recognition model is trained using the second speech sample to obtain a trained speech intent recognition model.

[0060] Figure 3 FIG. is a schematic flow chart showing the second-stage training of a speech intent recognition model according to an exemplary embodiment of the present disclosure.

[0061] Referring to Figure 3 , first, the speech features of the second speech sample are extracted, and then the speech features are input into the pre-trained speech intent recognition model. The pre-trained speech intent recognition model is specifically obtained by importing a pre-trained semantic extraction network into the speech intent recognition model to be trained, that is, it includes a pre-trained semantic extraction network and an intent recognition network to be trained. The pre-trained speech intent recognition model can output an estimated intent. Based on the estimated intent and the intent label, a loss value can be determined, and then the parameters of the pre-trained speech intent recognition model can be adjusted according to the loss value to train the pre-trained speech intent recognition model.

[0062] Specifically, step 104 may include, for example: inputting the speech features of the second speech sample into the pre-trained semantic extraction network to obtain the second speech semantic features of the second speech sample; inputting the second speech semantic features into the intent recognition network to be trained to obtain an estimated intent; determining a loss value according to the estimated intent and the intent label; and adjusting the parameters of the intent recognition network to be trained or adjusting the parameters of the pre-trained semantic extraction network and the intent recognition network to be trained based on the loss value to train the pre-trained speech intent recognition model. The training method and training end condition of this step are the same as those of conventional model training and will not be elaborated here. It should be noted that after the semantic extraction network is pre-trained, it already has a certain accuracy. In step 104, it can be selectively determined whether to only adjust the parameters of the intent recognition network to be trained or to simultaneously adjust the parameters of the pre-trained semantic extraction network and the intent recognition network to be trained. The former can greatly reduce the number of parameters to be adjusted and improve the training efficiency. Although the number of parameters to be adjusted in the latter is relatively larger than that in the former, since the semantic extraction network has been pre-trained, the entire parameter adjustment calculation will be more likely to converge, that is, compared with the case without pre-training, the training time can be shortened, so it also helps to improve the training efficiency. As an example, the parameters of the intent recognition network can be adjusted only when the determined loss value is small, and the parameters of the semantic extraction network and the intent recognition network can be adjusted simultaneously when the determined loss value is large. Of course, other judgment conditions can also be configured, or the network whose parameters are to be adjusted can be determined manually by the staff, and the present disclosure does not limit this.

[0063] Figure 4It is a flowchart showing a speech intent recognition method according to an exemplary embodiment of the present disclosure. It should be understood that the speech intent recognition method according to the exemplary embodiment of the present disclosure can be implemented in a terminal device such as a smart phone, a tablet computer, or a personal computer (PC), or can also be implemented in a device such as a server.

[0064] Referring to Figure 4 , in step 401, the speech to be recognized is acquired. The speech to be recognized is the speech whose intent needs to be recognized, and each speech to be recognized can be a speech of a sentence.

[0065] In step 402, the speech features of the speech to be recognized are input into the speech intent recognition model to obtain the estimated intent of the speech to be recognized. Among them, the speech intent recognition model is trained by using the above-mentioned training method of the speech intent recognition model. Therefore, this speech intent recognition method has all the beneficial effects of the training method of the speech intent recognition model of the exemplary embodiment of the present disclosure, which will not be elaborated here. It should be understood that when the speech intent recognition model performs calculations, the speech semantic features of the speech to be recognized will be extracted by the semantic extraction network first, and then the estimated intent will be calculated by the intent recognition network.

[0066] Figure 5 It is a block diagram showing a training device for a speech intent recognition model according to an exemplary embodiment of the present disclosure. It should be understood that the training device for the speech intent recognition model according to the exemplary embodiment of the present disclosure can be implemented in a terminal device such as a smart phone, a tablet computer, or a personal computer (PC) in a software, hardware, or software-hardware combination manner, or can also be implemented in a device such as a server.

[0067] Referring to Figure 5 , the training device 500 for the speech intent recognition model includes an acquisition unit 501, a first training unit 502, and a second training unit 503.

[0068] The acquisition unit 501 can acquire a text sample and a first speech sample carrying a semantic label. Among them, the first speech sample corresponds to the content of the text sample, and the semantic label is the text semantic feature of the text sample.

[0069] The first training unit 502 can use the first speech sample to pre-train the semantic extraction network in the speech intent recognition model to be trained, and obtain a pre-trained speech intent recognition model. Among them, the pre-trained speech intent recognition model includes a pre-trained semantic extraction network and an intent recognition network to be trained.

[0070] Optionally, the first training unit 502 may also input the speech features of the first speech sample into the semantic extraction network in the speech intent recognition model to be trained, and obtain the first speech semantic features of the first speech sample; determine the semantic similarity between the first speech semantic features and the text semantic features; and based on the semantic similarity, adjust the parameters of the semantic extraction network in the speech intent recognition model to be trained, so as to pre-train the semantic extraction network in the speech intent recognition model to be trained.

[0071] Optionally, the first training unit 502 may also respectively determine the speech representation vector corresponding to the first speech semantic features and the text representation vector corresponding to the text semantic features; and determine the semantic similarity between the speech representation vector and the text representation vector as the semantic similarity between the first speech semantic features and the text semantic features.

[0072] Optionally, the first training unit 502 may also perform pooling processing on the first speech semantic features in the time dimension to obtain the speech representation vector.

[0073] Optionally, the first training unit 502 may also use the first character at the beginning of the sentence of the text semantic features as the text representation vector; or perform pooling processing on the text semantic features in the time dimension to obtain the text representation vector.

[0074] The acquisition unit 501 may also acquire a second speech sample carrying an intent label.

[0075] The second training unit 503 may use the second speech sample to train the pre-trained speech intent recognition model to obtain a trained speech intent recognition model.

[0076] Optionally, the second training unit 503 may also input the speech features of the second speech sample into the pre-trained semantic extraction network to obtain the second speech semantic features of the second speech sample; input the second speech semantic features into the intent recognition network to be trained to obtain a predicted intent; determine a loss value according to the predicted intent and the intent label; and based on the loss value, adjust the parameters of the intent recognition network to be trained, or adjust the parameters of the pre-trained semantic extraction network and the intent recognition network to be trained, so as to train the pre-trained speech intent recognition model.

[0077] It can be understood in combination with Figures 1 to 3 the training method of the speech intent recognition model described, the operations of the corresponding units in the training apparatus 500 of the speech intent recognition model.

[0078] Figure 6It is a block diagram showing a voice intent recognition device according to an exemplary embodiment of the present disclosure. It should be understood that the voice intent recognition device according to the exemplary embodiment of the present disclosure can be implemented in a software, hardware, or software-hardware combination manner in terminal devices such as smartphones, tablets, and personal computers (PCs), or can also be implemented in devices such as servers.

[0079] Referring to Figure 6 , the voice intent recognition device 600 includes an acquisition unit 601 and a recognition unit 602.

[0080] The acquisition unit 601 can acquire the voice to be recognized.

[0081] The recognition unit 602 can input the voice features of the voice to be recognized into the voice intent recognition model to obtain the estimated intent of the voice to be recognized, where the voice intent recognition model is trained using the above training method. It should be understood that when the recognition unit 602 runs the voice intent recognition model to perform calculations, the voice semantic features of the voice to be recognized will first be extracted by the semantic extraction network, and then the estimated intent will be calculated by the intent recognition network.

[0082] It can be understood in combination with referring to Figure 4 the voice intent recognition method described, the operations of the corresponding units in the voice intent recognition device 600.

[0083] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.

[0084] Referring to Figure 7 , the electronic device 700 includes at least one memory 701 and at least one processor 702. A set of computer-executable instructions is stored in the at least one memory 701. When the set of computer-executable instructions is executed by the at least one processor 702, the training method or the voice intent recognition method of the voice intent recognition model according to the exemplary embodiment of the present disclosure is executed.

[0085] As an example, the electronic device 700 can be a PC computer, a tablet device, a personal digital assistant, a smartphone, or other devices capable of executing the above instruction set. Here, the electronic device 700 does not have to be a single electronic device, and can also be a collection of any devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 700 can also be a part of an integrated control system or a system manager, or can be configured to be interconnected with a local or remote (e.g., via wireless transmission) interface as a portable electronic device.

[0086] In the electronic device 700, the processor 702 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.

[0087] The processor 702 may execute instructions or code stored in the memory 701, where the memory 701 may also store data. The instructions and data may also be sent and received over a network via the network interface device, where the network interface device may employ any known transmission protocol.

[0088] The memory 701 may be integrated with the processor 702, for example, by arranging RAM or flash memory within an integrated circuit microprocessor or the like. Additionally, the memory 701 may include a separate device, such as an external disk drive, a storage array, or other storage devices that may be used by any database system. The memory 701 and the processor 702 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 702 can read files stored in the memory.

[0089] In addition, the electronic device 700 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 700 may be connected to each other via a bus and / or a network.

[0090] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When the instructions in the computer-readable storage medium are run by at least one processor, at least one processor is caused to execute the training method or the speech intent recognition method of the speech intent recognition model according to the exemplary embodiment of the present disclosure. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0091] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided. The computer program product includes computer instructions that, when run by at least one processor, cause at least one processor to execute the training method or the speech intent recognition method of the speech intent recognition model according to the exemplary embodiment of the present disclosure.

[0092] A method for training a speech intent recognition model, a speech intent recognition method, and an apparatus according to an exemplary embodiment of the present disclosure. The speech intent recognition model is an end-to-end model with a consistent global optimization objective, which is easy to ensure the recognition accuracy. Moreover, only one model is required, and the system architecture is simple, which can reduce the consumption of computing resources. In addition, by adopting a two-stage model training method, in the first stage, a large amount of easily obtainable speech-text paired data, that is, the first speech samples and text samples, are used to pre-train the semantic extraction network, which can improve the accuracy of the pre-trained semantic extraction network. In the second stage, by using a small amount of manually annotated speech-intent paired data, that is, the second speech samples carrying intent labels, the training of the speech intent recognition model can be achieved. Further, the cost of manually annotating training data is relatively high, so it is difficult to obtain a large amount of manually annotated paired data. According to the exemplary embodiment of the present disclosure, the preparation cost of training data can be reduced while ensuring the model accuracy, and the feasibility of model training can be improved.

[0093] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0094] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for a voice intent recognition model, characterized in that The speech intent recognition model includes a semantic extraction network and an intent recognition network, and the training method includes: Obtain a text sample and a first speech sample with a semantic label, where the first speech sample corresponds to the content of the text sample, and the semantic label is the text semantic feature extracted from the text sample by a text semantic extraction network; Use the first speech sample to pre-train the semantic extraction network to obtain a pre-trained semantic extraction network, where the parameters of the text semantic extraction network are fixed and not updated during the pre-training process; Obtain a second speech sample with an intent label; Use the second speech sample to train the intent recognition network, or train the pre-trained semantic extraction network and the intent recognition network to obtain a trained speech intent recognition model.

2. The training method according to claim 1, characterized in that The using the first speech sample to pre-train the semantic extraction network includes: Input the speech feature of the first speech sample into the semantic extraction network to obtain the first speech semantic feature of the first speech sample; Determine the semantic similarity between the first speech semantic feature and the text semantic feature; Based on the semantic similarity, adjust the parameters of the semantic extraction network to pre-train the semantic extraction network.

3. The training method according to claim 2, characterized in that The determining the semantic similarity between the first speech semantic feature and the text semantic feature includes: Respectively determine the speech representation vector corresponding to the first speech semantic feature and the text representation vector corresponding to the text semantic feature; Determine the semantic similarity between the speech representation vector and the text representation vector as the semantic similarity between the first speech semantic feature and the text semantic feature.

4. The training method according to claim 3, characterized in that The determining the speech representation vector corresponding to the first speech semantic feature includes: Perform pooling processing on the first speech semantic feature in the time dimension to obtain the speech representation vector.

5. The training method according to claim 3, characterized in that, Determining the text representation vector corresponding to the text semantic feature includes: Use the first character of the sentence of the text semantic feature as the text representation vector; or Perform pooling processing on the text semantic feature in the time dimension to obtain the text representation vector.

6. The training method according to any one of claims 1 to 5, characterized in that, The using the second speech sample to train the intent recognition network, or train the pre-trained semantic extraction network and the intent recognition network includes: Input the speech feature of the second speech sample into the pre-trained semantic extraction network to obtain the second speech semantic feature of the second speech sample; Input the second speech semantic feature into the intent recognition network to obtain an estimated intent; Determine a loss value according to the estimated intent and the intent label; Based on the loss value, adjust the parameters of the intent recognition network to train the intent recognition network, or adjust the parameters of the pre-trained semantic extraction network and the intent recognition network to train the pre-trained semantic extraction network and the intent recognition network.

7. A method for speech intention recognition, characterized in that, The speech intent recognition method includes: Obtain the speech to be recognized; Input the speech features of the speech to be recognized into a speech intent recognition model to obtain the predicted intent of the speech to be recognized. Among them, the speech intent recognition model is trained by using the training method described in any one of claims 1 to 6.

8. A training device for a voice intention recognition model, characterized in that The speech intent recognition model includes a semantic extraction network and an intent recognition network. The training device includes: An acquisition unit configured to: acquire a text sample and a first speech sample carrying a semantic label, where the first speech sample corresponds to the content of the text sample, and the semantic label is a text semantic feature extracted from the text sample by a text semantic extraction network; A first training unit configured to: pre-train the semantic extraction network by using the first speech sample to obtain a pre-trained semantic extraction network, where the parameters of the text semantic extraction network are fixed and not updated during the pre-training process; The acquisition unit is further configured to: acquire a second speech sample carrying an intent label; A second training unit configured to: train the intent recognition network by using the second speech sample, or train the pre-trained semantic extraction network and the intent recognition network to obtain a trained speech intent recognition model.

9. The training device according to claim 8, characterized in that The first training unit is further configured to: Input the speech features of the first speech sample into the semantic extraction network to obtain the first speech semantic features of the first speech sample; Determine the semantic similarity between the first speech semantic features and the text semantic features; Based on the semantic similarity, adjust the parameters of the semantic extraction network to pre-train the semantic extraction network.

10. The training device according to claim 9, wherein, The first training unit is further configured to: Respectively determine the speech representation vector corresponding to the first speech semantic features and the text representation vector corresponding to the text semantic features; Determine the semantic similarity between the speech representation vector and the text representation vector as the semantic similarity between the first speech semantic features and the text semantic features.

11. The training device according to claim 10, wherein, The first training unit is further configured to: Perform pooling processing on the first speech semantic features in the time dimension to obtain the speech representation vector.

12. The training device according to claim 10, characterized in that, The first training unit is further configured to: Use the first character at the beginning of the text semantic features as the text representation vector; or Perform pooling processing on the text semantic features in the time dimension to obtain the text representation vector.

13. The training device according to any one of claims 8 to 12, characterized in that, The second training unit is further configured to: Input the speech features of the second speech sample into the pre-trained semantic extraction network to obtain second speech semantic features; Input the second speech semantic features into the intent recognition network to obtain a predicted intent; Determine a loss according to the predicted intent and the intent label; Based on the loss, adjust the parameters of the intent recognition network to train the intent recognition network, or adjust the parameters of the pre-trained semantic extraction network and the intent recognition network to train the pre-trained semantic extraction network and the intent recognition network.

14. A voice intention recognition device, characterized in that, The speech intent recognition device includes: An acquisition unit configured to: acquire a speech to be recognized; An identification unit, configured to: input the speech features of the speech to be identified into a speech intent recognition model to obtain the estimated intent of the speech to be identified, wherein the speech intent recognition model is trained by using the training method described in any one of claims 1 to 6.

15. An electronic device, characterized in that, Comprising: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method of the speech intent recognition model described in any one of claims 1 to 6 or the speech intent recognition method described in claim 7.

16. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the training method of the speech intent recognition model described in any one of claims 1 to 6 or the speech intent recognition method described in claim 7.

17. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by at least one processor, the training method of the speech intent recognition model described in any one of claims 1 to 6 or the speech intent recognition method described in claim 7 is implemented.