Training method, intention recognition method, device and medium for intention recognition model
By obtaining the frequency of intent categories and adjusting the model parameters, the problem of uneven learning of intent categories feature in the intent recognition model during the training process is solved, which improves the recognition accuracy and training efficiency of the model, and reduces storage requirements.
Patent Information
- Application Number
- CN202210325806.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-03-30
AI Technical Summary
During the training process of existing intent recognition models, due to random acquisition of sample data, the feature learning level of different consent categories is uneven, resulting in the model's low accuracy in identifying certain intent categories.
By obtaining the intent category frequency in the sample data, using the classification loss function to update the model parameters with the sample frequency, and initializing the weight parameters of the industry glossary in the initial network model, using pre-trained embedding layer and hidden layer, adjusting the sample weight and number of hidden layers, and combining with the data generator for training.
It improves the accuracy of the intent recognition model for identification of intent categories, shortens training time, reduces the number of model loading and unloading times, saves storage space, and improves recognition speed.
Smart Images

Figure CN114610851B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a training method, an intention recognition method, a device and a medium of an intention recognition model, and belongs to the field of computer technology.
Background Art
[0002] With the continuous development of natural language processing (NLP) technology and the rapid improvement of computer computing power, natural language processing technology has been widely applied to scenarios such as sentiment analysis, intention recognition, and machine translation.
[0003] Taking intention recognition as an example, traditional intention recognition methods include: first training an initial network model with sample data to obtain an intention recognition model. During the process of intention recognition, inputting target text data into the intention recognition model to obtain the intention category corresponding to the target text data.
[0004] However, during the process of training the initial network model with sample data, the sample data is usually randomly obtained, which may lead to the problem that the initial network model learns different degrees of features of different intention categories, resulting in the problem that the trained intention recognition model has low accuracy in recognizing certain intention categories.
Summary of the Invention
[0005] The present application provides a training method, an intention recognition method, a device and a medium of an intention recognition model, which can solve the problem that the initial network model learns different degrees of features of different intention categories, resulting in the problem that the trained intention recognition model has low accuracy in recognizing certain intention categories. The present application provides the following technical solutions:
[0006] In a first aspect, a training method of an intention recognition model is provided, where the intention recognition model is used to recognize the intention of text data, and the method includes:
[0007] Obtain first sample data, where the first sample data includes first text data and the intention category corresponding to the first text data;
[0008] Obtain the sample frequency corresponding to each intention category;
[0009] Input the first text data into a pre-created initial network model to obtain category prediction information;
[0010] Input the category prediction information, the intention category and the sample frequency into a classification loss function to obtain a classification loss value;
[0011] Update the model parameters of the initial network model based on the classification loss value to train the intention recognition model.
[0012] Optionally, obtaining the sample frequencies corresponding to each intent category includes:
[0013] For each target intent category, determining the number of category samples of the first text data with the intent category being the target intent category in the first sample data;
[0014] Determining the ratio of the number of category samples to the total number of samples of the first text data in the first sample data as the sample frequency.
[0015] Optionally, inputting the category prediction information, the intent category, and the sample frequency into a classification loss function to obtain a classification loss value, which is represented by the following formula:
[0016]
[0017] Where L is the classification loss value; y is the intent category; p(y) is the sample frequency corresponding to the intent category y; f y (x; θ) is the probability that the first text data indicated by the category prediction information is of category y; f i (x; θ) is the probability that the first text data indicated by the category prediction information is of the i-th intent category; p(i) is the sample frequency corresponding to the i-th intent category; K is the number of intent categories; x is the first text data; θ is the model parameter of the initial network model.
[0018] Optionally, the initial network model includes an embedding layer; the embedding layer is used to convert the words in the text data into word vectors;
[0019] Before inputting the first text data into a pre-created initial network model to obtain category prediction information, it further includes:
[0020] Obtaining an embedding layer pre-trained using second text data, where the weight parameter matrix of the pre-trained embedding layer corresponds to the initial words in a general vocabulary; the general vocabulary includes the words in the second text data; the second text data is different from the first text data;
[0021] Initializing the weight parameters corresponding to the industry words in the industry vocabulary based on the weight parameter matrix; the industry vocabulary includes the words in the first text data, and the industry vocabulary is partially the same as the general vocabulary;
[0022] Establishing the initial network model based on the weight parameters corresponding to the industry words.
[0023] Optionally, initializing the weight parameters corresponding to the industry words in the industry vocabulary based on the weight parameter matrix includes:
[0024] Obtain new words that are in the industry vocabulary table but not in the general vocabulary table;
[0025] For each new word, determine the first frequency of the new word in a preset general corpus;
[0026] For each initial word, determine the second frequency of the initial word in the general corpus;
[0027] Determine the target initial word corresponding to the second frequency with the smallest difference from the first frequency;
[0028] Initialize the weight parameter corresponding to the new word based on the weight parameter corresponding to the target initial word.
[0029] Optionally, the method further includes:
[0030] Determine the third frequency of each word appearing in a preset industry corpus;
[0031] Add words with the third frequency greater than a preset frequency threshold to the industry vocabulary table.
[0032] Optionally, the inputting the first text data into a pre-created initial network model includes:
[0033] Based on the first text data, the intent category, and the number of intent categories, obtain the sample weights corresponding to the first text data of each intent category;
[0034] Extract the first text data from the first sample data according to the sample weights and input it into a pre-created initial network model.
[0035] Optionally, the initial network model includes at least one hidden layer, and the number of hidden layers in the initial network model is less than the number of hidden layers in the BERT model.
[0036] In a second aspect, an intent recognition method is provided, and the method includes:
[0037] Obtain target text data;
[0038] Input the target text data into a pre-trained intent recognition model to obtain the intent category corresponding to the target text data;
[0039] Wherein, the intent recognition model is obtained by updating the model parameters of a pre-created initial network model based on a classification loss value; the classification loss value is obtained by inputting category prediction information, the intent category corresponding to the first text data, and the sample frequency corresponding to the intent category into a classification loss function; the category prediction information is obtained by inputting the first text data into the initial network model.
[0040] Optionally, inputting the target text data into a pre-trained intent recognition model includes:
[0041] Inputting the target text data into a pre-trained intent recognition model through a data generator.
[0042] In a third aspect, an electronic device is provided. The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement the training method of the intent recognition model provided in the first aspect, or the intent recognition method provided in the second aspect.
[0043] In a fourth aspect, a computer-readable storage medium is provided. A program is stored in the storage medium, and when the program is executed by a processor, it is used to implement the training method of the intent recognition model provided in the first aspect, or the intent recognition method provided in the second aspect.
[0044] The beneficial effects of this application at least include: by obtaining first sample data, the first sample data includes first text data and an intent category corresponding to the first text data; obtaining the sample frequency corresponding to each intent category; inputting the first text data into a pre-created initial network model to obtain category prediction information; inputting the category prediction information, the intent category, and the sample frequency into a classification loss function to obtain a classification loss value; updating the model parameters of the initial network model based on the classification loss value to train an intent recognition model, which can solve the problem that the initial network model has different learning degrees for the features of different intent categories, resulting in a lower accuracy of the trained intent recognition model for some intent categories; since the sample frequencies corresponding to different intent categories are input into the classification loss function, the classification loss function can fuse the distribution of data of different intent categories in the first sample data, so as to calculate a loss value with the prior knowledge of this distribution, which can make the initial network model have the same learning degree for the features of different intent categories from the level of the loss function. Therefore, the accuracy of the trained intent recognition model for intent recognition can be improved.
[0045] In addition, since the sample frequency is calculated based on the first sample data, the distribution of the first text data of each intent category determined based on the sample frequency is the same as the distribution of the first text data of each category in the first sample data, thereby improving the accuracy of the determined classification loss value and the accuracy of the trained intent recognition model for intent category recognition.
[0046] In addition, since the weight parameters corresponding to the industry words in the industry vocabulary are initialized based on the weight parameter matrix, and the initial network model is established based on the weight parameters corresponding to the industry words, the weight parameters in the pre-trained embedding layer can be fully utilized, the training time of the initial network model can be shortened, and the accuracy of the intention recognition model obtained by training in recognizing intention categories can be improved.
[0047] In addition, since the weight parameters of the words with the smallest word frequency difference are also similar, initializing the weight parameters of the newly added words based on the weight parameters of the initial words with the smallest word frequency difference can make the initial values of the weight parameters of the newly added words as close as possible to the actual values. Therefore, the training difficulty of the initial network model can be reduced, and the training speed of the initial network model can be improved.
[0048] In addition, since the words with the third frequency greater than the preset frequency threshold in the industry corpus are added to the industry vocabulary, the industry vocabulary can include industry feature words, so that the initial network model can learn the features of the industry feature words, and the accuracy of the intention recognition model obtained by training in intention recognition can be improved.
[0049] In addition, since the first text data is extracted from the first sample data according to the sample weight and input into the pre-created initial network model, the sample data of different intention categories can be balanced at the data level, the model training speed can be accelerated, and the accuracy of the intention recognition model obtained by training in recognizing intention categories can be improved.
[0050] In addition, since the number of hidden layers in the initial network model is less than that in the BERT model, and the model structure of the intention recognition model is the same as that of the initial network model, the speed of intention recognition using the intention recognition model obtained by training can be improved.
[0051] In addition, since the data generator can continuously maintain variables and return results during one call, the number of times of loading and unloading the intention recognition model during the intention recognition process can be reduced, and the speed of intention recognition can be improved.
[0052] In addition, since using the data generator to input the target data can achieve intention recognition while inputting, without generating a very large set of all the target text data at one time, the storage space of the central processing unit can be saved.
[0053] The above description is only an overview of the technical solution of the present application. In order to understand the technical means of the present application more clearly and implement it according to the content of the specification, the following is a detailed description with reference to the preferred embodiments of the present application and the accompanying drawings.
Description of the Drawings
[0054] Figure 1It is a flowchart of a method for training an intent recognition model provided by an embodiment of the present application;
[0055] Figure 2 It is a schematic diagram of the model structures of a BERT model and a RoBERTa - tiny - clue model provided by an embodiment of the present application;
[0056] Figure 3 It is a flowchart of an intent recognition method provided by an embodiment of the present application;
[0057] Figure 4 It is a block diagram of a training device for an intent recognition model provided by an embodiment of the present application;
[0058] Figure 5 It is a block diagram of an intent recognition method provided by an embodiment of the present application;
[0059] Figure 6 A block diagram of an electronic device provided by an embodiment of the present application.
Specific Embodiments
[0060] Next, with reference to the accompanying drawings and embodiments, the specific embodiments of the present application will be further described in detail. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0061] First, several terms related to the embodiments of the present application will be introduced.
[0062] Intent recognition: It is to extract the intent expressed in a text, that is, to recognize the intent of text data. Intent recognition mainly includes two steps. First, different intent categories need to be divided. Then, through the classification algorithm of natural language processing (NLP), the intent of the text is classified to obtain the corresponding intent category of the text.
[0063] BERT model: A type of large - scale pre - trained language model used for natural language processing. Such pre - trained models usually have high computational costs and large memory requirements, so it is difficult to execute on some devices with limited resources.
[0064] TinyBERT model: Using the BERT - base model as the teacher model, a small - scale student TinyBERT model is obtained through the method of knowledge distillation (transformer). The model size of TinyBERT is only 13.3% of that of BERT, the number of 12 - layer hidden layers is reduced to 4 layers, and the inference speed is 9.4 times that of BERT. The TinyBERT model includes: the ALBERT - tiny model and RoBERTa - tiny - clue.
[0065] ALBERT-tiny model: When pre-training, it uses a large-scale training corpus of 30G, adopts the Google vocabulary, and significantly reduces the vector dimensions such as the hidden layer dimension (hiden_size) by reducing the number of hidden layers from 12 to 4. The model size is 1 / 25 of the BERT model, and the training and inference speeds are about 10 times faster than those of the BERT model, with a slight decrease in accuracy.
[0066] RoBERTa-tiny-clue model: When pre-training, it uses a large-scale training corpus of 100G, adopts the clue_vocab vocabulary, and significantly reduces the vector dimensions such as the hidden layer dimension (hiden_size) by reducing the number of hidden layers from 12 to 4. The model size is 1 / 10 of the BERT model, and the training and inference speeds are 7-8 times faster than those of the BERT model.
[0067] Optionally, this application takes the training method and intention recognition method of the intention recognition model provided in each embodiment used in an electronic device as an example for illustration. The electronic device is a terminal or a server. The terminal can be a video conferencing terminal, a mobile phone, a computer, a tablet, a scanner, an electronic eye, etc. This embodiment does not limit the type of the electronic device.
[0068] Figure 1 It is a flowchart of the training method of the intention recognition model provided by an embodiment of this application. The intention recognition model is used to recognize the intention of text data. The method at least includes the following steps:
[0069] Step 101, obtain the first sample data.
[0070] Among them, the first sample data includes the first text data and the intention category corresponding to the first text data.
[0071] Optionally, the first text data can be Chinese or English. When the first text data is Chinese, it can be simplified Chinese or traditional Chinese. This embodiment does not limit the type of the first text data.
[0072] In an example, the first text data is a sentence composed of one or more than two characters.
[0073] Optionally, the first sample data is sample data in a specific field. For example: The first sample data is sample data in the video conferencing field. At this time, the first sample data is collected during the video conferencing process.
[0074] In this embodiment, obtaining the first sample data includes: obtaining the first text data; annotating the intention of the first text data to obtain the intention category corresponding to the first text data.
[0075] Optionally, obtaining the first text data includes: obtaining audio data; converting the audio data into text data.
[0076] In one example, the audio data is collected during a video conference.
[0077] In actual implementation, other methods can also be adopted to obtain the first text data. For example, scanning a book to obtain the first text data, or the first text data input by an input component. This embodiment does not limit the method of obtaining the first text data.
[0078] Optionally, the method of annotating the intention of the first text data can be manual annotation, or it can also be machine annotation. This embodiment does not limit the method of annotating the intention of the first text data.
[0079] Optionally, the classification methods of intention categories include but are not limited to the following:
[0080] First, classify based on the content of control. At this time, the intention categories can be classified into categories such as audio control, video control, and meeting process control. Specifically, audio control can be classified into categories such as increasing volume, decreasing volume, muting, enabling microphone permission, and disabling microphone permission; video control can be classified into categories such as switching cameras, enabling camera permission, and disabling camera permission; meeting process control can be classified into categories such as ending the meeting, starting the meeting, and joining the meeting.
[0081] For example: If the first text data is "Please increase the volume" or "The sound is too low and I can't hear clearly", at this time, the intention category corresponding to the first text data is "Increase volume".
[0082] Another example: If the first text data is "Please turn off the camera", "My speech is over", or "The meeting is over", the intention category corresponding to the first text data is "Turn off the camera".
[0083] Second, classify based on the content of the query. At this time, the intention categories can be classified into querying participants, querying the meeting agenda, querying the meeting duration, querying the meeting end time, etc.
[0084] For example: If the first text data is "Query the participants", "Is everyone here", or "Who else hasn't come yet", at this time, the intention category corresponding to the first text data is "Query the participants".
[0085] Another example: If the first text data is "What time does the meeting end" or "I have another meeting later", the intention category corresponding to the first text data is "Query the meeting end time".
[0086] In other embodiments, other ways can also be adopted to divide intention categories. For example, intention categories can be divided according to different fields. This embodiment does not limit the division method of intention categories and the types of intention categories.
[0087] Step 102: Obtain the sample frequencies corresponding to each intention category.
[0088] Optionally, the sample frequency can be calculated based on the first sample data, or it can also be a preset empirical value. This embodiment does not limit the method of obtaining the sample frequency corresponding to the intention category.
[0089] In one example, the sample frequency is calculated based on the first sample data. At this time, obtaining the sample frequencies corresponding to each intention category includes: for each target intention category, determining the category sample quantity of the first text data with the intention category being the target intention category in the first sample data; and determining the ratio of the category sample quantity to the total sample quantity of the first text data in the first sample data as the sample frequency.
[0090] Optionally, determining the ratio of the category sample quantity to the total sample quantity of the first text data in the first sample data as the sample frequency is represented by the following formula:
[0091]
[0092] Among them, p(y) is the sample frequency corresponding to the intention category y; n y is the category sample quantity of the intention category y; N is the total sample quantity of the first text data in the first sample data.
[0093] Since the sample frequency is calculated based on the first sample data, the distribution of the first text data of each intention category determined based on the sample frequency is the same as the distribution of the first text data of each category in the first sample data. Therefore, the accuracy of the determined classification loss value can be improved, and the accuracy of the intention recognition model trained for intention category recognition can be improved.
[0094] Step 103: Input the first text data into the initially created network model to obtain category prediction information.
[0095] In this embodiment, the initial network model includes an Embedding layer, an ENCODER hidden layer, and a classification layer. Among them, the Embedding layer is used to convert the words in the text data into word vectors to obtain the first word vector corresponding to the text data; the hidden layer is used to enhance the first word vector to obtain the second word vector corresponding to the text data; and the classification layer is used to classify the text data based on the second word vector to obtain the intention category corresponding to the text data.
[0096] Optionally, the initial network model can be built based on the TinyBERT model, or it can be built based on the BERT model, or it can also be built based on other natural language processing models. The type of the initial network model is not limited in this embodiment.
[0097] In one example, the initial network model includes at least one hidden layer, and the number of hidden layers in the initial network model is less than the number of hidden layers in the BERT model.
[0098] Since the number of hidden layers in the initial network model is less than the number of hidden layers in the BERT model, and the model structure of the intent recognition model is the same as that of the initial network model, the speed of intent recognition using the trained intent recognition model can be improved.
[0099] In this embodiment, an example is given with the initial network model built based on the RoBERTa-tiny-clue model.
[0100] Refer to Figure 2 , Figure 2 a is a schematic diagram of the model structure of the BERT model, Figure 2 b is a schematic diagram of the model structure of the RoBERTa-tiny-clue model. According to Figure 2 a, it can be seen that the BERT model includes an embedding layer and twelve hidden layers, while according to Figure 2 b, it can be seen that the RoBERTa-tiny-clue model includes an embedding layer and four hidden layers. The number of hidden layers of the RoBERTa-tiny-clue model is only one-third of the number of hidden layers of the BERT model. Therefore, building the initial network model based on the RoBERTa-tiny-clue model can improve the speed of intent recognition using the trained intent recognition model.
[0101] Since the sample data input into the model is usually randomly obtained during the process of training the initial network model with sample data, this will lead to slow training speed and low accuracy of the trained intent recognition model in recognizing certain intent categories during the process of training the initial network model with sample data.
[0102] Based on the above technical problems, in this embodiment, inputting the first text data into the pre-created initial network model includes: obtaining the sample weights corresponding to the first text data of each intent category based on the first text data, the intent categories, and the number of intent categories; extracting the first text data from the first sample data according to the sample weights and inputting it into the pre-created initial network model.
[0103] Since the first text data is extracted from the first sample data according to the sample weights and input into the pre-created initial network model, it is possible to balance the sample data of different intent categories at the data level, accelerate the model training speed, and improve the accuracy of the intent recognition model obtained by training in identifying intent categories.
[0104] In addition, since the influence of the number of intent categories on the sample weights is considered in the process of calculating the sample weights, it is possible to avoid the problem that the sample weights corresponding to the first text data of the intent categories with fewer sample numbers are too different from the sample weights corresponding to the first text data of other intent categories when the number of different intent categories is large. Therefore, it is possible to balance the sample weight differences of the first sample data of each intent category and improve the accuracy of the intent recognition model obtained by training in identifying intent categories.
[0105] Optionally, based on the first text data, the intent category, and the number of intent categories, obtaining the sample weights corresponding to the first text data of each intent category includes: for each target intent category, determining the category sample number of the first text data in the first sample data whose intent category is the target intent category; determining the total sample number of the first text data in the sample data; and obtaining the sample weights corresponding to the first text data of each intent category based on the total sample number, the category sample number, and the number of intent categories.
[0106] In one example, the sample weight of the first text data of the target intent category is negatively correlated with the category sample number corresponding to the target intent category, that is, the larger the category sample number corresponding to the target intent category, the smaller the sample weight of the first text data of the target intent category.
[0107] Correspondingly, the probability that the first text data is drawn is positively correlated with the sample weight, that is, the larger the sample weight, the greater the probability that the first text data is drawn.
[0108] Optionally, obtaining the sample weights corresponding to the first text data of each intent category based on the total sample number, the category sample number, and the number of intent categories is represented by the following formula:
[0109]
[0110] where w(y) is the sample weight of the first text data of the intent category y; n y is the category sample number of the intent category y; N is the total sample number; and M is the number of intent categories.
[0111] In actual implementation, other methods can also be used to calculate the sample weights corresponding to the first sample data of each intent category. For example, the ratio of the category sample number to the total sample number can be determined as the sample weight. The present embodiment does not limit the calculation method of the sample weights.
[0112] In the case where the number of the first samples is small, in order to improve the accuracy of the intention recognition model trained for recognizing intention categories, second text data different from the first text data can be used to pre-train the embedding layer and the hidden layer, and then an initial network model is established using the pre-trained embedding layer and hidden layer. In this way, since the pre-trained embedding layer and hidden layer have already pre-learned the knowledge of converting words into word vectors and enhancing the word vectors, during the process of training the initial network model, it is only necessary to fine-tune the parameters of the embedding layer and the hidden layer and train the parameters of the classification layer. This can reduce the training difficulty of the initial network model, reduce the number of the first samples required in the training process, and at the same time can also accelerate the training speed of the initial network model and improve the accuracy of the intention recognition model trained for recognizing intention categories.
[0113] Optionally, the weight parameter matrix of the pre-trained embedding layer corresponds to the initial words in the general vocabulary, and the general vocabulary includes the words in the second text data.
[0114] Optionally, the weight parameter matrix is a word vector matrix for storing word vectors corresponding to different words.
[0115] In one example, the weight parameter corresponding to a word is the word vector corresponding to the word in the word vector matrix.
[0116] Since the second text data is different from the first text data, and the industry vocabulary includes the words in the first text data, the industry vocabulary is partially the same as the general vocabulary, that is, the industry vocabulary includes words not in the general vocabulary. This will result in that the weight parameter matrix of the pre-trained embedding layer is not exactly the same as the weight parameter matrix of the embedding layer required by the initial network model. Therefore, during the process of establishing the initial network model based on the pre-trained embedding layer, it is necessary to initialize the weight parameter matrix of the pre-trained embedding layer.
[0117] However, the traditional initialization method is to randomly initialize the weight parameter matrix, which will lead to the problems of slow training speed of the intention recognition model and low accuracy of the intention recognition model trained for recognizing intention categories.
[0118] Based on the above technical problems, in this embodiment, before inputting the first text data into the initially created initial network model to obtain category prediction information, it further includes: obtaining an embedding layer pre-trained using the second text data, where the weight parameter matrix of the pre-trained embedding layer corresponds to the initial words in the general vocabulary; initializing the weight parameters corresponding to the industry words in the industry vocabulary based on the weight parameter matrix; and establishing an initial network model based on the weight parameters corresponding to the industry words.
[0119] Since the weight parameters corresponding to the industry words in the industry vocabulary table are initialized based on the weight parameter matrix, and an initial network model is established based on the weight parameters corresponding to the industry words, it is possible to make full use of the weight parameters in the pre-trained embedding layer, shorten the training time of the initial network model, and improve the accuracy of the intention recognition model obtained by training in recognizing intention categories.
[0120] In one example, initializing the weight parameters corresponding to the industry words in the industry vocabulary table based on the weight parameter matrix includes: obtaining the new words that are in the industry vocabulary table but not in the general vocabulary table; for each new word, determining the first frequency of the new word in the preset general corpus; for each initial word, determining the second frequency of the initial word in the general corpus; determining the target initial word corresponding to the second frequency with the smallest difference from the first frequency; initializing the weight parameters corresponding to the new words based on the weight parameters corresponding to the target initial word.
[0121] Since the weight parameters of the words with the smallest difference in word frequency are also similar, initializing the weight parameters of the new words based on the weight parameters of the initial words with the smallest difference in word frequency can make the initial values of the weight parameters of the new words as close as possible to the actual values. Therefore, the training difficulty of the initial network model can be reduced, and the training speed of the initial network model can be improved.
[0122] Optionally, the general corpus database can be a pre-collected corpus, or it can also be an open-source corpus. The type of the general corpus is not limited in this embodiment.
[0123] Optionally, the determination methods of the difference between the first frequency and the second frequency include but are not limited to the following several types:
[0124] First, determine the difference between the first frequency and the second frequency based on the absolute value of the difference between the first frequency and the second frequency. At this time, the smaller the absolute value of the difference between the second frequency and the first frequency, the smaller the difference between the second frequency and the first frequency.
[0125] Second, determine the difference between the first frequency and the second frequency based on the ratio between the first frequency and the second frequency. The closer the ratio between the second frequency and the first frequency is to 1, the smaller the difference between the second frequency and the first frequency.
[0126] In other embodiments, the difference between the first frequency and the second frequency can also be determined by other methods. The determination method of the difference between the first frequency and the second frequency is not limited in this embodiment.
[0127] In one instance, the general corpus is the Baidu Encyclopedia corpus, the general vocabulary is the Vocabulary dictionary of RoBERTa-tiny-clue, and the corresponding relationship between the new words and the target initials is as follows:
[0128] Corresponding relationship between newly added words and target words in Table 1
[0129] New character First frequency Target initial character Second frequency V 18716 ##bee 18713 I 18361 ##data 18320 R 16181 Xi 16182 G 9165 ##iki 9157 H 7010 Mi 7011 Yu 478 Qian 478 Yun 396 ##skip 419 Qian 232 ##onsored 214 Fu 222 ##onsored 214 Kai 206 ③ 209 Xie 22 ##α 19 Xing 21 ##α 19
[0130] In another example, initializing the weight parameters corresponding to the industry words in the industry vocabulary table based on the weight parameter matrix includes: obtaining the newly added words that are in the industry vocabulary table but not in the general vocabulary table; for each newly added word, determining the similarity of the pronunciation and / or meaning between the newly added word and each initial word; determining the target initial word with the greatest similarity of pronunciation and / or meaning to the newly added word; initializing the weight parameters corresponding to the newly added word based on the weight parameters corresponding to the target initial word.
[0131] Since the weight parameters of words with high similarity in pronunciation and / or meaning are similar, therefore, initializing the weight parameters of the newly added word based on the weight parameters of the initial word with the greatest similarity of pronunciation and / or meaning to the newly added word can make the initial value of the weight parameters of the newly added word as close as possible to the actual value. Therefore, the training difficulty of the initial network model can be reduced and the training speed of the initial network model can be improved.
[0132] Optionally, initializing the weight parameters corresponding to the newly added word based on the weight parameters of the target initial word includes: determining the weight parameters of the target initial word as the weight parameters of the newly added word.
[0133] Optionally, the method for obtaining the industry vocabulary table includes: determining the third frequency of each word appearing in the preset industry corpus; adding the words with the third frequency greater than the preset frequency threshold to the industry vocabulary table.
[0134] Optionally, the preset frequency threshold is pre-stored in the electronic device.
[0135] Optionally, the corpus in the industry corpus includes industry feature corpus. For example, the industry corpus in the video conferencing industry can include the conversation information in video conferencing. The industry corpus database can be a pre-collected industry corpus, or it can also be an open-source industry corpus. This embodiment does not limit the type of the industry corpus.
[0136] Optionally, the industry corpora of different industries are the same or different.
[0137] In one example, the first sample data includes the data in the industry corpus.
[0138] Since the words with the third frequency greater than the preset frequency threshold in the industry corpus are added to the industry vocabulary table, therefore, the industry vocabulary table can include industry feature words, so that the initial network model can learn the features of the industry feature words, and the accuracy of intent recognition of the trained intent recognition model can be improved.
[0139] In one example, the industry vocabulary table is obtained by modifying the general vocabulary table.
[0140] Optionally, the methods for modifying the general vocabulary to obtain the industry vocabulary include, but are not limited to, the following:
[0141] First, new words are added to the general vocabulary. For example, common words in the industry field that are not in the general vocabulary are added to the general vocabulary to obtain the industry vocabulary.
[0142] Correspondingly, an initial network model is established based on the weight parameters corresponding to the industry words, including: obtaining the new words that are in the industry vocabulary but not in the general vocabulary; adding the weight parameters corresponding to the new words to the weight parameter matrix of the pre-trained embedding layer to obtain the weight parameter matrix of the embedding layer of the initial network model, so that the weight parameter matrix of the embedding layer of the initial network model corresponds to the industry words in the industry vocabulary.
[0143] Optionally, adding common words in the industry field that are not in the general vocabulary to the industry vocabulary includes: determining the third frequency of each word appearing in the preset industry corpus; adding the words whose third frequency is greater than the preset frequency threshold and are not in the general vocabulary to the general vocabulary to obtain the industry vocabulary.
[0144] Second, some initial words in the general vocabulary are deleted. For example, unnecessary words in the industry field in the general vocabulary are deleted to obtain the industry vocabulary.
[0145] Correspondingly, an initial network model is established based on the weight parameters corresponding to the industry words, including: obtaining the deleted words that are in the general vocabulary but not in the industry vocabulary; deleting the weight parameters corresponding to the deleted words from the weight parameter matrix of the pre-trained embedding layer, or setting the weight parameters corresponding to the deleted words to unknown (UnKnown) to obtain the weight parameter matrix of the embedding layer of the initial network model, so that the weight parameter matrix of the embedding layer of the initial network model corresponds to the industry words in the industry vocabulary.
[0146] In an example, the intent recognition model only needs to recognize Chinese and English. At this time, deleting unnecessary words in the industry field in the general vocabulary from the industry vocabulary includes: deleting other languages and other language symbols (Other Tokens) except Chinese characters and English words in the general vocabulary to obtain the industry vocabulary.
[0147] In one example, the vocabulary distributions of different vocabularies are shown in Table 2. Among them, the number of vocabularies in the first general vocabulary is 21,128, including useless vocabularies in Chinese such as Korean and Japanese. The second general vocabulary table improves the first general vocabulary table for Chinese texts to make the second general vocabulary table more suitable for the needs of Chinese general texts and reduces the number of vocabularies. The industry vocabulary table is obtained by combining the words with frequencies greater than the preset frequency threshold in the industry Q&A texts stored over the years on the basis of the second general vocabulary table. And, considering that Chinese and English texts are used in the process of intent recognition and punctuation marks are removed, other language symbols in the second general vocabulary table are removed, and words for video conferencing scenarios are added, thus obtaining the industry vocabulary table. The number of vocabularies in the industry vocabulary table is 7,345, which is about one-third of the number of vocabularies in the first general vocabulary table. Therefore, the training speed of the initial network model can be improved, and at the same time, the intent recognition speed using the trained intent recognition model can also be enhanced.
[0148] Table 2 Vocabulary distributions of different vocabularies
[0149]
[0150] Step 104: Input the category prediction information, intent category, and sample frequency into the classification loss function to obtain the classification loss value.
[0151] In one example, the classification loss function is the softmax loss function. At this time, input the category prediction information, intent category, and sample frequency into the classification loss function to obtain the classification loss value, which is represented by the following formula:
[0152]
[0153] where L is the classification loss value; y is the intent category; p(y) is the sample frequency corresponding to the intent category y; f y (x; θ) is the probability that the first text data indicated by the category prediction information is of category y; f i (x; θ) is the probability that the first text data indicated by the category prediction information is of the i-th intent category; p(i) is the sample frequency corresponding to the i-th intent category; K is the number of intent categories; x is the first text data; θ is the model parameter of the initial network model.
[0154] As can be seen from the above classification loss function, taking the logarithm of the sample frequency and adding it to the classification loss function is equivalent to adding the sample distribution of each intention category in the sample data as a bias to the original classification loss function. This enables the trained intention recognition model to "rely on prior knowledge to solve classifications that can be solved by prior knowledge, and use the intention recognition model to solve the parts that cannot be solved by prior knowledge". Therefore, the speed of intention recognition using the trained intention recognition model can be improved.
[0155] Since inputting the sample frequencies corresponding to different intention categories into the classification loss function can enable the classification loss function to fuse the distribution of data of different intention categories in the first sample data, and thus calculate the loss value with the prior knowledge of this distribution, the initial network model can learn the features of different intention categories equally well at the level of the loss function. Therefore, the accuracy of intention recognition of the trained intention recognition model can be improved.
[0156] Step 105: Update the model parameters of the initial network model based on the classification loss value to train an intention recognition model.
[0157] In one example, updating the model parameters of the initial network model based on the classification loss value includes: in response to the classification loss value being greater than or equal to a preset loss degree threshold, updating the model parameters of the initial network model using the stochastic gradient descent method based on the classification loss value; and then executing again the step of inputting the category prediction information, intention category, and sample frequency into the classification loss function to obtain the classification loss value, that is, step 104, until the total loss value is less than the preset loss degree threshold and then stop to obtain the intention recognition model.
[0158] Optionally, the preset loss degree threshold is pre-stored in the electronic device.
[0159] In another example, updating the model parameters of the initial network model based on the classification calculation includes: in response to the number of iterative training not reaching the preset number of iterations, updating the model parameters of the initial network model using the stochastic gradient descent method based on the classification loss value; and then executing again the step of inputting the category prediction information, intention category, and sample frequency into the classification loss function to obtain the classification loss value, that is, step 104, until the number of iterations reaches the preset number of iterations and then stop to obtain the intention recognition model.
[0160] Optionally, the preset number of iterations is pre-stored in the electronic device.
[0161] In summary, for the training method of the intention recognition model provided in this embodiment, by obtaining first sample data, where the first sample data includes first text data and the intention category corresponding to the first text data; obtaining the sample frequency corresponding to each intention category; inputting the first text data into a pre-created initial network model to obtain category prediction information; inputting the category prediction information, the intention category, and the sample frequency into a classification loss function to obtain a classification loss value; and updating the model parameters of the initial network model based on the classification loss value to train the intention recognition model, it is possible to solve the problem that the initial network model has different learning degrees for the features of different intention categories, resulting in a lower accuracy of the trained intention recognition model in recognizing certain intention categories; since the sample frequencies corresponding to different intention categories are input into the classification loss function, the classification loss function can fuse the distribution of data of different intention categories in the first sample data, thereby calculating a loss value with the prior knowledge of this distribution, and making the initial network model have the same learning degree for the features of different intention categories at the level of the loss function. Therefore, the accuracy of the trained intention recognition model in intention recognition can be improved.
[0162] In addition, since the sample frequency is calculated based on the first sample data, the distribution of the first text data of each intention category determined based on the sample frequency is the same as the distribution of the first text data of each category in the first sample data, thus improving the accuracy of the determined classification loss value and the accuracy of the trained intention recognition model in intention category recognition.
[0163] In addition, since the weight parameters corresponding to the industry words in the industry vocabulary are initialized based on the weight parameter matrix, and the initial network model is established based on the weight parameters corresponding to the industry words, the weight parameters in the pre-trained embedding layer can be fully utilized, shortening the training time of the initial network model and improving the accuracy of the trained intention recognition model in intention category recognition.
[0164] In addition, since the weight parameters of the words with the smallest word frequency difference are also similar, initializing the weight parameters of the newly added words based on the weight parameters of the initial words with the smallest word frequency difference can make the initial values of the weight parameters of the newly added words as close as possible to the actual values. Therefore, the training difficulty of the initial network model can be reduced and the training speed of the initial network model can be improved.
[0165] In addition, since the words with the third frequency greater than the preset frequency threshold in the industry corpus are added to the industry vocabulary, the industry vocabulary can include industry feature words, so that the initial network model can learn the features of the industry feature words, and the accuracy of the trained intention recognition model in intention recognition can be improved.
[0166] In addition, since the first text data is extracted from the first sample data according to the sample weights and input into the pre-created initial network model, it is possible to balance the sample data of different intent categories at the data level, accelerate the model training speed, and improve the accuracy of the intent recognition model obtained by training in recognizing intent categories.
[0167] In addition, since the number of hidden layers in the initial network model is less than that in the BERT model, and the model structure of the intent recognition model is the same as that of the initial network model, it is possible to improve the speed of intent recognition using the trained intent recognition model.
[0168] Figure 3 It is a flowchart of an intent recognition method provided by an embodiment of the present application. The method at least includes the following steps:
[0169] Step 301, obtain target text data.
[0170] Optionally, obtaining target text data includes: obtaining target audio data; converting the target audio data into target text data.
[0171] In one example, the target audio data is collected during a video conference.
[0172] In actual implementation, other methods can also be adopted to obtain target text data. For example: scanning a book to obtain target text data, or target text data input by an input component. The present embodiment does not limit the method of obtaining target text data.
[0173] Step 302, input the target text data into the pre-trained intent recognition model to obtain the intent category corresponding to the target text data.
[0174] Among them, the intent recognition model is obtained by updating the model parameters of the pre-created initial network model based on the classification loss value; the classification loss value is obtained by inputting the category prediction information, the intent category corresponding to the first text data, and the sample frequency corresponding to the intent category into the classification loss function; the category prediction information is obtained by inputting the first text data into the initial network model.
[0175] In one example, inputting the target text data into the pre-trained intent recognition model includes: inputting the target text data into the pre-trained intent recognition model through a data generator.
[0176] Since the data generator can continuously maintain variables and return results during a single call, it is possible to modify the variables maintained by the data generator into target text data, and continuously obtain the intent categories corresponding to the target text data, thereby avoiding the problem of slow intent recognition speed caused by repeatedly loading and unloading the intent recognition model during the intent recognition of multiple target text data. Since only the variables of one generator need to be maintained to complete the intent recognition of multiple target document data, the number of times of loading and unloading the intent recognition model during the intent recognition process can be reduced, and the speed of intent recognition can be improved.
[0177] In addition, since using the data generator to input target data can achieve intent recognition while inputting, without generating a very large set of all target text data at once, it is possible to save the storage space of the Central Processing Unit (CPU).
[0178] Optionally, inputting the target text data into the pre-created initial network model through the data generator includes: calling the first function; using the second function to pass in the target text data, so that the data is input into the intent recognition model in the form of a generator.
[0179] In one example, the first function is from_generator of TensorFlow, and the second function is estimator.predict.
[0180] In actual implementation, other methods can also be used to input the target text data into the intent recognition model. For example, input the target text data into the intent recognition model based on the format of the file data, that is, use the same method as training, but only set batch_size = 1, that is, the number of samples fetched at one time is 1. This embodiment does not limit the method of inputting the target text data into the intent recognition model.
[0181] For related details, refer to the above method embodiments.
[0182] In summary, the intent recognition method provided in this embodiment obtains target text data; inputs the target text data into a pre-trained intent recognition model to obtain the intent category corresponding to the target text data. Among them, the intent recognition model is obtained by updating the model parameters of a pre-created initial network model based on a classification loss value. The classification loss value is obtained by inputting class prediction information, the intent category corresponding to the first text data, and the sample frequency corresponding to the intent category into a classification loss function. The class prediction information is obtained by inputting the first text data into the initial network model. This can solve the problem that the initial network model has different learning degrees for the features of different intent categories, resulting in a lower recognition accuracy of the trained intent recognition model for some intent categories. Since inputting the sample frequencies corresponding to different intent categories into the classification loss function can enable the classification loss function to fuse the distribution of data of different intent categories in the first sample data, and thus calculate the loss value with the prior knowledge of this distribution, it can make the initial network model have the same learning degree for the features of different intent categories at the level of the loss function. Therefore, the recognition accuracy of the trained intent recognition model for intent recognition can be improved.
[0183] In addition, since the data generator can continuously maintain variables and return results during one call, the number of times of loading and unloading the intent recognition model during the intent recognition process can be reduced, and the speed of intent recognition can be improved.
[0184] In addition, since using the data generator to input the target data can realize intent recognition while inputting, without generating a very large set of all target text data at one time, the storage space of the central processing unit can be saved.
[0185] Figure 4 It is a block diagram of a training device for an intent recognition model provided by an embodiment of the present application. The device at least includes the following modules: a sample acquisition module 410, a frequency acquisition module 420, a class prediction module 430, a loss calculation module 440, and a parameter update module 450.
[0186] The sample acquisition module 410 is used to acquire first sample data, where the first sample data includes first text data and the intent category corresponding to the first text data.
[0187] The frequency acquisition module 420 is used to acquire the sample frequencies corresponding to each intent category.
[0188] The class prediction module 430 is used to input the first text data into a pre-created initial network model to obtain class prediction information.
[0189] The loss calculation module 440 inputs the class prediction information, the intent category, and the sample frequency into a classification loss function to obtain a classification loss value.
[0190] The parameter update module 450 updates the model parameters of the initial network model based on the classification loss value to train an intent recognition model.
[0191] For related details, refer to the above method embodiments.
[0192] It should be noted that when training the intent recognition model by the training device for the intent recognition model provided in the above embodiments, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the training device for the intent recognition model is divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the intent recognition model provided in the above embodiments and the method embodiments for training the intent recognition model belong to the same concept. For the specific implementation process, refer to the method embodiments and will not be elaborated here.
[0193] Figure 5 It is a block diagram of an intent recognition device provided by an embodiment of the present application. The device at least includes the following modules: a text acquisition module 510 and an intent recognition module 520.
[0194] The text acquisition module 510 is used to acquire target text data;
[0195] The intent recognition module 520 is used to input the target text data into a pre-trained intent recognition model to obtain the intent category corresponding to the target text data;
[0196] Among them, the intent recognition model is obtained by updating the model parameters of a pre-created initial network model based on the classification loss value; the classification loss value is obtained by inputting the category prediction information, the intent category corresponding to the first text data, and the sample frequency corresponding to the intent category into a classification loss function; the category prediction information is obtained by inputting the first text data into the initial network model.
[0197] For related details, refer to the above method embodiments.
[0198] It should be noted that when performing intent recognition by the intent recognition device provided in the above embodiments, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the intent recognition device is divided into different functional modules to complete all or part of the functions described above. In addition, the intent recognition device provided in the above embodiments and the method embodiments for intent recognition belong to the same concept. For the specific implementation process, refer to the method embodiments and will not be elaborated here.
[0199] Figure 6Block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 601 and a memory 602.
[0200] The processor 601 may include one or more processing cores, such as: a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0201] The memory 602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 601 to implement the training method of the intent recognition model provided by the method embodiment in the present application, or the intent recognition method.
[0202] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 601, the memory 602, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0203] Of course, the electronic device may also include fewer or more components, and this embodiment does not limit this.
[0204] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the training method of the intention recognition model in the above method embodiments, or the intention recognition method.
[0205] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the training method of the intention recognition model in the above method embodiments, or the intention recognition method.
[0206] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0207] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A training method for an intention recognition model, characterized in that, The intention recognition model is used to recognize the intention of text data, and the method includes: Obtain first sample data, where the first sample data includes first text data and the intention category corresponding to the first text data; Obtain the sample frequency corresponding to each intention category; Obtain an embedding layer pre-trained using second text data, where the weight parameter matrix of the pre-trained embedding layer corresponds to the initial words in a general vocabulary; the general vocabulary includes the words in the second text data; the second text data is different from the first text data; Initialize the weight parameters corresponding to the industry words in the industry vocabulary based on the weight parameter matrix; the industry vocabulary includes the words in the first text data, and the industry vocabulary is partially the same as the general vocabulary; establish the initial network model based on the weight parameters corresponding to the industry words; the initial network model includes an embedding layer; the embedding layer is used to convert the words in text data into word vectors; Input the first text data into the pre-created initial network model to obtain category prediction information; Input the category prediction information, the intention category, and the sample frequency into a classification loss function to obtain a classification loss value; Update the model parameters of the initial network model based on the classification loss value to train and obtain the intention recognition model.
2. The method according to claim 1, wherein The obtaining the sample frequency corresponding to each intention category includes: For each target intention category, determine the number of category samples of the first text data whose intention category is the target intention category in the first sample data; Determine the ratio of the number of category samples to the total number of samples of the first text data in the first sample data as the sample frequency.
3. The method according to claim 1, wherein The inputting the category prediction information, the intention category, and the sample frequency into a classification loss function to obtain a classification loss value is represented by the following formula: where, L is the classification loss value; y is the intent category; p(y) is the sample frequency corresponding to the intent category y; f y (x; θ) is the probability that the first text data indicated by the category prediction information is of category y; f i (x; θ) is the probability that the first text data indicated by the category prediction information is of the i-th intent category; p(i) is the sample frequency corresponding to the i-th intent category; K is the number of intent categories; x is the first text data; θ is the model parameter of the initial network model.
4. The method according to claim 1, characterized in that The initializing the weight parameters corresponding to the industry words in the industry vocabulary based on the weight parameter matrix includes: Obtain the new words that are in the industry vocabulary and not in the general vocabulary; For each new word, determine the first frequency of the new word in a preset general corpus; For each initial word, determine the second frequency of the initial word in the general corpus; Determine the target initial word corresponding to the second frequency with the smallest difference from the first frequency; Initialize the weight parameters corresponding to the new words based on the weight parameters corresponding to the target initial word.
5. The method according to claim 1, wherein The method further includes: Determine the third frequency of each word appearing in a preset industry corpus; Add the words whose third frequency is greater than a preset frequency threshold to the industry vocabulary.
6. The method according to claim 1, characterized in that The inputting the first text data into the pre-created initial network model includes: Based on the first text data, the intention category, and the number of intention categories, obtain the sample weights corresponding to the first text data of each intention category; Extract the first text data from the first sample data according to the sample weights and input it into the pre-created initial network model.
7. The method according to claim 1, wherein The initial network model includes at least one hidden layer, and the number of hidden layers in the initial network model is less than the number of hidden layers in the BERT model.
8. A method for intention recognition, characterized in that, The method includes: Obtain target text data; Input the target text data into a pre-trained intent recognition model to obtain the intent category corresponding to the target text data; Among them, the intent recognition model is obtained by updating the model parameters of a pre-created initial network model based on a classification loss value; the classification loss value is obtained by inputting category prediction information, the intent category corresponding to the first text data, and the sample frequency corresponding to the intent category into a classification loss function; the category prediction information is obtained by inputting the first text data into the initial network model.
9. The method according to claim 8, wherein Inputting the target text data into a pre-trained intent recognition model includes: Input the target text data into a pre-trained intent recognition model through a data generator.
10. An electronic device, characterized in that, The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement the training method of the intent recognition model according to any one of claims 1 to 7, or to implement the intent recognition method according to claim 8 or 9.
11. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, and when the program is executed by a processor, it is used to implement the training method of the intent recognition model according to any one of claims 1 to 7, or to implement the intent recognition method according to claim 8 or 9.
Citation Information
Patent Citations
Sample data processing method, sample data processing device and electronic equipment
CN111198938A
Neural network training method and apparatus, electronic device and storage medium
WO2021174739A1