A slot recognition method and an electronic device
By jointly optimizing the intention recognition and slot filling in the human-computer dialogue system, and using the joint training model to recognize voice conversations, the problem of low recognition accuracy caused by independent intention recognition and slot filling in the prior art is solved, and the accuracy and user experience of voice recognition are improved.
Patent Information
- Application Number
- CN202010210034.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-03-23
AI Technical Summary
In the prior art, intention recognition and slot filling tasks are carried out independently, resulting in the trained model having a low recognition accuracy in speech recognition, which reduces the user experience.
By jointly optimizing the intention recognition and slot filling, the joint training model is used to recognize speech conversations to improve the accuracy of speech recognition. The specific method includes preprocessing the user command, generating an intent vector and a hidden state vector through BERT encoding processing, and combining the attention vector and the slot probability vector for slot recognition.
Through joint optimization training, the slot prediction results are corrected, and the accuracy of the dialogue system to understand the user request information is improved, and the user experience is improved.
Smart Images

Figure CN113505591B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a slot identification method and an electronic device. Background Art
[0002] With the rapid development of the Internet, the application of human-computer dialogue systems is becoming more and more extensive. Taking the task-oriented human-computer dialogue system as an example, the user's input can be questions such as asking about the weather, ordering air tickets, and medical consultations. The human-computer dialogue system can feedback the response dialogue information to the user based on the user's question information. For example, the question information input by the user can be "What is the weather in Beijing tomorrow?" The human-computer dialogue system can search the preset database and feedback the response dialogue information to the user as "The weather in Beijing tomorrow will be sunny and cloudy." It can be seen that for the human-computer dialogue system, it is crucial to accurately identify the question information input by the user. In the process of speech recognition, intent recognition and slot filling technology are the key to ensuring the accuracy of speech recognition results.
[0003] For intent recognition, it can be abstracted as a classification problem, and then the intent recognition model can be trained using convolution and knowledge representation classifiers. In addition to embedding the user's voice questions into words, the semantic representation of knowledge is also introduced to increase the generalization ability of the representation layer in the intent recognition model. However, in actual applications, it is found that the model has the defect of slot filling bias, which affects the accuracy of the intent recognition model. For slot filling, its essence is to formalize the sentence sequence into annotated sequence. There are many commonly used methods for annotating sequences, such as hidden Markov models or conditional random field models. However, these slot filling models, in specific application scenarios, lack of contextual information will lead to ambiguity in slots under different semantic intents, and thus cannot meet the actual application requirements. It can be seen that the training of the two models in the prior art is carried out independently, and there is no combined optimization for the intent recognition task and the slot filling task, which ultimately leads to the problem of low recognition accuracy of the trained model in speech recognition, which reduces the user experience. Summary of the invention
[0004] The present application provides a slot recognition method and an electronic device for jointly optimizing the training of intent recognition and slot filling, and using a joint training model to recognize voice dialogues to improve the accuracy of voice recognition.
[0005] In a first aspect, an embodiment of the present application provides a slot recognition method, which can be executed by a human-machine dialogue system or a human-machine dialogue device. The method includes: preprocessing a user command to obtain an original word sequence, and performing BERT encoding processing on the original word sequence to obtain an intent vector and a hidden state vector of each word segment. For any word segment, that is, the first word segment, the following processing is performed: determining an attention vector of the first word segment according to the hidden state vector and the intent vector of the first word segment. Then, splicing the hidden state vector of the first word segment and the attention vector of the first word segment to determine a slot probability vector of the first word segment; determining the slot corresponding to the first word segment according to K probability values in the slot probability vector of the first word segment.
[0006] It can be seen that in the embodiment of the present application, the intent is used as the input of the slot filling task to correct the slot prediction result, which helps to improve the accuracy of the dialogue system in understanding the user's request information and enhances the user experience.
[0007] In a possible design, for any one of the T-1 word segments in the original word sequence except the first word segment, the attention vector of the word segment can be further determined in combination with the slot probability vector corresponding to the previous word segment, that is, the attention vector of the word segment can be determined according to the hidden state vector of the word segment, the intent vector, and the slot probability vector corresponding to the previous word segment of the word segment.
[0008] In the embodiment of the present application, using the intent and the slot of the previous word segment as the input of the slot filling task of the next word segment helps to further correct the slot prediction result, thereby helping to further improve the accuracy of the dialogue system in understanding the user's request information.
[0009] In a possible design, the method of preprocessing the user command to obtain the original word sequence includes: first generating a Token sequence according to the user command, then randomly sorting the Token sequence, and dividing the Token sequence into multiple batches of Token sequences according to batch_size; finally, performing truncation or padding operations on each batch of Token sequences to obtain the preprocessed original word sequence. Preprocessing the user command in the above manner in this method helps to filter out invalid information.
[0010] In a possible design, the specific way to generate the hidden state vector can be: first performing BERT semantic encoding on the original word sequence to generate a vector sequence h0, h1, ……, h T , where h0 is the sentence vector encoding information of the user command, h1, ……, h Tare the hidden state vectors corresponding to T word segments respectively; then, according to the sentence vector encoding information h0 of the user command, an intention vector of the user command is generated, where the intention vector satisfies where, y I ∈R 1 ×I , I represents the number of possible intentions of the user command, and the intention corresponding to the maximum probability value in y I is the intention of the user command, h0 is the sentence vector encoding information of the user command, is the bias term, is the weight matrix.
[0011] In a possible embodiment, the specific calculation method of the slot probability vector may include: first, concatenate the hidden state vector h i of the first word segment and the attention vector of the first word segment to generate deep vector encoding information, and the deep vector encoding information satisfies where, concat is the concatenation operation function, represents the deep vector encoding information after concatenation. Then, perform a softmax transformation of the deep vector encoding information by a logistic regression model, so as to obtain the slot probability vector of the first word segment, and the slot probability vector of the first word segment satisfies where, softmax represents the normalized exponential function, represents the weight matrix, represents the deep vector encoding information, and the represents the bias term.
[0012] In a second aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory is used to store one or more computer programs; when the one or more computer programs stored in the memory are executed by the processor, the electronic device can implement the method of any possible design in any of the above aspects.
[0013] In a third aspect, an embodiment of the present application further provides a device, and the device includes modules / units that execute the method of any possible design in any of the above aspects. These modules / units can be implemented by hardware or by hardware executing corresponding software.
[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium includes a computer program. When the computer program runs on an electronic device, the electronic device executes the method of any possible design in any of the above aspects.
[0015] In a fifth aspect, an embodiment of the present application further provides a computer program product. When the computer program product runs on a terminal, the electronic device is caused to execute the method of any possible design in any of the above aspects.
[0016] In a sixth aspect, an embodiment of the present application further provides a chip. The chip is coupled to a memory and is configured to execute a computer program stored in the memory to execute the method of any possible design in any of the above aspects. Description of the Drawings
[0017] Figure 1 It is a schematic diagram of a possible dialogue system architecture applicable to an embodiment of the present application;
[0018] Figure 2 It is a schematic diagram of a joint training model provided by an embodiment of the present application;
[0019] Figure 3 It is a schematic flowchart of a slot recognition method provided by an embodiment of the present application;
[0020] Figure 4 It is a schematic diagram of system module interaction provided by an embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of a dialogue interface of a dialogue system provided by an embodiment of the present application;
[0022] Figure 6 It is a schematic diagram of an example of joint intention speculation and slot filling provided by an embodiment of the present application;
[0023] Figure 7 It is an exemplary block diagram of a device provided by an embodiment of the present application;
[0024] Figure 8 It is a schematic diagram of a device provided by an embodiment of the present application. Detailed Embodiments
[0025] First, some terms involved in the present application are explained to facilitate understanding by those skilled in the art.
[0026] (1), User command
[0027] In the field of human-computer dialogue, a user command is what the user inputs, and it can also be referred to as a user requirement. In the embodiments of this application, a user command can be one or a combination of several types such as voice, image, video, audio-video, text, etc. For example, if the user command is the voice input by the user through a microphone, at this time, the user command can also be called a "voice command"; another example is that the user command is the text input by the user through a keyboard or a virtual keyboard, and at this time, the user command can also be called a "text command"; another example is that the user command is the image input by the user through a camera, and the user inputs "Who is the person in the image?" through a virtual keyboard. At this time, the user command is a combination of an image and text; another example is that the user command is an audio-video input by the user through a camera and a microphone, and at this time, the user command can also be called an "audio-video command".
[0028] (2) Speech recognition
[0029] Speech recognition technology, also known as automatic speech recognition (ASR), computer speech recognition, or speech to text (STT), is a method of converting human speech into corresponding text by a computer.
[0030] When the user command is a voice command or a command containing voice, the user command can be converted into text through ASR. Generally, the working principle of ASR is as follows: The first step is to split the audio signal input by the user into frames to obtain frame information; the second step is to recognize the obtained frame information into states, where several frame information correspond to one state; the third step is to combine the states into phonemes, where every three states are combined into one phoneme; the fourth step is to combine the phonemes into words, and several phonemes form one word. It can be seen that as long as it is known which state each frame of information corresponds to, the result of speech recognition will come out. And how to determine which state each frame of information corresponds to. Usually, it can be recognized which state the frame information corresponds to with the highest probability, and then this frame of information belongs to that state.
[0031] In the process of speech recognition, an acoustic model (AM) and a language model (LM) can be used to determine a set of character sequences corresponding to a piece of speech. Among them, the acoustic model can be understood as modeling the pronunciation. It can convert the speech input into an output of an acoustic representation, that is, decode the acoustic features of a piece of speech into units such as phonemes or words. More precisely, it gives the probability that the speech belongs to a certain acoustic symbol (such as a phoneme). The language model, on the other hand, gives the probability that a set of character sequences is this piece of speech, that is, decodes the words into a set of character sequences (that is, a complete sentence).
[0032] (3) Natural Language Understanding (NLU)
[0033] Natural language understanding aims to enable machines to have the language understanding ability of normal people like humans. Among them, an important function is intent recognition. For example, if the user's command is "How far is Hilton Hotel from Baiyun Airport?", the intent of the user's command is "query distance". The slots configured for this intent are "starting point" and "destination". The information of the slot "starting point" is "Hilton Hotel", and the information of the slot "destination" is "Baiyun Airport". With the information of the intent and slots, the machine can give an answer.
[0034] (4) Intent and Intent Recognition
[0035] Intent refers to identifying what the user's command specifically wants to do. Intent recognition can be understood as a problem of semantic expression classification. That is to say, intent recognition is a classifier (also called an intent classifier in the embodiments of this application), which determines which intent the user's command is. Commonly used intent classifiers for intent recognition are Support Vector Machine (SVM), decision tree, and deep neural network (DNN). Among them, the deep neural network can be a convolutional neural network (CNN) or a recurrent neural network (RNN), etc. The RNN can include a long short-term memory (LSTM) network, a stacked recurrent neural network (SRNN), etc.
[0036] The general process of intent recognition includes: First, preprocess the corpus (i.e., a sequence of words), such as removing punctuation marks and stop words from the corpus, etc.; Second, use the word embedding algorithm, such as the word2vec algorithm, to generate word vectors (word embedding) for the preprocessed corpus; Furthermore, use an intent classifier (such as LSTM) to perform feature extraction, intent classification, etc. In the embodiments of this application, the intent classifier is a trained model that can recognize intents in one or more scenarios, or recognize any intent. For example, the intent classifier can recognize intents in the flight reservation scenario, including booking a flight, filtering flights, querying flight prices, querying flight information, canceling a flight, changing a flight, querying the distance to the airport, etc. Another example is that the intent classifier can recognize intents in multiple scenarios.
[0037] (5) Slot
[0038] After the user's intention is determined, the NLU module needs to further understand the content in the user's command. For simplicity, the most core part can be selected for understanding, and the rest can be ignored. Those most important parts can be called slots. That is to say, a slot is the definition of key information in the user's expression (such as a sequence of words recognized from the user's command). One or more slots can be configured for the intention of the user's command to obtain the information of that slot, and the machine can respond to the user's command. For example, in the intention of booking a flight ticket, the slots are "departure time", "departure place", and "destination". These three key pieces of information need to be recognized during natural language understanding. To accurately recognize slots, slot types are required. Still taking the above example, if we want to accurately recognize the three slots of "departure time", "departure place", and "destination", we need the corresponding slot types behind, which are "time" and "city name" respectively. It can be said that slot types are structured knowledge bases of specific knowledge, used to recognize and transform the slots in the user's colloquial expressions. From the perspective of programming languages, intent + slot can be regarded as using a function to describe the user's needs, where "intent corresponds to the function", "slot corresponds to the parameters of the function", and "slot_type corresponds to the type of the parameters". The slots configured for different intents can be divided into required slots and optional slots. Among them, the required slots are the slots that must be filled to execute the user's command, and the optional slots are the slots that can be optionally filled or not filled to execute the user's command. Without further explanation, in this application, a slot can be a required slot or an optional slot, or it can be a required slot.
[0039] In the above example of "booking a flight ticket", three core slots are defined, namely "departure time", "departure place", and "destination". If we want to comprehensively consider the content that needs to be input by the user when booking a flight ticket, we can think of more slots, such as the number of passengers, the airline company, the departure airport, the arrival airport, etc. For slot designers, slots can be designed based on the granularity of the intention.
[0040] (6) Slot filling
[0041] Slot filling is to extract the structured fields in the user command, which can also be said to read some semantic components in the sentence (referred to as the user command in the embodiments of this application). Therefore, slot filling can be regarded as a sequence labeling problem. Sequence labeling problems include word segmentation, part-of-speech tagging, named entity recognition (NER), keyword extraction, semantic role labeling, etc. in natural language processing. When performing sequence labeling, given a specific set of tags, sequence labeling can be carried out. The methods for solving sequence labeling problems include the maximum entropy Markov model (MEMM), conditional random field (CRF), and recurrent neural network (RNN), etc.
[0042] Sequence labeling is to assign a label to each character in the given text. Essentially, it is a problem of classifying each element in the linear sequence according to the context. That is, for a one-dimensional linear input sequence, assign a certain label in the set of labels to each element in the linear input sequence. In the embodiments of this application, the text annotation slots of the user command can be implemented through a slot extraction classifier. In the NLU involved in the embodiments of this application, the linear sequence is the text of the user command (the text input by the user or the text recognized from the input speech). Usually, a Chinese character can be regarded as an element of the linear sequence. For different tasks, the meanings represented by the set of labels are different. Sequence labeling is to assign a suitable label to the Chinese character according to the context of the Chinese character, that is, to determine its slot.
[0043] Exemplarily, when the filling information of the slot is missing in the user command, for example, the user command is "How far is this hotel from Hongqiao Airport?" When the machine responds to this user command, it needs to know which hotel "this hotel" refers to. In the prior art, the machine may ask the user "Which hotel do you want to query the distance from to Hongqiao Airport?" to obtain the information of this slot. It can be seen that the machine needs to interact with the user multiple times to obtain the information of the missing slot in the user command.
[0044] In order to make the purpose, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings.
[0045] Figure 1 It is a schematic diagram of a possible human-machine dialogue system architecture applicable to the embodiments of this application. As Figure 1 shown, the human-machine dialogue system architecture may include a server 101, one or more user-side devices (such as Figure 1 the user-side devices 1021 and 1022 shown). Optionally, the dialogue system architecture may further include one or more customer service-side devices (such as Figure 1 the customer service-side devices 1023 and 1024 shown).
[0046] Among them, the user-side device or the customer service-side device can be a terminal device, such as a mobile phone, a tablet computer, or a desktop computer, etc., and there is no specific limitation. In the embodiments of the present application, based on the differences in the operation interfaces of the client devices, the client devices are divided into user-side devices (client devices for users to operate) and customer service-side devices (client devices for human customer service to operate). In other embodiments, the user-side device can also be called a user client or other names, and the customer service-side device can also be called an artificial seat client or other names, and there is no specific limitation.
[0047] The user-side device can be used to obtain the information input by the user and send the information input by the user to the server. For example, if the user enters text information in the dialog box, the user-side device can obtain the text information and send it to the server; if the user enters voice information in the dialog box, the user-side device can convert the voice into text information through speech recognition technology and then send the text information to the server. Optionally, the user-side device can also communicate with the customer service-side device. For example, the user-side device sends the input information of the user to the customer service-side device and receives the information returned by the customer service-side device, so as to realize that the human customer service provides services for the user.
[0048] The server is used to process various calculations required by the human-machine dialogue system, such as question-answer matching, that is, searching a preset database according to the user's request information to obtain the corresponding answer information of the request information. Among them, the preset database can include a question bank and a corresponding answer bank for the question bank. The question bank includes multiple preset request information, and the answer bank includes the corresponding answer information of multiple preset request information. The server can compare the user's request information with multiple preset request information, and then feedback the answer information corresponding to the preset request information with the highest similarity to the user's request information in the answer bank to the user-side device, and then present it to the user by the user-side device. There may be various presentation methods, and there is no specific limitation.
[0049] Taking a task-based multi-turn dialogue system as an example, its application scenarios can be virtual personal assistants, intelligent customer service, etc. Currently, usually, when a virtual personal assistant or intelligent customer service conducts a task-based multi-turn dialogue with a user, it can return response information that meets the user's expectations based on the request information input by the user. However, due to the strong flexibility of language expression, in the current process of virtual personal assistants or intelligent customer service understanding the semantics of the sentences input by users, the intent recognition task and the slot filling task are carried out independently. Since the slots corresponding to different intents may be different, there may be a problem that the intent and the slot in the recognition result are not aligned, which may cause the virtual personal assistant or intelligent customer service to misunderstand the user's semantics and return response information that does not meet the user's expectations to the user. For example, when the voice content input by the user is "Buy a plane ticket to Shanghai tomorrow morning", the intent recognition result of the intent recognition task in the human-machine dialogue system is determined to be "book a ticket", while the slot recognition of the slot filling task may be to mark "Shanghai" as the "navigation destination". It can be seen that since the intent recognition task and the slot filling task are carried out independently, it is very likely that the intelligent customer service thinks that the slot of the "navigation departure place" needs to be filled, so it inputs incorrect response information such as "Is the navigation departure place the current location?" to the user. That is to say, the traditional slot filling process fails to consider the intent recognition result, resulting in the misalignment of intent and slot, which affects the semantic parsing result.
[0050] Generally, it is found in actual applications that slot filling is related to intent. For example, the same location entity (such as Chenghuang Temple) represents a restaurant under the food search intent, while under the navigation intent, it may represent the starting point or the ending point. Based on this analysis, the embodiment of the present application provides a slot recognition method, which improves the training model in the existing human-machine dialogue system. In the improved joint training model, the intent recognition result is used as an input parameter for slot filling. That is to say, the intent recognition result is associated with slot filling to achieve the correction of slot filling by using the intent recognition result. Using the joint training model trained by this method to perform semantic parsing on the input information of the user can improve the accuracy of the semantic parsing result.
[0051] Among them, the improved joint training model is as Figure 2As shown in the figure, the joint training model includes a Bert encoding layer, a dense layer, a Masked Attention Layer, and a softmax (logistic regression model) layer. When a dialogue corpus is input into the joint training model, the logistic regression model of the joint training model can output the information of the intent and slot corresponding to the dialogue corpus. Among them, from the connection relationship between the dense layer and the Masked Attention Layer, it can be seen that the intent is an input of the Masked Attention Layer, and the dense layer and the Masked Attention Layer are used to associate the intent recognition result with the slot filling.
[0052] It should be noted that during the training process of the joint training model, the target loss function for the joint optimization of the intent and slot is the sum of the intent classification loss function, the slot filling loss function, and the regularization term of the weight. Among them, the intent classification loss function can adopt the binary Cross Entropy Loss function, and the slot filling loss function can adopt the multi-class Cross Entropy Loss function. The training termination condition for triggering the joint training model can be: when the number of training epochs reaches the set threshold or the batch interval from the previous optimal model is greater than the set threshold, the training terminates, and the final joint training model is generated.
[0053] It should be noted that the slot recognition method provided in the embodiments of this application can be applied to a variety of possible human-machine dialogue systems, especially task-based multi-turn dialogue systems. To facilitate the understanding of the solution of the embodiments of this application, the following takes a task-based multi-turn dialogue system as an example to illustrate the slot recognition method.
[0054] Figure 3 It is a schematic flowchart of a slot recognition method provided in the embodiments of this application, including:
[0055] Step 301, after the server obtains the user command, it preprocesses the user command to obtain the original word sequence.
[0056] Among them, the user command can be one or a combination of multiple types such as voice, image, video, audio-video, text, etc. For example, the user command is the voice input by the user through the microphone. At this time, the user command can also be called a "voice command"; another example is that the user command is the text input by the user through the keyboard or virtual keyboard. At this time, the user command can also be called a "text command". Specifically, the server can first mark the user command to generate a Token (mark) sequence, then randomly sort the Token sequence, and divide the Token sequence into multiple batches of Token sequences according to the batch_size (batch size). Finally, perform truncation or padding operations on each batch of Token sequences to obtain the preprocessed original word sequence. Optionally, the server can also create a mask with the same dimension for the Token sequence after the truncation or padding operation, where <pad>Element position, where the element value in the mask sequence is 0, otherwise 1.
[0057] Exemplarily, a mobile phone user can run the voice assistant software program in the mobile phone. When the user inputs the voice "play red breast" through the microphone, the voice assistant software program converts the voice into English text and sends it to the server corresponding to the program. First step, the server performs serialization processing on the English text. For example, the WordPiece (word fragment) technology can be used to generate a Token sequence from the text. It should be noted that if the text is Chinese text, a character-based method can be used to generate a Token sequence from the text. Second step, the server randomly sorts the serialized corpus and divides the corpus into multiple batches according to the batch_size. Then, a truncation or padding operation is performed on the Token sequence of each Batch. Specifically, for each Token sequence in each batch, if its length + 2 is less than the predetermined maximum sequence length (usually maxLength = 512), then padding <pad>, if its length + 2 is greater than the predetermined maximum sequence length, truncate the extra tokens; after truncation / padding, pad at the beginning of the token sequence with <cls>, used for marking as a classification task, padding at the end of the Token sequence <sep>, used for sentence segmentation, indicating that the previous part is a complete sentence. For example, as shown in Table 1.
[0058] Table 1
[0059]
[0060] Step 302, the joint training model in the server first performs BERT encoding on the original word sequence to obtain the intent vector and the hidden state vectors corresponding to T word segments respectively.
[0061] Among them, BERT (bidirectional encoder representations from Transformers), BERT is a deep bidirectional pre-trained language understanding model used as a feature extractor. Specifically, after the server performs BERT semantic encoding on the original word sequence input into the joint training model, a sequence of hidden state vectors h0, h1, ……, h is generated T , where h0 is the sentence vector encoding information corresponding to the user command (that is, corresponding to <cls>(encoding vectors of positions), h1, ……, h T are the hidden state vectors corresponding to the T word segments respectively (that is, the encoding vectors corresponding to the remaining positions). Further, the BERT encoding layer inputs the sentence vector h0 into the logistic regression model (Softmax) layer to generate an intent vector, that is y I ∈R 1×I , I represents the number of intents corresponding to the user command, and the intent corresponding to the maximum probability value in y I is the intent corresponding to the user command, h0 is the sentence vector encoding information of the user command, is the bias term, is the weight matrix.
[0062] Exemplarily, when the user inputs the voice command "play red breast" as shown in Table 1, after the original word sequence corresponding to the voice command is input into the jointly trained model, the output intent is "play music", as Figure 2 shown. In addition, the hidden state vector h1 corresponding to "play", the hidden state vector h2 corresponding to "red", and the hidden state vector h3 corresponding to "breast" are input into the fully connected layer.
[0063] Step 303, the jointly trained model in the server performs the following processing on the first word segment in the original word sequence, where the first word segment is any one of the T word segments:
[0064] Determine the attention vector of the first word segment according to the hidden state vector and the intent vector of the first word segment; then splice the hidden state vector of the first word segment and the attention vector of the first word segment to determine the slot probability vector of the first word segment; determine the slot corresponding to the first word segment according to K probability values in the slot probability vector of the first word segment.
[0065] Among them, the way to splice the hidden state vector of the first word segment and the attention vector of the first word segment can be:
[0066] Splice the hidden state vector h i of the first word segment and the attention vector of the first word segment to generate deep vector encoding information, and the deep vector encoding information satisfies where concat is the splicing operation function, represents the deep vector encoding information after splicing. Further, the server inputs the deep vector encoding information into the softmax layer for softmax conversion to obtain the slot probability vector of the first word segment, and the slot probability vector of the first word segment Meet Among them, softmax represents the weight matrix, Represents the weight matrix, Represents the deep vector encoding information, the Represents the bias term.
[0067] Combine Figure 2 In terms of, for the first token "play" in the voice command "play red breast", c1 in the fully connected layer is the intent vector y I , the intent vector y I And the hidden state vector h1 are used as the inputs of the masked attention layer, so as to generate the attention vector corresponding to the first token "play". For the second token "red" in the voice command "play red breast", c2 in the fully connected layer is the intent vector y I , the intent vector y I And the hidden state vector h2 are used as the inputs of the masked attention layer, so as to generate the attention vector corresponding to the second token "red". For the second token "breast" in the voice command "play red breast", c3 in the fully connected layer is the intent vector y I , the intent vector y I And the hidden state vector h2 are used as the inputs of the masked attention layer, so as to generate the attention vector corresponding to the third token "breast".
[0068] In another possible implementation, assuming that the first token is any one of the T - 1 tokens except the first token, the attention vector of the first token can be determined in the following way: Determine the attention vector of the first token according to the hidden state vector, intent vector of the first token, and the slot probability vector corresponding to the previous token of the first token, and finally obtain the slot probability vector corresponding to "play" For example, in terms of Figure 2 , for the first token "play" in the voice command "play red breast", c1 in the fully connected layer is the intent vector y I , the intent vector y I And the hidden state vector h1 are used as the inputs of the masked attention layer, so as to generate the attention vector corresponding to the first token "play". For the second token "red" in the voice command "play red breast", c2 in the fully connected layer is the intent vector y I , the intent vector y I And the hidden state vector h2, and the slot probability vector corresponding to "play" As the input of the masked attention layer, an attention vector corresponding to the second token "red" is generated For the second token "breast" in the voice command "play red breast", c3 in the fully connected layer is the intent vector y I , the intent vector y I and the hidden state vector h2, as well as the attention vector corresponding to the second token "red" As the input of the masked attention layer, an attention vector corresponding to the third token "breast" is generated
[0069] Specifically, in the process of calculating the attention vector, the masked attention layer can calculate the attention vector according to the following calculation method:
[0070] Step a, for the hidden vector h at each moment i (i = 1,…,T), the query vector It only receives the intent vector information y I , the hidden vector information h at the current moment i and the slot output information predicted at the previous moment (time t ranges from 1 to i - 1) where q i ∈R 1×d ;
[0071] Step b, the Key vector information is linearly transformed into where k i ∈R 1×d , and the key vector information at all moments forms a matrix K=(k1,k2…,k T )∈R n×d ;
[0072] Step c, the Value vector information is linearly transformed into where v i ∈R 1×d , and the value vector information at all moments forms a matrix V=(v1,v2…,v T )∈R n×d ;
[0073] Step d, calculate the attention vector information at the current moment, where The M matrix is a mask matrix, and its form is an upper triangular identity matrix, that is, when i ≤ j, m ij = 1, while when i > j, m ij = -∞.
[0074] It can be seen that in the embodiment of the present application, when the server receives a user command, it first performs preprocessing according to step 301 to generate an original word sequence, and then uses the original word sequence as the input of the joint training model. After model prediction and inference, an intent vector y is obtained. I and a slot vector sequence y i S . For the intent vector y I , the intent corresponding to the maximum probability can be selected as the predicted intent; for the i-th slot vector y i S , the slot corresponding to the maximum probability can be selected as the i-th predicted slot. In this method, the intent is used as the input of the slot filling task to correct the slot prediction result, thereby improving the accuracy of the dialogue system in understanding the user's request information and enhancing the user experience.
[0075] The slot recognition method provided by the present application can be specifically applied to the Figure 4 shown system architecture, where the NLU module integrates the joint training model. The following combines this system architecture to illustrate the specific application process of the above method through examples. The specific steps are as follows.
[0076] Step 401, the user opens the voice assistant software program on the mobile phone and sends the voice message "Book a flight ticket from Shenzhen to Shanghai for me" in the dialog box.
[0077] Step 402, the ASR module in the mobile phone voice assistant software program converts the voice message into text information, as Figure 5 shown, and sends the converted text information to the DM (Dialog manage) module in the voice assistant software program.
[0078] Step 403, the DM module in the voice assistant software program obtains the context information corresponding to the voice message from the dialog box (such as Figure 5 the historical dialogue information in, and the current dialogue information), as well as the status information, etc. The DM module sends the voice message and other relevant information to the NLU (Natural Language Understanding) module.
[0079] Step 404, the NLU module identifies the intent and slots in the text "Book a flight ticket from Shenzhen to Shanghai for me" according to the method provided by the embodiment of the present application.
[0080] Specifically, as Figure 6 shown, after "Book a flight ticket from Shenzhen to Shanghai for me" is encoded semantically by BERT, a hidden state vector sequence h0, h1, ……, h8 is generated. The BERT encoding layer inputs the sentence vector h0 into the Softmax layer to generate an intent vector y I , y I The intention corresponding to the maximum probability value is to book a flight ticket. When performing slot filling, the hidden state vector h1 and the intention vector y I are used as inputs to generate the slot probability vector corresponding to the first token "help". Among them, the slot corresponding to the maximum probability value is empty, so it is marked as "o". In addition, when performing slot filling, the hidden state vector h4 and the intention vector y I , and the slot probability vector of the previous token "book" are used as inputs to generate the slot probability vector corresponding to the fourth token "Shenzhen". Among them, the slot corresponding to the maximum probability value is "departure location", so it is marked as "departure location" or "FromLoc". And so on, when performing slot filling, the hidden state vector h6 and the intention vector y I , and the slot probability vector of the previous token "book" are used as inputs to generate the slot probability vector corresponding to the sixth token "Shenzhen". Among them, the slot corresponding to the maximum probability value is "destination", so it is marked as "destination" or "ToLoc". Therefore, the joint training model outputs that the intention corresponding to "Help me book a flight from Shenzhen to Shanghai" is "book a flight ticket", the slot corresponding to "Shenzhen" is "departure location", and the slot corresponding to "Shanghai" is "destination".
[0081] Step 405, the NLU module returns the intention and slot recognition results to the DM module.
[0082] Step 406, the DM module inputs the intention and slot recognition results into the NLG (Natural Language Generation) module.
[0083] Among them, the DM module is divided into two sub-modules, namely Dialogue State Tracking (DST) and Dialogue Policy Learning (DPL). Its main function is to update the state of the dialogue system according to the recognition results of the NLU module and generate corresponding system actions, such as querying flight tickets.
[0084] Step 407, the NLG module texturizes the system action output by the DM, expresses the system action in text form, and sends other relevant information to the DM module.
[0085] Step 408, the DM module sends the system action execution result to the TTS (Text-to-Speech) module.
[0086] Step 409, the TTS module converts the text into speech and outputs the speech to the user. For example, output the speech content corresponding to the queried flight information.
[0087] Further, if the server determines that there are still unfilled slots, for example, the "time" slot in the request information "book a flight from Shenzhen to Shanghai for me" is unfilled. Therefore, the server can further send guiding information to the user-side device, such as "Which day do you want to book the flight?", as Figure 5 shown. This guiding information is used to guide the user to provide associated information of the request information. In this way, by using the guiding information to guide the user to provide the associated information of the request information, it is convenient for the server to query the response information that meets the user's expectations in the preset database based on the associated information of the request information, thereby avoiding unnecessary transfer to manual operation and improving user satisfaction.
[0088] In the embodiment of the present application, the guiding information may further include third historical conversation information, and the similarity between the historical request information in the third historical conversation information and the request information is greater than a sixth threshold. That is to say, the guiding information may further append historical request information similar to the user's request information, so as to remind the user that they have asked similar questions.
[0089] It can be understood that in other possible situations, the guiding information may further include other possible contents. Those skilled in the art can set the contents included in the guiding information according to actual experience and needs. However, any information that plays a guiding, prompting, and soothing role in actively seeking response information that meets the user's expectations and is sent to the user is within the protection scope of the present invention.
[0090] It can be seen that in the embodiment of the present application, intent recognition and slot filling are jointly performed. The intent is considered during the slot filling process, so the granularity of slot filling will be finer and more accurate. In this way, only one joint model result is needed to well complete the two tasks.
[0091] It should be noted that: (1) The above step numbers are only an example of the execution process of the embodiment of the present application. There is no strict execution order between the steps that do not have a timing dependence relationship with each other among the above steps. Each step from step 401 to step 410 is not a necessary execution step. In specific implementation, some of the steps can be selectively executed according to actual needs.
[0092] The above mainly introduces the solution provided by this application from the perspective of the interaction between various devices. It can be understood that, in order to implement the above functions, each of the above devices includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed in this article, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0093] In the case of adopting an integrated unit, Figure 7 FIG. shows a possible exemplary block diagram of the device involved in the embodiments of this application, and the device 700 may exist in the form of software. The device 700 may include: a processing unit 702 and a communication unit 703. The processing unit 702 is used to control and manage the operations of the device 700. The communication unit 703 is used to support the communication between the device 700 and other devices (such as user-side devices or customer service-side devices). The device 700 may further include a storage unit 701, which is used to store the program code and data of the device 700.
[0094] Among them, the processing unit 702 may be a processor or a controller. For example, it may be a general central processing unit (CPU), a general processor, a digital signal processing (DSP), an application specific integrated circuits (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present invention. The processor may also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of DSP and a microprocessor, and so on. The communication unit 703 may be a communication interface, a transceiver or a transceiver circuit, etc. Among them, the communication interface is a general term, and in specific implementations, the communication interface may include multiple interfaces. The storage unit 701 may be a memory.
[0095] The device 700 may be the server in the above embodiments, or may also be a semiconductor chip disposed in the server. The processing unit 702 may support the device 700 to perform the actions of the server in the method examples above, and the communication unit 703 may support the communication between the device 700 and the user-side device or the customer service-side device; for example, the processing unit 702 is used to support the device 700 to perform Figure 3 Steps 301 to 303 in Figure 4 and the communication unit 703 is used to support the device 700 to perform
[0096] Step 405 in
[0097] Specifically, in one embodiment, after obtaining the user command, the processing unit 702 preprocesses the user command to obtain the original word sequence. The joint training model in the server first performs BERT encoding processing on the original word sequence to obtain the intent vector and the hidden state vectors corresponding to T word segments respectively. The joint training model in the server performs the following processing on the first word segment in the original word sequence, where the first word segment is any one of the T word segments: determining the attention vector of the first word segment according to the hidden state vector and the intent vector of the first word segment; then concatenating the hidden state vector of the first word segment and the attention vector of the first word segment to determine the slot probability vector of the first word segment; and selecting the slot corresponding to the maximum probability value from the K probability values in the slot probability vector of the first word segment as the slot corresponding to the first word segment.
[0098] For the relevant specific implementation, reference may be made to the content in the above method, and details will not be repeated here.
[0099] Refer to Figure 8 As shown in the figure, it is a schematic diagram of a device provided by the present application. The device can be the above-mentioned server, or it can also be a chip set in the server. The device 800 includes: a processor 802, a communication interface 803, and a memory 801. Optionally, the device 800 may further include a communication line 804. Among them, the communication interface 803, the processor 802, and the memory 801 can be interconnected through the communication line 804; the communication line 804 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication line 804 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 8 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0100] The processor 802 can be a CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application solution.
[0101] The communication interface 803 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), wired access networks, etc.
[0102] The memory 801 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory can exist independently and be connected to the processor through the communication line 804. The memory can also be integrated with the processor.
[0103] Among them, the memory 801 is used to store computer execution instructions for implementing the solution of this application, and is controlled by the processor 802 for execution. The processor 802 is used to execute the computer execution instructions stored in the memory 801, so as to implement the method provided in the above embodiments of this application.
[0104] Optionally, the computer execution instructions in the embodiments of this application may also be referred to as application code, and this application does not make specific limitations on this.
[0105] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available media may be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid state disks (SSDs)).
[0106] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to this application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes
[0109] An embodiment of the present application also provides a computer storage medium storing computer instructions, which, when running on an electronic device, cause the electronic device to execute the above-mentioned related method steps to implement the method in the above embodiment
[0110] An embodiment of the present application also provides a computer program product, which, when running on a computer, causes the computer to execute the above-mentioned related steps to implement the method in the above embodiment
[0111] In addition, an embodiment of the present application also provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected thereto; wherein, the memory is used to store computer-executable instructions, and when the device runs, the processor may execute the computer-executable instructions stored in the memory, so that the chip executes the methods in the above method embodiments
[0112] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the above-mentioned division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above
[0113] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be discarded or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0114] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0115] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0116] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.
[0117] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.< / cls> < / sep> < / cls> < / pad> < / pad>
Claims
1. A slot recognition method, characterized in that, Including: Preprocessing the user command to obtain an original word sequence; Through encoding and processing the original word sequence with a deep bidirectional pre-trained language understanding model BERT, obtaining an intent vector and hidden state vectors corresponding to T word segments in the original word sequence respectively; For the first word segment among the T word segments, where the first word segment is any one of the T word segments, perform the following processing: Determine the attention vector of the first word segment according to the hidden state vector of the first word segment and the intent vector; Concatenate the hidden state vector of the first word segment and the attention vector of the first word segment to determine the slot probability vector of the first word segment; Determine the slot corresponding to the first word segment according to K probability values in the slot probability vector of the first word segment.
2. The method according to claim 1, wherein Determining the attention vector of the first word segment according to the hidden state vector of the first word segment and the intent vector includes: Determine the attention vector of the first word segment according to the hidden state vector of the first word segment, the intent vector, and the slot probability vector corresponding to the previous word segment of the first word segment; Wherein, the first word segment is any one of the T - 1 word segments except the first word segment among the T word segments.
3. The method according to claim 1 or 2, characterized in that, The preprocessing the user command to obtain an original word sequence includes: Generating a token sequence according to the user command; Randomly sorting the token sequence and dividing the token sequence into multiple batches of token sequences according to the batch size batch_size; Performing truncation or padding operations on the token sequence of each batch to obtain the preprocessed original word sequence.
4. The method according to claim 1, wherein The encoding and processing the original word sequence with BERT to obtain an intent vector and hidden state vectors corresponding to T word segments respectively includes: After performing BERT semantic encoding on the original word sequence, a vector sequence h0, h1, ……, h T is generated, where h0 is the sentence vector encoding information of the user command, and h1, ……, h T are the hidden state vectors corresponding to the T word segments respectively; Generate an intent vector for the user command according to the sentence vector encoding information h0 of the user command, where the intent vector satisfies where y I ∈R 1×I , I represents the number of possible intents of the user command, and the intent corresponding to the maximum probability value in y I is the intent of the user command, h0 is the sentence vector encoding information of the user command, is the bias term, is the weight matrix.
5. The method according to claim 1, wherein Concatenating the hidden state vector of the first word segment and the attention vector of the first word segment to determine the slot probability vector of the first word segment includes: Concatenate the hidden state vector h of the first word segmentation i and the attention vector of the first word segmentation to generate deep vector encoding information, where the deep vector encoding information satisfies where concat is a concatenation operation function, represents the deep vector encoding information after concatenation; Encode the deep vector encoding information Perform a softmax transformation on the logistic regression model to obtain the slot probability vector of the first word segment. The slot probability vector of the first word segment Satisfy where softmax represents the normalized exponential function represents the weight matrix represents the deep vector encoding information, the represents the bias term 6. An electronic device, characterized in that, The electronic device includes a processor and a memory; The memory stores program instructions; The processor is used to run the program instructions stored in the memory, so that the electronic device executes: Preprocessing the user command to obtain an original word sequence; Through encoding and processing the original word sequence with a deep bidirectional pre-trained language understanding model BERT, obtaining an intent vector and hidden state vectors corresponding to T word segments in the original word sequence respectively; For the first word segment among the T word segments, where the first word segment is any one of the T word segments, perform the following processing: Determine the attention vector of the first word segment according to the hidden state vector of the first word segment and the intent vector; Concatenate the hidden state vector of the first word segment and the attention vector of the first word segment to determine the slot probability vector of the first word segment; Determine the slot corresponding to the first word segment according to K probability values in the slot probability vector of the first word segment.
7. The electronic device according to claim 6, wherein When the processor determines the attention vector of the first word segment according to the hidden state vector of the first word segment and the intent vector, it specifically executes: Determine the attention vector of the first token according to the hidden state vector of the first token, the intent vector, and the slot corresponding to the previous token of the first token; Wherein, the first token is any one of the T - 1 tokens except the first token among the T tokens.
8. The electronic device according to claim 6 or 7, characterized in that, When the processor preprocesses the user command to obtain the original word sequence, it specifically executes: Generate a token sequence according to the user command; Randomly sort the token sequence, and divide the token sequence into multiple batches of token sequences according to the batch size batch_size; Perform truncation or padding operations on the token sequence of each batch to obtain the preprocessed original word sequence.
9. The electronic device according to claim 6, wherein When the processor obtains the intent vector and the hidden state vectors corresponding to the T tokens respectively by performing BERT encoding processing on the original word sequence, it specifically executes: After performing BERT semantic encoding on the original word sequence, a vector sequence h0, h1, ……, h is generated T , where h0 is the sentence vector encoding information of the user command, and h1, ……, h T are the hidden state vectors corresponding to the T word segments respectively; Generate an intent vector for the user command according to the sentence vector encoding information h0 of the user command, where the intent vector satisfies where y I ∈R 1×I , I represents the number of possible intents of the user command, and the intent corresponding to the maximum probability value in y I is the intent of the user command, h0 is the sentence vector encoding information of the user command, is the bias term, is the weight matrix.
10. The electronic device according to claim 6, characterized in that, When the processor concatenates the hidden state vector of the first token and the attention vector of the first token to determine the slot probability vector of the first token, it specifically executes: Concatenate the hidden state vector h of the first word segmentation i and the attention vector of the first word segmentation to generate deep vector encoding information, where the deep vector encoding information satisfies where concat is the concatenation operation function, denotes the deep vector encoding information after concatenation; Encode the deep vector encoding information Perform a softmax transformation on the logistic regression model to obtain the slot probability vector of the first word segmentation. The slot probability vector of the first word segmentation Satisfy where softmax represents the normalized exponential function represents the weight matrix represents the deep vector encoding information, the represents the bias term 11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes program instructions, which when run on the processor, cause the processor to execute the method according to any one of claims 1 to 5.
12. A chip, characterized in that, The chip is coupled to the memory and is used to execute the computer program stored in the memory to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Man-machine conversation understanding method and system for specific field and relevant equipment
CN108334496A
Human-computer interactive speech recognition method and system used for intelligent equipment
CN109785833A