Electronic device, semantic parsing method, medium and human-computer dialogue system thereof
By identifying and utilizing multiple intention prediction slot information in user voice, the accuracy and inefficiency of intention recognition and slot filling in the prior art are solved, and higher semantic analytical accuracy and ability to handle complex scenarios are achieved.
Patent Information
- Application Number
- CN202010970477.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-09-15
AI Technical Summary
The prior art cannot effectively identify multiple intentions in user voice when processing multi-intention statements, resulting in inaccuracy and inefficiency of intention recognition and slot filling, especially when facing single-intention and multi-intention mixed scenarios.
The accuracy and efficiency of slot filling are improved by identifying multiple intentions close to the user's true intention from the user's voice and using these intentions to predict slot information. The specific method includes obtaining the corpus data to be parsed, calculating the degree of correlation between words and intentions and words and slots, and making predictions based on semantic information and the above semantic information.
It improves the accuracy and efficiency of slot filling, thereby improving the accuracy of semantic analysis in human-computer dialogue, and better dealing with multi-intention and single-intention mixed scenarios.
Smart Images

Figure CN114186563B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human - machine dialogue, and particularly relates to an electronic device, a semantic parsing method, a medium and a human - machine dialogue system thereof. Background Art
[0002] With the continuous development of artificial intelligence technology and the deep popularization of various intelligent terminal electronic devices, human - machine dialogue systems are increasingly applied in various intelligent terminal electronic devices, such as: intelligent speakers, smart phones, in - vehicle intelligent voice navigation and other in - vehicle intelligent systems, as well as robots, etc. A human - machine dialogue system uses technologies such as speech recognition, semantic parsing, and language generation to realize dialogue and information exchange between humans and machines.
[0003] Among them, the spoken language understanding task in semantic parsing technology includes two subtasks, intention recognition and slot filling. At present, intention recognition and slot filling mainly focus on the recognition of single - intention single - slot, that is, selecting one of the closest intentions from multiple intention recognition result options for the same piece of speech as the recognition result. However, in actual applications, the corpora of multiple intentions may have the same sentence patterns as those of single - intention corpora, and the single - intention classification model cannot distinguish the corpora of multiple intentions, ultimately resulting in a high false - positive rate of the model, that is, a high error rate of the intention recognition and slot filling results. Moreover, the existing intention - slot recognition architecture cannot explicitly model the relationship between intention slots, has poor accuracy in multi - label intention recognition and slot filling, and cannot even be compatible with single - intention and multi - intention mixed scenarios. Summary of the Invention
[0004] Embodiments of the present application provide an electronic device, a semantic parsing method, a medium and a human - machine dialogue system thereof. By identifying multiple intentions close to the user's true intention from the user's speech, and then using the identified multiple intentions to predict slot information, the accuracy of slot filling is improved, and accordingly, the speed or efficiency of slot filling is also improved, thereby improving the accuracy of semantic parsing in human - machine dialogue.
[0005] In a first aspect, embodiments of the present application provide a semantic parsing method, the method comprising: obtaining to - be - parsed corpus data; calculating the intention correlation degree between the characters included in the to - be - parsed corpus data and the intention represented by the to - be - parsed corpus data, and the slot correlation degree between the characters and the slot represented by the to - be - parsed corpus data; predicting the slot of the to - be - parsed corpus data based on the semantic information of the characters, the above - mentioned semantic information of the characters, as well as the intention correlation degree and slot correlation degree of the characters.
[0006] For example, the corpus data can be obtained by converting a user's voice command through speech recognition.
[0007] The degree of intention relevance between the characters included in the corpus data to be parsed and the intention represented by the corpus data to be parsed can be represented by an intention attention vector, and the degree of slot relevance between the characters and the slot represented by the corpus data to be parsed can be represented by a slot attention vector.
[0008] The semantic information of the character can be understood as the word meaning information of the character, that is, the literal meaning of the character and its referential meaning. For example, the word "you" can be a pronoun (indicating the meaning of the title of the other party) or a noun in a specific sentence (such as in the song title, Hello Old Times).
[0009] The above semantic information of the word can be the word meaning information of the previous word that is continuous with the current word in the corpus data. If the current word being processed is the first word, the above semantic information can be the sentence semantic information of the corpus data. The above semantic information is mainly used because it is of great significance for the slot prediction of the current word. The above semantic information can be expressed in the hidden state vector output at the previous moment (relative to the current moment).
[0010] In a possible implementation of the first aspect above, the method above also includes: predicting multiple intents from the corpus data to be parsed; and determining, from the predicted slots, a slot corresponding to each intent in the multiple intents.
[0011] For example, multiple intents are obtained by parsing the corpus data converted from a user's voice command. If the corpus data only contains a single intent, the present application can also be applied to parsing a single intent in such single-intent corpus data. It has certain versatility and provides a better user experience.
[0012] For the multiple intents obtained by parsing, each intent should have at least one slot corresponding to it, and some intents may have three or more slots corresponding to it. This application can accurately sort out the correspondence between multiple intents and multiple slots.
[0013] In a possible implementation of the first aspect, the method further includes: the preceding semantic information includes semantic information of at least one word preceding the word in the corpus data to be parsed.
[0014] For example, when performing semantic parsing on a certain piece of corpus data, at the first moment when predicting the slot for the first character, the semantic information of the context of the first character is the sentence semantic information of this piece of corpus data. At the second moment when predicting the slot for the second character, the semantic information of the context of the second character is the semantic information of the first character, and at this time, the semantic information of the first character contains the information passed to the first character by the sentence semantic information of the corpus data at the first moment. And so on, the semantic information of the context of each subsequent character is the semantic information of the previous character. At the same time, the semantic information of the previous character includes the semantic information passed by the character before the previous character. This transmission relationship is progressive. The correlation degree of the semantic information of two adjacent characters is the largest, and the correlation degree of the semantic information of two non-adjacent characters is smaller or the correlation degree gradually approaches 0 as the number of intervening characters increases.
[0015] In a possible implementation of the above first aspect, the above method further includes: generating the sentence semantic information of the corpus data to be parsed and the semantic information of each character in the corpus data to be parsed.
[0016] For example, the encoder encodes the sentence characters representing the sentence in the corpus data so that the sentence characters can express specific semantic information, and this specific semantic information is the same as or close to the semantic information obtained by human understanding of this sentence. Also, the encoder encodes the word characters of each character in the corpus data so that the word characters can express specific semantic information, and this specific semantic information is the same as or close to the semantic information of each character understood after human understanding of this sentence. In some embodiments, the sentence semantic information of the corpus data can be represented by a sentence vector, and the semantic information of each character in the corpus data can be represented by a word vector.
[0017] In a possible implementation of the above first aspect, the above method further includes: the above method is implemented through a neural network model. The neural network model includes a fully connected layer and a long short-term memory network model.
[0018] For example, a semantic parsing model is trained through a neural network model combined with a BERT model, an attention mechanism, a slot gate mechanism, and a Sigmoid activation function so that it can implement the above method.
[0019] In a possible implementation of the above first aspect, the above method further includes: the sentence semantic information of the corpus data to be parsed, the semantic information of the context of the character, the degree of relevance of the character's intention, and the degree of relevance of the slot are represented in the form of vectors in the neural network model.
[0020] For example, the sentence semantic information of the corpus data to be parsed is represented by a sentence vector, the semantic information of the context of the character is represented by the hidden state vector at the previous moment, and the degree of relevance of the character's intention and the degree of relevance of the slot are represented by an intention attention vector and a slot attention vector respectively.
[0021] In a second aspect, an embodiment of the present application provides a human-machine dialogue method, including: receiving a user voice instruction; converting the user voice instruction into a text-form corpus to be parsed; parsing out the intent and the slots corresponding to each intent in the corpus to be parsed through the above semantic parsing method; and performing the operation corresponding to the user voice instruction or generating a response voice based on the parsed intent and the slots corresponding to each intent.
[0022] In a possible implementation of the above second aspect, the above method further includes: the operation includes one or more of sending an instruction to a smart home device, opening an application software, searching the web, making a phone call, and sending and receiving text messages.
[0023] For example, for the corpus to be parsed obtained by converting a certain user voice instruction through a smart phone, the parsed intents are booking a train ticket and booking a hotel, and the slots corresponding to these two intents are the departure place, the destination, the (hotel) location, and the (hotel) star rating. Then the operations performed by the smart phone may be to open a certain train ticket and hotel reservation software, query the train ticket information corresponding to the departure place and the destination for the user to select, and recommend five-star hotels at a certain location for the user to select. The smart home may include, but is not limited to, a laptop computer, a desktop computer, a tablet computer, a smart phone, a wearable device, a portable music player, a reader device, or other electronic devices capable of accessing the network.
[0024] In a third aspect, an embodiment of the present application provides a human-machine dialogue system, the system includes: a voice recognition module, configured to convert a user voice instruction into a text-form corpus data; a semantic parsing module, configured to execute the above semantic parsing method; a problem-solving module, configured to find a solution for the result parsed by the semantic parsing module; a language generation module, configured to generate a natural language sentence corresponding to the solution; a voice synthesis module, configured to synthesize the natural language sentence into a response voice; and a dialogue management module, configured to schedule the voice recognition module, the semantic parsing module, the problem-solving module, the language generation module, and the voice synthesis module to cooperate with each other to implement human-machine dialogue.
[0025] In a fourth aspect, an embodiment of the present application provides a readable medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device is caused to execute the above semantic parsing method or the above human-machine dialogue method.
[0026] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a memory, configured to store instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, configured to execute the above semantic parsing method or the above human-machine dialogue method. Brief Description of the Drawings
[0027] Figure 1 It is a schematic software block diagram of a common human - machine dialogue system;
[0028] Figure 2 It is a schematic diagram of a human - machine dialogue scenario applicable to the embodiments of the present application;
[0029] Figure 3 It is an exemplary structural schematic diagram of a semantic parsing model in the embodiments of the present application;
[0030] Figure 4 It is a schematic diagram of the processing results of corpus data at different stages in the semantic parsing method of the embodiments of the present application;
[0031] Figure 5 It is a schematic diagram of the training process of a semantic parsing model in the semantic parsing method of the embodiments of the present application;
[0032] Figure 6 It is a schematic diagram of the interaction process between the electronic device 100 and the user in the embodiments of the present application;
[0033] Figure 7 It is a schematic diagram of the interface where the electronic device 100 in the embodiments of the present application executes corresponding operations according to the user's voice command;
[0034] Figure 8 It is an exemplary structural diagram of an electronic device 100 in the embodiments of the present application. Detailed Description of the Embodiments
[0035] The illustrative embodiments of the present application include, but are not limited to, electronic devices and their semantic parsing methods and media.
[0036] As described above, in the prior art, when processing multi - intent statements, there is a problem that multiple intents in the user's voice cannot be recognized, resulting in a high error rate in slot filling when filling slots based on the intent recognition result. To solve this problem, the embodiments of the present application first recognize multiple intents close to the user's true intent from the user's voice, and then use the recognized multiple intents to predict slot information, thereby improving the accuracy of slot filling, and correspondingly improving the speed or efficiency of slot filling, and further improving the accuracy of semantic parsing in human - machine dialogue.
[0037] To facilitate a clear understanding of the embodiments of the present application, the following briefly introduces the technical terms that may be involved in the embodiments of the present application and the related terms of neural networks.
[0038] (1) Natural Language Processing (NLP)
[0039] Natural language is the language of humans. Natural language processing (NLP) is the processing of human language. Natural language processing is a process of systematically analyzing, understanding, and extracting information from text data in an intelligent and efficient manner. By using NLP and its components, we can manage very large amounts of text data, perform a large number of automated tasks, and solve various problems, such as automatic summarization, machine translation (MT), named entity recognition (NER), relation extraction (RE), information extraction (IE), sentiment analysis, speech recognition, question answering, and topic segmentation. Exemplarily, natural language processing tasks can be classified into the following categories.
[0040] Sequence labeling: For each word in a sentence, the model is required to give a classification category based on the context. Such as Chinese word segmentation, part-of-speech tagging, named entity recognition, and semantic role labeling.
[0041] Classification task: The entire sentence outputs a classification value, such as text classification.
[0042] Sentence relation inference: Given two sentences, determine whether these two sentences have a certain nominal relationship. For example, entailment, QA, semantic rewriting, natural language inference.
[0043] Generative task: Output a piece of text to generate another piece of text. Such as machine translation, text summarization, writing poems and sentences, and describing pictures in words.
[0044] (2) Intent: The voice commands input by the user all correspond to the user's intent. It can be understood that the so-called intent is the expression of the user's will. In a human-computer dialogue system, the intent is generally named after "verb + noun", such as querying the weather, booking a hotel, etc. Intent recognition, also known as intent classification, mainly extracts the intent corresponding to the current voice command according to the voice command input by the user. An intent is a set of one or more expression forms. For example, "I want to watch a movie" and "I want to watch an action movie shot by a certain star in a certain year" can belong to the same intent of playing a video. One or more slots can be configured under one intent. Figure 1
[0045] (3) The slot is the key information used to express the user's intention. The accuracy of slot filling directly affects whether the electronic device can match the correct intention. A slot corresponds to a keyword of a certain type of attribute, and the information in the slot can be filled with keywords of the same type, that is, slot filling. For example, the query sentence pattern corresponding to the intention of playing a song can be "I want to listen to {singer}'s {song}". Among them, {singer} is the slot for the singer, and {song} is the slot for the song. Then, if the voice command "I want to listen to the song Red Bean by Singer A" is received from the user, the electronic device (or server) can extract the slot information filled in the {singer} slot from this voice command as: Singer A, and the slot information filled in the {song} slot as: Red Bean. In this way, the electronic device (or server) can identify the user's intention of this voice input as: playing the song Red Bean by Singer A according to these two slot information.
[0046] It can be understood that the semantic parsing method of this application is applicable to various scenarios that require semantic parsing. For example, when the user issues a voice command to a smart electronic device, or when the user has a human-computer conversation with the voice assistant of a smart electronic device, etc. For the convenience of description, the semantic parsing solution of this application will be introduced below based on a human-computer conversation system.
[0047] Currently, as Figure 1 shown, the common human-computer conversation system 110 mainly includes the following 6 technical modules: a speech recognition module 111; a semantic parsing module 112; a problem-solving module 113; a language generation module 114; a dialogue management module 115; a speech synthesis module 116. Among them,
[0048] The speech recognition module 111 is used to implement the recognition conversion from speech to text through Automatic Speech Recognition (ASR) technology, and the recognition result generally outputs corpus data in the form of the top n (n≥1) sentences or word lattices with the highest scores.
[0049] The semantic parsing module 112, also known as the Natural Language Understanding (NLU) module, is mainly used to perform natural language processing (NLP) tasks, including semantic parsing of the corpus data output by the speech recognition module, and identifying the intention (intent) and corresponding slot (slot) expressed by the user. In the embodiments of this application, the function of the semantic parsing module is implemented by a pre-trained semantic parsing model 121, which will be described in detail below and will not be elaborated here.
[0050] The problem-solving module 113 is mainly used to perform reasoning or queries based on the intent and corresponding slots recognized by semantic parsing, so as to feedback to the user a solution corresponding to their intent and corresponding slots.
[0051] The language generation module 114 mainly generates natural language sentences for the solution that needs to be output to the user found by the problem-solving module 113, and feeds them back to the user in text or further converted into speech.
[0052] The dialogue management module 115 is the central hub in the human-machine dialogue system, and is used to schedule the mutual cooperation of other modules in the human-machine interaction system based on the dialogue history, assist the language parsing module in correctly understanding the results of speech recognition, provide help for the problem-solving module, and guide the natural language generation process of the language generation module.
[0053] The speech synthesis module 116 is used to convert the natural language sentences generated by the language generation module into speech output.
[0054] To make the purpose, technical solutions and advantages of this application clearer, the technical solutions of the embodiments of this application will be further described in detail below by combining the accompanying drawings and implementation examples.
[0055] Figure 2 According to the embodiments of this application, a schematic diagram of a human-machine dialogue scenario is shown.
[0056] Specifically, as Figure 2 shown, this application scenario includes an electronic device 100 and a server 200. Among them, the electronic device 100 is a terminal intelligent device for interacting with users, and an application system capable of performing semantic parsing is installed thereon, such as the above-mentioned human-machine dialogue system 110. The electronic device 100 can identify the user's voice commands through the human-machine dialogue system 110 and perform corresponding operations or answer questions raised by the user according to the voice commands. It can be understood that in this application, the electronic device 100 may include but is not limited to smart speakers, smartphones, wearable devices, head-mounted displays, in-vehicle intelligent systems such as in-vehicle intelligent voice navigation, and intelligent robots, portable music players, reader devices, and other electronic devices installed with a human-machine dialogue system or other voice recognition application programs.
[0057] The server 200 can be used to train the semantic parsing model 121, and transplant the trained semantic parsing model 121 to the electronic device 100 for the electronic device 100 to perform semantic parsing and execute corresponding operations. In addition, the server 200 can also perform semantic parsing on the corpus data sent by the electronic device 100 through the trained semantic parsing model 121 and feedback the results to the electronic device 100, and the electronic device 100 further executes corresponding operations.
[0058] It can be understood that the server 200 may include, but is not limited to, the cloud, a server, a laptop computer, a desktop computer, a tablet computer, and other network-accessible electronic devices in which one or more processors are embedded or coupled.
[0059] For the sake of convenience of description, hereinafter, the electronic device 100 is taken as a mobile phone and the server 200 is taken as a server as an example to detail the technical solution of this application. Among them, the electronic device 100 is installed with the above-mentioned human-computer dialogue system 110, and the semantic parsing module 112 in the human-computer dialogue system 110 has a semantic parsing model 121, and this semantic parsing model 121 can perform semantic parsing on the user voice based on the technical solution of this application.
[0060] The semantic parsing model 121 of this application is introduced in detail below.
[0061] The semantic parsing model 121 is a natural language processing model pre-trained by the server 200 based on natural language processing and the above various neural network structures and models. The pre-trained semantic parsing model 121 can extract multiple intents in a single piece of corpus data and predict slots based on multiple intents, so as to accurately identify the intents and corresponding slots in the corpus data, which can greatly improve the accuracy of slot filling.
[0062] Data preprocessing
[0063] The data input into the semantic parsing model 121 is the data obtained after the corpus data is preprocessed. Among them, the corpus data is obtained after the user voice command is recognized and transformed. The preprocessing of the corpus data is a conventional operation for understanding text in the human-computer dialogue system 110 and is one of the natural language processing tasks executed by the semantic parsing module 112. For example, the preprocessing generally includes performing word segmentation on the corpus data, filling the token sequence and the segmentation mark, and creating a mask. The data preprocessing finally obtains a token sequence including the sentence text characters and each word text character in the sentence, a segmentation mark representing the corresponding sentence position of each word, and a mask corresponding to indicating whether each character position in the token sequence is a valid character.
[0064] Among them, the word segmentation process mainly uses a word segmentation tool (such as a Chinese vocabulary table) to divide the corpus data into sentences and individual words that make up the sentences, and assign all possible intent labels to the obtained sentences and all possible slot labels to each word that makes up the sentences. The purpose of the word segmentation process is to prepare data for the next step of filling the token sequence.
[0065] For example, as Figure 4 shown, after the word segmentation process on the corpus data "Please play Hello, My Old Times for me" obtained by converting the voice command, we get:
[0066] 3 possible intent tags: PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE;
[0067] A sentence: Please play Hello, My Old Good Days for me;
[0068] The 10 characters that make up the sentence: Please, for, me, play, you, good, old, time, light;
[0069] Among them, the slot tags assigned to each character are:
[0070] The slot tags corresponding to the five words Please, for, me, play, and you are all O;
[0071] The 3 slots corresponding to you are songName - B, videoName - B, and mediaName - B;
[0072] The 3 slot tags corresponding to the four characters good, old, time, and light are songName - I, videoName - I, and mediaName - I respectively.
[0073] To fill the Token sequence, mainly use the data obtained from word segmentation to obtain a Token sequence that meets the character length requirements by truncating or padding characters in the sentence. Usually, the maximum character length requirement for each sentence in the Token sequence is maxLength = 32. If the character length of the sentence obtained by word segmentation + 2 is greater than maxLength, the sentence needs to be truncated; if the character length of the sentence obtained by word segmentation + 2 is less than maxLength, blank characters need to be padded at the end of the sentence <pad>Increase the sentence character length by 2 to reach the maxLength. The token sequence contains the sentence characters of the entire sentence corresponding to the voice command, as well as the word characters for each character in the corresponding sentence.
[0074] Among them, the +2 in the character length calculation is mainly because the first character in the token sequence is generally <cls>, which marks the sentences obtained by word segmentation (for example, characters <cls>The marked sentence is: Please play "Hello, My Old Times" for me. The ending character in the Token sequence is generally a truncation character <sep> , <sep>Indicates that the sentence preceding it is a complete sentence that meets the single-sentence character length requirement, characters <cls>With <sep>Put the punctuation mark "Sentence 1" on each character between them, indicating that these characters are all the characters that make up Sentence 1. If there are two <sep>, which indicates the first <sep>To the front <cls>Between is the first sentence, two <sep>Between them is the second sentence, and the character length of each sentence + 2 is required to meet the maximum character length requirement. Generally, the character length + 2 included in the user instruction is within the 32-bit maximum character length range.
[0075] Create a mask, mainly to create a mask for each character corresponding to the Token sequence obtained by the above padding. The purpose of creating the mask is to mark the computer-readable marker code for whether each character in the Token sequence expresses valid information. Among them, the characters in the Token sequence <pad>The value of the corresponding created mask element is 0, not a character <pad>The character corresponding to it creates a mask element value of 1.
[0076] Such as Figure 4 As shown, after preprocessing the corpus data, three main data are obtained, namely the Token sequence, the sentence segmentation marker, and the mask generated corresponding to the Token sequence. For Figure 4 the corpus data shown, the above three data are respectively:
[0077] Token sequence: <cls>Please play "Hello, My Old Times" for me <pad> … <pad> <sep>;
[0078] Sentence break marker: Sentence 1 (Please, play, for me, you, good, old, times);
[0079] Mask: {11111111111000000000000000000001}.
[0080] For another example, if the corpus data obtained by recognizing the voice command input by the user is "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station", the three data obtained after the above data preprocessing are respectively:
[0081] Token sequence: <cls>Book me a train ticket from Shanghai to Beijing and reserve a five-star hotel near Beijing Railway Station <sep>;
[0082] Sentence segmentation marker: Sentence 1 (Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station);
[0083] Mask: {11111111111111111111111111111111}.
[0084] The three data obtained after the above data preprocessing of the corpus data can be input into the semantic parsing model 121 for semantic parsing. The semantic parsing model 121 will be introduced in detail below.
[0085] Semantic parsing model 121
[0086] Specifically, as Figure 3 shown, the semantic parsing model 121 includes a BERT encoding layer 1211, an intent classification layer 1212, an attention layer 1213, a slot filling layer 1214, and a post-processing layer 1215.
[0087] 1) BERT encoding layer 1211
[0088] The BERT encoding layer 1211 takes the Token sequence, sentence segmentation marker, and mask obtained after the data preprocessing of the expected data as inputs, and outputs an encoded vector sequence after encoding. Among them, the encoded vector sequence includes sentence vectors and word vectors. The sentence vector represents the semantic information of the corpus data to be parsed, and the word vector contains the semantic information of each word in the corpus data to be parsed. It can be understood that the semantic information and semantic information of the word are the meaning expressions of the corpus data based on natural language understanding. These semantic information and semantic information of the word can express the true intention of the user and the true slot corresponding to the true intention of the user.
[0089] For example, as Figure 4 shown, if the corpus data to be parsed is "Please play Hello, Old Time for me", then in the encoded vector sequence {h0, h1, h2,..., h t} output by the BERT encoding layer 1211, the semantic information represented by the sentence vector h0 may include PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE, Hello, Old Time, Hello, Old Time, etc. The semantic information of the word vectors h1, h2,..., h t may include songName, videoName, mediaName, and the literal meaning of each word that makes up the sentence, where the word corresponding to h1 is Please, the word corresponding to h2 is For, the word corresponding to h3 is Me,..., h 10 corresponds to the word Light.
[0090] For another example, if the corpus data to be parsed is "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station", then the encoding vector sequence {h0,h1,h2,……,h t }, the semantic information represented by the sentence vector h0 may include booking a ticket, booking a hotel, departure place, destination, Shanghai, Beijing, hotel, star, five-star, etc. The word vectors h1,h2,...,h t The word meaning information may include departure place, destination, Shanghai, Beijing, hotel, star, five-star and the literal meaning of each word in the sentence, where h1 corresponds to the word "help", h2 corresponds to the word "me", h3 corresponds to the word "pre", ..., h 30 The corresponding word is store.
[0091] Specifically, the working process of BERT encoding layer 1211 is as follows: Figure 3 As shown:
[0092] The Token sequence, sentence markers, and masks generated by the corresponding Token sequence obtained after data preprocessing will serve as input to the BERT encoding layer 1211.
[0093] BERT encoding layer 1211 identifies valid characters in the Token sequence in sequence by identifying the mask element value <cls>, x1, x2, ……, x t-1 , <sep>and blank characters (invalid characters) (the valid character mask element value is 1, and the blank character mask element value is 0). Among them, t represents the moment when each character in the sentence is processed, hereinafter referred to as time step t or the moment t. For example, at the moment t = 1, the character corresponding to x1 is processed, and at the moment t = 2, the character corresponding to x2 is processed.
[0094] Characters in the Token sequence that mark the sentence <cls>After inputting the trained BERT encoding layer 1211 for semantic encoding, for the character <cls>Assign semantic information to the corpus data to generate a high-dimensional sentence vector h0.
[0095] Characters in the Token sequence <cls>With truncation characters <sep>The characters x1, x2, ……, x among t-1 correspond to each character that makes up a sentence in the corpus data. The characters x1, x2, ……, x t-1 After inputting into the trained BERT encoding layer 1211 for semantic encoding, for the characters x1, x2, ……, x t-1 endow the semantic information of the corpus data, and correspondingly generate high-dimensional word vectors h1, h2, ……, h t .
[0096] The whitespace characters in the Token sequence <pad>The corresponding mask element value is 0, and no word is marked, so it is not used as the input of the BERT encoding layer 1211.
[0097] Based on the sentence vector h0 and word vectors h1, h2, ……, h t Generate an encoded vector sequence as the output of the BERT encoding layer 1211.
[0098] The BERT encoding layer 1211 can be trained based on the BERT model. The specific training process is described in detail below and will not be elaborated here. Among them, the BERT model is a multi-layer bidirectional transformer encoder model based on fine-tuning. The key technical innovation of the BERT model is to apply the bidirectional training of the transformer to language modeling. There are two stages to train the BERT encoding layer using the BERT model: pre-training and fine-tuning. After the BERT model is trained on unlabeled data for different pre-training tasks during pre-training, first, the BERT model is initialized with the pre-training parameters, and then all parameters are fine-tuned using the labeled data from downstream tasks. A remarkable feature of the BERT model is its unified architecture across different tasks, so the difference between its pre-training architecture and the final downstream architecture is very small. The BERT model can further enhance the generalization ability of the word vector model and fully describe the features at the character level, word level, sentence level, and even the inter-sentence relationship level.
[0099] In some other embodiments, the BERT encoding layer 1211 can also be trained by other encoders or encoding models, which are not limited here.
[0100] 2) Intent Classification Layer 1212
[0101] The intent classification layer 1212 is used to predict the candidate intents in the corpus data. Among them, the intent classification layer 1212 can extract multiple intent labels from the corpus data and retain the intent labels that meet the conditions as the candidate intent output.
[0102] Specifically, the intent classification layer 1212 takes the sentence vector h0 obtained from the above BERT encoding layer 1211 as the input. Based on the semantic information represented by the sentence vector h0, the intent classification layer 1212 can extract all possible intent labels and calculate the intent confidence for each extracted intent label to determine whether the intent label meets the output conditions.
[0103] It can be understood that here, the intention confidence represents the degree of proximity between the extracted intention label and the true intention expressed by the corpus data, and can also be called the intention reliability. The higher the intention confidence, the closer the intention is to the true intention expressed by the corpus data. In the intention classification layer 1212, a certain threshold can be set for the intention confidence. For example, the threshold of the intention confidence is set to 0.5. The intention label with an intention confidence greater than or equal to the threshold meets the output condition, and the corresponding intention label will be output as a candidate intention; the intention label with an intention confidence less than the threshold does not meet the output condition, and its corresponding intention label will be deleted and will not be output from the intention classification layer 1212.
[0104] For example, as Figure 4 shown, if the corpus data to be parsed is "Please play Hello, Old Times for me", then the semantic information represented by the sentence vector h0 output by the BERT encoding layer may include 3 possible intention labels: PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE. After this sentence vector h0 is input into the intention classification layer 1212, the intention classification layer 1212 extracts the above 3 possible intention labels and calculates the intention confidence of each intention label to be 0.8, 0.75, and 0.5 respectively. Assuming that the intention confidence threshold set in the intention classification layer 1212 is 0.5, then the intention confidence of the above 3 intention labels all meet the condition of being greater than or equal to 0.5, that is, the above 3 intention labels meet the output condition, and finally the intention classification layer 1212 outputs 3 candidate intentions: PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE.
[0105] Another example, if the corpus data to be parsed is "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station", then the semantic information represented by the sentence vector h0 output by the BERT encoding layer may include 4 possible intention labels: CHECK_TRAIN_NUMBER, BOOK_TRAIN_TICKET, FIND_HOTEL, BOOK_HOTEL. After this sentence vector h0 is input into the intention classification layer 1212, the intention classification layer 1212 extracts the above 4 possible intention labels and calculates the intention confidence of each intention label to be 0.48, 0.87, 0.45, and 0.7 respectively. Assuming that the intention confidence threshold set in the intention classification layer 1212 is 0.5, then among the above 4 intention labels, the intention labels with an intention confidence greater than or equal to 0.5 are BOOK_TRAIN_TICKET and BOOK_HOTEL, which meet the output condition. Then the intention classification layer 1212 outputs 2 candidate intentions: BOOK_TRAIN_TICKET, BOOK_HOTEL. The two intention labels with an intention confidence less than 0.5: CHECK_TRAIN_NUMBER, FIND_HOTEL do not meet the output condition and will not be output from the intention classification layer 1212.
[0106] Specifically, the working process of the intention classification layer 1212 is as Figure 3 shown:
[0107] The intent classification layer 1212 takes the sentence vector h0 in the encoded vector sequence output by the BERT encoding layer 1211 as input. The intent classification layer 1212 extracts all possible intent labels in the semantic information represented by the sentence vector h0 by decoding and activating the sentence vector h0, and calculates the intent confidence y of each intent label. I . Among them, the intent confidence y I After passing through the Sigmoid activation function, the calculation formula is as follows:
[0108] y I = Sigmoid(W I h0 + b I ) (1)
[0109] Among them, I represents the number of intents, W I is the random weight coefficient of the sentence vector h0, and b I represents the bias value.
[0110] The intent classification layer 1212 can be trained by using a fully connected layer (dense) and the Sigmoid function as the activation function. The specific training process is referred to the detailed description below and will not be elaborated here. It can be understood that in some other embodiments, other deep neural networks with the same function as the fully connected layer can be used as the decoder, and other functions with the same function as the Sigmoid function can be used as the activation function of the corresponding deep neural network decoder, which is not limited here.
[0111] 3) Attention layer 1213
[0112] The attention layer 1213 is used to quantify the correlation degree between each word in the corpus data and the intent expressed by the sentence. For example, it can be represented by an intent attention vector, and the intent attention vector can also be understood as an intent context vector; and the attention layer 1213 is also used to quantify the correlation degree between each word in the corpus data and the slot expressed by the sentence. For example, it is represented by a slot attention vector. Among them, the intent attention vector output by the attention layer 1213 will be used as the input of the slot filling layer 1214 to guide slot prediction, so as to improve the accuracy of slot prediction; the slot attention vector output by the attention layer 1213 is used as the bias value of slot calculation to correct the bias of slot prediction calculation.
[0113] Specifically, the attention layer 1213 takes the encoded vector sequence output by the BERT encoding layer 1211 as input, and based on the semantic information represented by the sentence vector h0 and the word vectors h1, h2, ……, h t The semantic information represented, the intent attention vector output by the attention layer can correspondingly be understood as quantifying the degree of relevance of each word vector corresponding character to the sentence expression intent corresponding to the sentence vector, and the slot attention vector output by the attention layer can correspondingly be understood as quantifying the degree of relevance of each word vector corresponding character to the slot expressed by the sentence corresponding to the sentence vector.
[0114] For example, as Figure 4 shown, if the corpus data to be parsed is "Please play Hello, Old Times for me", then in the encoded vector sequence output by the BERT encoding layer, the semantic information represented by the sentence vector h0 may include 3 possible intent labels: PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE, and the semantic information represented by the word vectors h1, h2, ……, h t may include songName, videoName, mediaName, and the literal meaning of each character that makes up the sentence. After inputting the above encoded vector sequence into the attention layer 1213, for example, the intent attention vector C output by the attention layer 1213 I (corresponding to: Please play Hello, Old Times for me, play), where the intent expressed by the sentence "Please play Hello, Old Times for me" may be PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE. The degree of relevance of "play" to the intent expressed by the sentence is relatively high, while the degree of relevance of "you, good, old, time, light, please, for, me" to the intent expressed by the sentence is relatively low or irrelevant. And the slot expressed by the sentence "Please play Hello, Old Times for me" is songName, videoName, mediaName. Then at time t = 1, the slot attention vector output by the attention layer 1213 represents the degree of relevance of "please" to the above 3 slots, for example, the degree of relevance is 0, that is, irrelevant; at time t = 2, the slot attention vector output by the attention layer 1213 represents the degree of relevance of "for" to the above 3 slots, for example, the degree of relevance is 0, that is, irrelevant; and so on. At time t = 6, the slot attention vector output by the attention layer 1213 represents the degree of relevance of "you" to the above 3 slots, for example, the degree of relevance is 0.9, that is, the degree of relevance is relatively large; finally, it can be obtained that the degree of relevance of "you, good, old, time, light" to the above three slots is relatively high, while the degree of relevance of "play, please, for, me" to the slot expressed by the sentence is relatively low or irrelevant.
[0115] For another example, if the corpus data to be parsed is "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station", then the intent attention vector output by the attention layer 1213 is (Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station, book, ticket, train, ticket, hotel), where the intent expressed by the sentence "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station" may be booking a train ticket and booking a hotel. The relevance of "book, ticket, train, ticket, hotel" to the intent expressed by the sentence is relatively high, while the relevance of "Shanghai, Beijing, railway station, five-star, help, me" to the intent expressed by the sentence is relatively low or irrelevant.
[0116] The slots expressed by the sentence "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station" may be the departure place, destination, location, and star rating. Then at time t = 1, the slot attention vector output by the attention layer 1213 represents the relevance of "help" to the above 4 slots. For example, the relevance is 0, which means it is irrelevant. At time t = 2, the slot attention vector output by the attention layer 1213 represents the relevance of "me" to the above 4 slots. For example, the relevance is 0, which means it is irrelevant. And so on. At time t = 6, the slot attention vector output by the attention layer 1213 represents the relevance of "Shanghai" to the above 4 slots. For example, the relevance of "Shanghai" to the slot "departure place" is 0.9, while the relevance to the other 3 slots (destination, location, star rating) is 0.3, indicating that the relevance of "Shanghai" to the slot "departure place" is relatively high, and the relevance to the other 3 slots (destination, location, star rating) is relatively low. And so on. Finally, it can be obtained that the relevance of "Shanghai" to the slot "departure place" is relatively large, the relevance of "Beijing" to the slot "destination" is relatively large, the relevance of "Beijing Railway Station" to the slot "location" is relatively large, the relevance of "five-star" to the slot "star rating" is relatively large, while the relevance of other words in the sentence, such as "help, me" to the above 4 slots is relatively low or irrelevant.
[0117] Specifically, the working process of the attention layer 1213 is as Figure 3 shown:
[0118] The attention layer 1213 takes the encoded vector sequence {h0, h1, h2, ……, h t} output by the BERT encoding layer 1211 as input. The attention layer 1213 extracts the semantic information represented by the sentence vector h0 and the word vectors h1, h2, ……, h t The semantic information represented, and the hidden state vector is output at each time step t, which represents the semantic information and word meaning information extracted before the previous moment (time step t-1) of the corresponding time step t. Among them, the hidden state vector output by the attention layer 1213 at the t-1 moment (time step 0) relative to the first moment (t = 1) is the semantic information corresponding to the sentence vector h0, and the hidden state vector output at the previous moment (time step 1) at t = 2 is the word meaning information of the first character corresponding to the word vector h1. The word meaning information corresponding to the word vector h1 also includes the semantic information passed from the sentence vector h0 at t = 0; the hidden state vector output at the previous moment (time step 2) at t = 3 is the word meaning information of the second character corresponding to the word vector h2. Among them, the word meaning information corresponding to the word vector h2 includes the word meaning information passed from the word vector h1 at t = 1, and the word meaning information corresponding to the word vector h1 also includes the semantic information passed from the sentence vector h0 at t = 0; and so on.
[0119] Further, in the attention layer 1213, the attention vector calculation formula based on the attention mechanism is as follows:
[0120] Attention = W u *tanh(W q *Q + W v *V) (2)
[0121] Among them, when calculating the intention attention vector C I , Q in the above formula (2) represents the sentence vector h0 in the encoded vector sequence input to the attention layer 1213, and V represents the word vectors h1, h2,..., h t in the encoded vector sequence input to the attention layer 1213 at each time step t. The attention vector obtained through the above formula (3) can quantify the correlation degree between each word vector and the sentence vector. Based on the semantic information represented by the sentence vector h0, all possible intention label information is included. Therefore, the sentence vector h0 is combined with the attention vector calculated through the above formula (2) to obtain the intention attention vector C I , and the obtained intention attention vector C I is used to quantify the correlation degree between the character corresponding to each word vector and the sentence expression intention corresponding to the sentence vector.
[0122] When calculating the slot attention vector, Q in the above formula (2) represents the hidden state vector C output by the attention layer 1213 at the previous moment (t-1 moment), and V represents the encoded vector sequence {h0, h1, h2,..., h t}, the attention vector obtained through the above formula (2) can combine the hidden state vector of the previous moment to learn the correlation degree of the word vector processed at the current moment t. The extracted semantic information and / or word meaning information represented by the hidden state vector of the previous moment contains all possible slot label information. Therefore, the hidden state vector C output at the t-1 moment is combined with the attention vector calculated through the above formula (2) to obtain a slot attention vector The obtained slot attention vector is used to quantify the correlation degree between the word corresponding to each word vector and the slot expressed by the sentence corresponding to the sentence vector.
[0123] As described above, the attention layer 1213 can be trained through a Long Short Term Memory (LSTM) model and an attention mechanism. The specific training process is referred to the detailed description below and will not be elaborated here. It can be understood that in some other embodiments, other neural network models with the same functions as the LSTM model and the attention mechanism, as well as other mechanisms for learning the correlation degree between the words in the sentence of natural language and the intention or slot expressed by the sentence, are not limited here.
[0124] 4) Slot filling layer 1214
[0125] The slot filling layer 1214 is used to predict the candidate slots in the corpus data and fill the slot values. Among them, the slot filling layer 1214 can predict multiple slot labels in the corpus data and retain the slot labels that meet the conditions as candidate slots for output.
[0126] Specifically, at the t moment, assuming that the slot filling layer 1214 processes a certain word at the current time, it uses the encoded vector h output by the BERT encoding layer 1211 t , the hidden state vector C output by the attention layer 1213 at the t-1 moment (i.e., the semantic information of the sentence before the word being processed currently or the word meaning information of the word), and the intention attention vector C output by the attention layer 1213 at the t moment I and the slot attention vector Take the input and output the candidate slots at time t. The slot filling layer 1214 predicts possible slot labels based on the above four vectors of the input at each time step t, and calculates the slot position confidence for the predicted slot labels to determine whether the slot labels meet the output conditions. That is, the slot filling layer 1214 obtains the possible slot labels of the corpus data to be parsed based on the encoded vector including the word vectors (containing the semantic information of each word in the corpus data to be parsed), the semantic information or word semantic information of the sentence before the word currently being processed, the degree of relevance between the word currently being processed and the intention expressed by the sentence, and the degree of relevance between the word currently being processed and the slot expressed by the sentence, calculates the slot position confidence for each slot label, and then selects the slot labels that meet the conditions or whose degree of relevance to the actual slot expressed by the corpus data to be parsed exceeds the threshold as the candidate slots for output.
[0127] It can be understood that here, the slot position confidence represents the degree of proximity between the predicted slot label and the actual slot expressed by the corpus data, and can also be called the slot reliability. The higher the slot position confidence of a slot label, the closer it is to the actual slot expressed by the corpus data. In the slot filling layer 1214, a certain threshold can be set for the slot position confidence. For example, set the threshold of the slot position confidence to 0.5. The slot labels with a slot position confidence greater than or equal to the threshold meet the output conditions, and the corresponding slot labels will be output as candidate slots; the slot labels with a slot position confidence less than the threshold do not meet the output conditions, and their corresponding slot labels will be deleted and will not be output from the slot filling layer 1214.
[0128] For example, as Figure 4 As shown in the figure, if the corpus data to be parsed is "Please play You Are the Apple of My Eye for me", and assume that the threshold for setting the slot position confidence in the slot filling layer 1214 is 0.5. Then, among the slot labels predicted by the slot filling layer 1214 for the 5 characters "Please, for, me, play, put", the slot position confidence of the O slot (for example, 0.7) is greater than or equal to 0.5, and the slot position confidence of other slots (such as songName) (for example, 0.3) is less than 0.5. Therefore, the candidate slots corresponding to the 5 characters "Please, for, me, play, put" are all O slots. Among the slot labels predicted by the slot filling layer 1214 for the 5 characters "You, are, the, apple, of", the slot position confidences of songName, videoName, and mediaName (for example, the slot position confidences are 0.86, 0.7, and 0.55 respectively) are greater than or equal to 0.5, while the slot position confidence of the O slot (for example, 0.3) is less than 0.5. Therefore, the candidate slots corresponding to "You" are songName-B, videoName-B, mediaName-B, and the candidate slots corresponding to "are, the, apple, of" are songName-I, videoName-I, mediaName-I. Among them, the B indicates the character at the start position of the name, that is, it means "You" is the first character in the name; the I indicates the character after the start position of the name. Since the O slot represents an empty slot or an unimportant slot, the output of the slot filling layer 1214 finally outputs 3 candidate slots songName, videoName, and mediaName, and fills the slot value "You Are the Apple of My Eye" for each candidate slot.
[0129] For another example, if the corpus data to be parsed is "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station", assume that the threshold for setting the slot position confidence in the slot filling layer 1214 is 0.5. Among them, in the slot labels predicted by the slot filling layer 1214 for "Shang, hai", the slot position confidence of the slot label "departure place" (for example, 0.7) is greater than or equal to 0.5. Therefore, the candidate slots corresponding to these 2 characters "Shang, hai" are all "departure place"; in the slot labels predicted by the slot filling layer 1214 for "Bei, jing", the slot position confidence of the slot label "destination" (for example, 0.8) is greater than or equal to 0.5. Therefore, the candidate slots corresponding to these 2 characters "Bei, jing" are all "destination"; in the slot labels predicted by the slot filling layer 1214 for "Bei, jing, Huo, che, zhan", the slot position confidence of the slot label "location" (for example, 0.75) is greater than or equal to 0.5. Therefore, the candidate slots corresponding to these 5 characters "Bei, jing, Huo, che, zhan" are all "location"; in the slot labels predicted by the slot filling layer 1214 for "Wu, xing, ji", the slot position confidence of the slot label "star rating" (for example, 0.75) is greater than or equal to 0.5. Therefore, the candidate slots corresponding to these 3 characters "Wu, xing, ji" are all "star rating". Therefore, the slot filling layer 1214 finally outputs 4 candidate slots: departure place, destination, location, star rating, and the slot value filled in for the slot (departure place) is (Shanghai), the slot value filled in for the slot (destination) is (Beijing), the slot value filled in for the slot (location) is (Beijing Railway Station), and the slot value filled in for the slot (star rating) is (five-star).
[0130] It should be noted that the slot filling layer 1214 uses the intention attention vector as the input for predicting the slot at each time step t, and when predicting the slot for the first character at t = 1, the sentence vector h0 is used as the initial value input. Since the semantic information represented by the intention attention vector and the sentence vector includes all possible intention labels, the slot filling layer 1214 predicts possible slot labels based on the possible intention labels, and the predicted slot labels are associated with the intention labels. This greatly improves the accuracy of slot prediction, and correspondingly also improves the speed or efficiency of slot prediction.
[0131] Specifically, the working process of the slot filling layer 1214 is as Figure 3 shown:[[]]
[0132] The slot filling layer 1214 uses the encoded vector h output by the BERT encoding layer 1211 at time step t t and the intention attention vector C output by the attention layer 1213 at time step t I and the slot attention vector And the hidden state vector C output by the attention layer 1213 at time t-1 is used as the input. The slot filling layer 1214 first explicitly models the relationship between the intent and the slots based on the slot gate mechanism to obtain the intent attention vector C I And the slot attention vector The fusion vector gS, and then further predicts the slot labels corresponding to each time step t, and calculates the slot position confidence of each slot label.
[0133] Among them, the calculation formula of the fusion vector gS of the intent attention vector and the slot attention vector is as follows:
[0134]
[0135] Among them, v represents the random weight coefficient of the hyperbolic tangent function tanh(x) in the above formula (3), and W represents the random weight coefficient of the intent attention vector C I If W is greater than 1, it means that the influence degree of the intent attention vector C I on slot prediction is greater than that of the slot attention vector If W is less than 1, it means that the influence degree of the intent attention vector C I on slot prediction is less than that of the slot attention vector If W is equal to 1, it means that the influence degree of the intent attention vector C I on slot prediction is the same as that of the slot attention vector
[0136] At each time step t, the slot filling layer 1214 can obtain the slot vector representing the slot label information based on the above four input vectors And then based on the slot vector Calculate the slot position confidence of the corresponding slot label The calculation formula after passing the slot position confidence through the Sigmoid activation function is as follows:
[0137]
[0138] Among them, S is the number of slots, and W S represents the random weight coefficient of the slot vector h i Dec And b S represents the bias value.
[0139] For example, in the above Figure 4 In the shown example, during the process of predicting slot positions for the corpus data "Please play You Are the Apple of My Eye for me", when predicting the slot position for "me" at time t = 3, the slot filling layer 1214 uses the encoded vector h3 (corresponding to: me), the hidden state vector C output by the attention layer 1213 at time t - 1 (corresponding to: for), the intent attention vector C I (corresponding to: Please play You Are the Apple of My Eye for me, play, for) and the slot attention vectors (corresponding to: for, me) as inputs; among them, the hidden state vector C (corresponding to: for) includes the semantic information passed from the word vector (corresponding to: Please), and the word vector (corresponding to: Please) in turn includes the semantic information passed from the sentence vector (corresponding to: Please play You Are the Apple of My Eye for me).
[0140] Since the semantic information corresponding to "me" is a self - referring term, and "me" is not related to the intent and slots expressed by the sentence "Please play You Are the Apple of My Eye for me", therefore, when predicting the slot position for "me", for example, it is calculated that: the slot position confidence of slot label songName is 0.2, the slot position confidence of slot label videoName is 0.3, and the slot position confidence of slot label O is 0.7. Then, the finally predicted slot position for "me" is the O slot, and the O slot generally represents an unimportant slot and will not be used as the output of the slot filling layer 1214 either.
[0141] For example, when predicting the slot position for "you" at time t = 6, the slot filling layer 1214 uses the encoded vector h6 (corresponding to: you), the hidden state vector C output by the attention layer 1213 at time t - 1 (corresponding to: play), the intent attention vector C I (corresponding to: Please play You Are the Apple of My Eye for me, play, for) and the slot attention vectors (corresponding to: play, you) as inputs; among them, the hidden state vector C (corresponding to: you) includes the semantic information passed from the word vector (corresponding to: play), the word vector (corresponding to: play) in turn includes the semantic information passed from its previous word vector (corresponding to: Please), and so on, and the word vector (corresponding to: Please) in turn includes the semantic information passed from the sentence vector (corresponding to: Please play You Are the Apple of My Eye for me).
[0142] Since the semantic information corresponding to "you" is a word in the name of a song or video, "you" has a lower correlation with the intention expressed by the sentence "Please play Hello Old Time for me" and a higher correlation with the slot expressed by the sentence "Please play Hello Old Time for me". Therefore, when predicting the slot for "you", for example, it is calculated that the slot position reliability of the slot label songName is 0.86, the slot position reliability of the slot label videoName is 0.7, the slot position reliability of the slot label mediaName is 0.55, and the slot position reliability of the slot label O is 0.2. Then, the slots finally predicted for "you" are songName, videoName, and mediaName, which are the outputs of the slot filling layer 1214.
[0143] The slot filling layer 1214 can be trained based on the slot-gate mechanism, the LSTM model and the Sigmoid activation function. The specific training process is described in detail below and will not be repeated here. Among them, the slot gate mechanism focuses on learning the relationship between the intention attention vector and the slot attention vector, and obtains a better semantic frame through global optimization. The slot gate mechanism mainly uses the intention context vector to model the relationship between the intention and the slot to improve the slot filling performance. In other embodiments, other deep neural network models with the same functions as the LSTM model can be used as decoders, and other functions with the same functions as the Sigmoid function can be used as the activation function of the corresponding deep neural network decoder, which is not limited here.
[0144] 5) Post-processing layer 1215
[0145] The slot filling layer 1214 is used to sort out the correspondence between the candidate intents and the candidate slots. The result obtained after the candidate intents correspond to the candidate slots is output from the post-processing layer 1215 as a semantic parsing result.
[0146] For example, Figure 4 As shown, if the input corpus is "Please play Hello Old Times for me", the candidate intents (PLAY_MUSIC, PLAY_VIDEO, PLAY_VOICE) output by the intent classification layer 1212 and the candidate slots (songName, videoName, mediaName) output by the slot filling layer 1214 are input into the post-processing layer 1215, and the semantic parsing result output after inference and prediction based on the intent slot mapping table in the post-processing layer 1215 is:
[0147] PLAY_MUSIC songName, videoName, mediaName, hello old times;
[0148] PLAY_VIDEO songName, videoName, mediaName, Hello, My Old Good Days;
[0149] PLAY_VOICE songName, videoName, mediaName, Hello, My Old Good Days.
[0150] Among them, the candidate intents PLAY_MUSIC, PLAY_VIDEO, and PLAY_VOICE are the intents recognized by parsing the corpus data, the candidate slots songName, videoName, and mediaName are the slots obtained by parsing the corpus data, and "Hello, My Old Good Days" is the filled slot value.
[0151] For another example, if the corpus data to be parsed is "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station", after the candidate intents (book train ticket, book hotel) output by the intent classification layer 1212 and the candidate slots (departure place, destination) output by the slot filling layer 1214 are input to the post-processing layer 1215, the semantic parsing result output after reasoning and prediction based on the intent-slot mapping table in the post-processing layer 1215 is:
[0152] Book train ticket, departure place, Shanghai;
[0153] Destination, Beijing;
[0154] Book hotel, location, Beijing Railway Station;
[0155] Star rating, five-star.
[0156] Among them, the candidate intents (book train ticket, book hotel) are the intents recognized by parsing the corpus data, the candidate slots (departure place, destination, location, star rating) are the slots obtained by parsing the corpus data, and Shanghai, Beijing, Beijing Railway Station, and five-star are the filled slot values corresponding to the slots (departure place, destination, location, star rating).
[0157] Specifically, the working process of the post-processing layer 1215 is as Figure 3 shown:
[0158] The post-processing layer 1215 takes the candidate intents obtained by the above-mentioned intent classification layer 1212 and the candidate slots obtained by the slot filling layer 1214 as inputs, and sorts out the corresponding relationship between the candidate intents and the candidate slots based on the intent-slot mapping table obtained during the pre-training process of the semantic parsing model 121. The intent-slot mapping table obtained during the pre-training process of the semantic parsing model 121 is described in detail below and will not be elaborated here.
[0159] It can be understood that the intent slot mapping table is the sorting result of candidate intents and candidate slots obtained based on a large number of sample trainings. Therefore, during the execution of the semantic parsing task, the intent slot mapping table can be continuously updated based on more corpus data in actual applications.
[0160] The above BERT encoding layer 1211, intent classification layer 1212, attention layer 1213, slot filling layer 1214, and post-processing layer 1215 together constitute the semantic parsing model 121. Among them, each layer in the structure of the semantic parsing model 121 needs to be pre-trained with a large amount of sample prediction data to enable it to have the corresponding functions of each layer above. As mentioned above, the semantic parsing model 121 is pre-trained by the server 200. After that, the trained semantic parsing model 121 can either be transplanted to the electronic device 100 to directly execute the semantic parsing task or continue to exist in the server 200 to execute the semantic parsing task requested by the electronic device 100.
[0161] The pre-training process of the semantic parsing model 121 will be introduced in detail below. The pre-training process of the semantic parsing model 121 can refer to the following example.
[0162] As Figure 5 shown, the pre-training process of the semantic parsing model 121 includes:
[0163] 501: The server 200 collects sample corpus data for training the semantic parsing model 121. Among them, the collected sample prediction data should cover as many fields as possible and as many verbs, proper nouns, common nouns, etc. as possible, so that the generalization performance of the trained semantic parsing model 121 will be better.
[0164] The sample corpus data used to train the semantic parsing model 121 needs to be input into each layer structure of the semantic parsing model 121 in batches for training, and each sample prediction data will go through the processing of each layer in the semantic parsing model 121. For the convenience of understanding, several concepts related to sample data are introduced below.
[0165] (a) batch: batch. In deep learning, the loss function required for each update of the parameters is not obtained from a single data label {data: label}, but is weighted by a group of data, and the number of this group of data is the batch size.
[0166] (b) batch size: batch size, the number of samples in a batch. Each time of training, batch size samples are taken from the training set for training.
[0167] (c) Iteration: The number of iterations is the number of times a batch needs to complete one epoch. One iteration is equal to training once using batchsize samples; in one epoch, the number of batches and the number of iterations are equal.
[0168] (d) Epoch: When a complete dataset passes through the neural network once and returns once, this process is called one epoch. That is, one epoch is equal to training once using all the samples in the training set.
[0169] For example, if the training set has 1000 samples and batchsize = 10, then: It takes 100 iterations to train the entire sample set, and 1 epoch. Another example, for a dataset with 2000 training samples. Divide the 2000 samples into batches of size 500, then it takes 4 iterations to complete one epoch.
[0170] 502: The server 200 performs data preprocessing on the sample corpus data to be input into the semantic parsing model 121 through the NLP module. The data preprocessing of the sample corpus data refers to the relevant description of data preprocessing in the BERT encoding layer 1211 above, which will not be elaborated here.
[0171] After data preprocessing, each sample corpus data obtains a Token sequence, a sentence segmentation mark, and a mask corresponding to the Token sequence.
[0172] 503: In one epoch of training, the server 200 inputs the Token sequence, the sentence segmentation mark, and the mask corresponding to the Token sequence obtained by data preprocessing of each sample corpus data into the BERT encoding layer 1211 in the semantic parsing model 121 for training, so that it can output the encoded vector sequence as described in the BERT encoding layer 1211 above.
[0173] The BERT encoding layer 1211 is trained based on the BERT model. During the training process, it is necessary to continuously fine-tune the upstream and downstream parameters of the semantic parsing model 121, so that after a sufficient long time or learning with sufficient sample prediction data, the BERT encoding layer can output the above encoded vector sequence {h0, h1, h2, ……, h t}.
[0174] 504: During one epoch of training, the server 200 respectively inputs the sentence vectors h0 output by the BERT encoding layer 1211 in the above process 503 into the intent classification layer 1212 in the semantic parsing model 121 for training, so that it can output candidate intents as described in the intent classification layer 1212 above, which will not be elaborated here.
[0175] The intent classification layer 1212 is trained based on a fully connected layer and the Sigmoid function as the activation function. During the training process, it is necessary to continuously fine-tune the upstream and downstream parameters of the semantic parsing model 121, so that after sufficient learning of a long enough time or a large amount of sample corpus data, the intent classification layer 1212 can extract all possible intent labels and the intent confidence corresponding to each intent label, and then extract multiple intent labels that meet the output conditions as candidate intents, which are output from the intent classification layer 1212. For specific details, please refer to the above formula (1) and related descriptions, which will not be elaborated here.
[0176] For each sample corpus data, the candidate intents output by the intent classification layer 1212 are input into the post-processing layer 1215.
[0177] 505: During one epoch of training, the server 200 respectively inputs the encoded vector sequence {h0, h1, h2, ……, h t} output by the BERT encoding layer 1211 trained in the above process 503 into the attention layer 1213 in the semantic parsing model 121 for training, so that it can output the intent attention vector C I and the slot attention vector as described in the attention layer 1213 above, which will not be elaborated here.
[0178] The attention layer 1213 is trained based on the attention mechanism and the LSTM model. During the training process, it is necessary to continuously fine-tune the upstream and downstream parameters in the semantic parsing model 121, so that the attention layer 1213 can quantify the correlation degree of each word vector corresponding to the word pair expressing the intent, and the correlation degree of the slot represented by the word pair corresponding to each word vector, and finally output the intent attention vector and the slot attention vector. For specific details, please refer to the above formula (2) and related descriptions, which will not be elaborated here.
[0179] Among them, the LSTM model is a special RNN model, which is proposed to solve the problem of gradient dispersion in the RNN model. Its core is the cell state, temporarily named the cell state, which can also be understood as a conveyor belt, actually the memory space in the whole model, which changes over time. The working principle of the LSTM model can be simply described as follows: (1) forget gate: select to forget some past information; (2) input gate: remember some current information; (3) merge the past and current memories; (4) output gate: select to output some information. The attention mechanism mimics the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external perception to increase the observation fineness of some areas, and can quickly screen out high-value information from a large amount of information using limited attention resources. The attention mechanism can quickly extract the important features of sparse data, and the essential idea of the attention mechanism can be rewritten as the following formula:
[0180]
[0181] Among them, Lx = ||Source|| represents the length of Source. The meaning of the formula is to imagine the constituent elements in Source as a series of <Key, Value> data pairs. At this time, given an element Query in the target Target, by calculating the similarity or correlation between Query and each Key, the weight coefficient of the Value corresponding to each Key is obtained, and then the Values are weighted and summed to obtain the final Attention value. Therefore, in essence, the Attention mechanism is to perform weighted summation on the Value values of the elements in Source, and Query and Key are used to calculate the weight coefficients of the corresponding Values.
[0182] 506: In one epoch of training, the server 200 respectively inputs the encoding vector h output by the BERT encoding layer 1211 trained in the above process 503 at time t t and the intent attention vector C output by the attention layer 1213 trained in the above process 505 at time t I and the slot attention vector and the hidden state vector C output by the LSTM model in the attention layer 1213 at time t - 1 (i.e., the semantic information of the sentence before the current processed word or the semantic information of the word) into the slot filling layer 1214 in the semantic parsing model 121 for training, so that it can output candidate slots as described in the slot filling layer 1214 above, which will not be elaborated here.
[0183] The slot filling layer 1214 is trained based on the slot gate mechanism, with the LSTM model as the decoder and the Sigmoid function as the activation function. During the training process, it is necessary to continuously fine-tune the upstream and downstream parameters of the semantic parsing model 121, so that after sufficient learning of a long enough time or a large amount of sample prediction data, the slot filling layer 1214 can predict all possible slot labels and the slot position confidence corresponding to each slot label for possible intent labels, and then extract multiple candidate slots that meet the output conditions and output them from the slot filling layer 1214. For specific details, refer to the above formulas (3) to (4) and related descriptions, which will not be elaborated here.
[0184] For each sample corpus data, the candidate slots output by the slot filling layer 1214 are input into the post-processing layer 1215.
[0185] 507: The server 200 determines whether the training results of the above processes 501 to 506 meet the training termination condition. If the training results meet the training termination condition, then proceed to 508; if the training results do not meet the training termination condition, then proceed to 509.
[0186] In the embodiments of the present application, the early stopping mechanism can be used to judge the termination of model training. That is, when the number of training epochs reaches the number threshold or the epoch interval from the previous optimal model is greater than the set interval threshold, the training results meet the training termination condition; otherwise, the training results do not meet the training termination condition.
[0187] The early stopping mechanism can enable the trained neural network model to have good generalization performance, that is, it can fit the data well. Its basic meaning is to calculate the performance of the model on the validation set during training. When the performance of the model on the validation set begins to decline, stop training, so as to avoid the problem of overfitting caused by continued training.
[0188] 508: The server 200 terminates the training of the BERT encoding layer 1211, the intent classification layer 1212, the attention layer 1213, and the slot filling layer 1214 in the semantic parsing model 121, and further inputs the large number of candidate intents and candidate slots accumulated during the training in the above processes 502 to 506 into the post-processing layer 1215 in the semantic parsing model 121 for relationship sorting, such as sorting the candidate slots based on the candidate intents, to obtain an intent-slot mapping table. The training of the semantic parsing model ends.
[0189] Among them, for each piece of sample corpus data trained through the above processes 502 to 506, candidate intents and candidate slots will be obtained. After being trained for a sufficient number of epochs, the candidate intents and candidate slots input to the post-processing layer 1215 are also sufficient. Before training the post-processing layer 1215, the relationship between the candidate intents and candidate slots is disordered and uncorresponding, that is, there is no mapping formed between the candidate intents and candidate slots. Based on a sufficient number of candidate intents and candidate slots, the post-processing layer 1215 is trained to enable it to sort out the candidate slots based on the candidate intents and output an ordered corresponding relationship between the intents and slots, for example, an intent-slot mapping table is trained. Based on the intent-slot mapping table, the post-processing layer 1215 can accurately and quickly find the corresponding relationship between the candidate intents and candidate slots for the candidate intents and candidate slots input thereto.
[0190] 509: The server 200 continues to input the sample corpus data of the next epoch and repeats the processes 502 to 507 to continue training the semantic parsing model 121.
[0191] It should be noted that in order to eliminate the differences between the candidate intents or candidate slots obtained by the semantic parsing model 121 and the true intents or slots caused by the intent classification or slot filling loss, it is necessary to introduce a joint optimization function during the training of the semantic parsing model 121 to perform joint optimization training on the output candidate intents and candidate slots using the intent classification loss function and the slot filling loss function.
[0192] Specifically, the target loss function for the joint optimization of intents and slots is the sum of the intent classification loss function, the slot filling loss function, and the regularization term of the weights. Among them, the intent classification loss function uses the multi-label Sigmoid cross-entropy loss (Cross Entropy Loss) function, and the slot filling loss function uses the serialized multi-label Sigmoid Cross Entropy Loss function. The calculation formula derivation of Sigmoid Cross Entropy Loss is as follows:
[0193]
[0194] Among them, P(t i = 1|x i ) is the Sigmoid function,
[0195] After adding L2 regularization to the weights, the joint optimization target loss function is obtained, and the formula is:
[0196]
[0197] Among them, L y (y, f(x)) is the intent classification loss function calculated according to the above formula (6), L c (y, f(x)) is the slot filling loss function calculated according to the above formula (6), λ is a hyperparameter, m is the number of data in a batch, and the reason for dividing by 2 is to cancel it out during the derivative calculation; represents the sum of the W parameters of the l-th layer; is a matrix, where k and j represent the rows and columns of the matrix.
[0198] It can be seen from this that the joint optimization function mainly jointly optimizes the intent classification loss or slot filling loss generated during the matrix transformation process in the neural network. After joint optimization through the above formula (7), the semantic parsing model 121 trained by the server 200 can parse the corpus data to be parsed into candidate intents and candidate slots that are closer to the true intent and true slots.
[0199] As described above, after the server 200 completes the pre-training of the semantic parsing model 121, the trained semantic parsing model 121 can either be transplanted to the electronic device 100 to directly perform the semantic parsing task, or continue to exist in the server 200 to perform the semantic parsing task requested by the electronic device 100. Specifically, as Figure 6 shown, the user wakes up the voice assistant of the electronic device 100 to input a voice command. The electronic device 100 extracts one or more intents and slots corresponding to the user's voice command based on the above semantic parsing model 121 through the internal human-computer dialogue system 110. The electronic device 100 further performs corresponding operations based on the recognized intents and slots, such as opening an application software or performing a web search, etc. The specific interaction process between the electronic device 100 transplanted with the semantic parsing model 121 and the user can refer to the following example:
[0200] 601: The electronic device 100 obtains the user's voice command.
[0201] A voice assistant is installed in the electronic device 100, and the user can wake up the voice assistant of the electronic device 100 to send a voice command to the electronic device 100. For example, the electronic device 100 obtains the user's voice command "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station".
[0202] 602: The voice recognition module 111 in the human-computer dialogue system 110 of the electronic device 100 recognizes the obtained user's voice command and converts it into corpus data in text form. For example, the above voice command is converted into corpus data in text form "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station".
[0203] 603: The semantic parsing module 112 in the human-machine dialogue system 110 of the electronic device 100 is used to perform semantic parsing on the corpus data to obtain the semantic parsing results of the slots corresponding to the intents.
[0204] Specifically, the semantic parsing module 112 first preprocesses the corpus data to obtain a Token sequence, sentence segmentation marks, and a mask created for the corresponding Token sequence. Then, the semantic parsing module 112 uses the Token sequence, sentence segmentation marks, and the mask created for the corresponding Token sequence as the input of the semantic parsing model 121 to perform semantic parsing, extract multiple candidate intents and multiple candidate slots. Finally, the semantic parsing model 121 outputs the semantic parsing results after sorting out the corresponding relationships between the multiple candidate intents and the multiple candidate slots. In some embodiments, for simple single-intent corpus, it can also be parsed by the semantic parsing model 121 to extract a single candidate intent and one or more corresponding candidate slots, which is not limited here.
[0205] For example, the semantic parsing result obtained by the semantic parsing module 112 in the human-machine dialogue system 110 for the above corpus data "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station" through the semantic parsing model 121 is:
[0206] Departure place for booking train ticket, Shanghai;
[0207] Destination, Beijing;
[0208] Location for booking hotel, Beijing Railway Station;
[0209] Star rating, five-star.
[0210] 604: The problem-solving module 113 in the human-machine dialogue system 110 of the electronic device 100 searches for corresponding application programs or network resources based on the semantic parsing results obtained by the semantic parsing module 112 to obtain solutions for the intents and slots in the semantic parsing results.
[0211] For example, in the above process 603, the intent and slot mapping results parsed from the user instruction "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station" include 2 intents of the user and 4 slots corresponding to each intent, as well as the slot values filled in each slot. Then, the solution searched by the above problem-solving module 114 is that through the electronic device 100, the installed ticket-booking service software application or travel software application can be opened to query train ticket information and hotel information for the user to select for booking, or according to the user's historical usage records, a certain train number ticket can be default selected to enter the booking interface for the user to confirm. The mobile phone interface is as Figure 7 shown.
[0212] For another example, for the user instruction "Please play 'Hello, My Old Times' for me", as Figure 4 shown, in the intention and slot mapping results parsed from the corpus data recognized for this instruction, it includes 3 intentions of the user and 3 slots corresponding to each intention, as well as the slot values filled in each slot. Then, the electronic device 100 can set to default open the music player software to play the local music "Hello, My Old Times" based on the user's usage habits, or open the audio player software to obtain music or video files about "Hello, My Old Times" for the user to select and play.
[0213] 605: The language generation module 114 in the human-computer dialogue system 110 of the electronic device 100 generates a natural language sentence for the solution found by the problem-solving module 113 and feeds it back to the user through the display interface of the electronic device 100.
[0214] For the user instruction "Help me book a train ticket from Shanghai to Beijing and book a five-star hotel near Beijing Railway Station" in the above process 603, after speech recognition and semantic parsing, the solution searched by the above problem-solving module 114 is that the electronic device 100 can open the installed ticket-booking service software application or travel software application to query train ticket information and hotel information for the user to select and book, or default select a train ticket of a certain train number to enter the booking interface for the user to confirm according to the user's historical usage records. The language generation module 114 can correspondingly generate the train number information of the train ticket or the introduction information of the hotel and feed it back to the user through the display interface of the electronic device 100, as Figure 7 shown.
[0215] For another example, the user voice instruction obtained by the electronic device 100 is to query the weather in the last three days. After speech recognition and semantic parsing, the solution searched by the problem-solving module 113 is to open the browser on the electronic device 100 or open the weather query software installed on the electronic device 100 to search for the weather conditions in the last three days. Correspondingly, the language generation module 114 generates the searched weather conditions into natural language text as follows:
[0216] Today's weather is sunny, 28 - 32°C;
[0217] Tomorrow's weather is sunny, 28 - 33°C;
[0218] The weather on Wednesday is sunny turning to cloudy, 28 - 32°C.
[0219] 606: The dialogue management module 115 in the human-machine dialogue system 110 of the electronic device 100 can schedule other modules based on the user's dialogue history to further improve the accurate understanding of the user's voice command. For example, during the process of the above-mentioned problem-solving module 113 searching for the weather, if the location is not clearly specified in the user's voice command, then the dialogue management module 115 can schedule the problem-solving module 113 to search for Beijing, which the user often queries, as the search address based on the user's dialogue history, and feedback the weather conditions in Beijing for the past three days to the user. The dialogue management module 115 can also schedule the problem-solving module 113 to search for the weather in the user's current location for the past three days based on the location information of the electronic device 100, and further schedule the language generation module 114 to generate the following natural language sentences:
[0220] Beijing area:
[0221] Today's weather is sunny, 28 - 32°C;
[0222] Tomorrow's weather is sunny, 28 - 33°C;
[0223] The weather on Wednesday is sunny turning to cloudy, 28 - 32°C.
[0224] It can be understood that the dialogue management module 115 in the human-machine dialogue system 110 of the electronic device 100 can flexibly schedule other modules in the human-machine dialogue system 110 to perform corresponding functions.
[0225] 607: The speech synthesis module 116 in the human-machine dialogue system 110 of the electronic device 100 further synthesizes and converts the natural language sentences generated by the language generation module 114 into speech and plays it back to the user through the electronic device 100. For example, for the weather conditions generated by the language generation module 114 in the above process 605, it is converted into speech and played for the user to listen to, so that the user can hear the weather conditions without looking at the mobile phone.
[0226] In some other embodiments, the trained semantic parsing model 121 can also continue to exist in the server 200 to execute the semantic parsing tasks requested by the electronic device 100. The user wakes up the voice assistant of the electronic device 100 and inputs a voice command. The electronic device 100 converts the user's voice command into corpus data through the internal human-machine dialogue system 110. The electronic device 100 sends the converted corpus data to the server 200 for semantic parsing by interacting with the server 200. The server 200 extracts multiple candidate intents and candidate slots corresponding to the intents in the user's voice command based on the above semantic parsing model 121. Further, the server 200 feeds back the extracted intent and slot correspondence results to the electronic device 100, and the electronic device 100 further performs corresponding operations based on the recognized intent and slot correspondence, such as opening an application software or performing a web search, etc.
[0227] An exemplary structure of the electronic device 100 will be given below in conjunction with the embodiments of the present application.
[0228] Figure 8 FIG. 4 shows a schematic structural diagram of the electronic device 100 according to an embodiment of the present application.
[0229] The electronic device 100 may include a processor 101, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0230] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0231] The electronic device 100 can obtain the user's voice commands and feedback the response voice to the user through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, the electronic device 100 obtains the user's voice commands through the receiver 170B or the microphone 170C, and sends the obtained user's voice commands to the human-machine dialogue system 110 for voice recognition and semantic parsing. According to the semantic parsing result, the corresponding solution is matched, and the corresponding operation is performed through the electronic device 100 to implement the solution corresponding to the semantic parsing result. The human-machine dialogue system 110 can also generate a response voice corresponding to the semantic parsing result and feedback the response voice to the user through the speaker 170A of the electronic device 100 or through the headphone plugged into the headphone interface 170D.
[0232] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 101, or some functional modules of the audio module 170 can be disposed in the processor 101.
[0233] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or a hands-free call through the speaker 170A.
[0234] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the user can listen to the voice by bringing the receiver 170B close to the ear.
[0235] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to implement functions such as collecting sound signals, noise reduction, identifying the sound source, and implementing a directional recording function.
[0236] Among them, the processor 101 can include one or more processing units. For example, the processor 101 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated in one or more processors. The processor 101 realizes the function of the semantic parsing model 121 by running a program. The human-machine dialogue system 110 converts the user's voice command recognition into text corpus data, which is input into the semantic parsing model 121 running on the processor 101 for semantic parsing after data preprocessing to obtain a semantic parsing result.
[0237] The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.
[0238] A memory can also be set in the processor 101 for storing instructions and data. In some embodiments, the memory in the processor 101 is a cache memory. This memory can save the instructions or data that the processor 101 has just used or recycled. If the processor 101 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 101, and thus improves the efficiency of the system.
[0239] In some embodiments, the processor 101 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a general-purpose input / output (GPIO) interface, a SIM interface, and / or a USB interface, etc.
[0240] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present invention is only for illustrative purposes and does not constitute a limitation on the structure of the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0241] The charging management module 140 is used to receive a charging input from a charger. The charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input from the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.
[0242] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 101. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 101, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc.
[0243] The wireless communication function of the electronic device 100 can be implemented by Antenna 1, Antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0244] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0245] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100.
[0246] The wireless communication module 160 can provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs), such as wireless fidelity (Wi-Fi) networks, Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0247] In some embodiments, Antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and Antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.
[0248] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor.
[0249] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0250] The SIM card interface 195 is used to connect the SIM card.
[0251] References in the specification to "one embodiment" or "an embodiment" mean that the particular features, structures, or characteristics described in connection with the embodiment are included in at least one exemplary embodiment or technique disclosed in accordance with the present application. Appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0252] The present application disclosure also relates to an apparatus for performing operations in the text. The apparatus may be specifically constructed for the required purpose or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, application specific integrated circuit (ASIC) or any type of medium suitable for storing electronic instructions, and each may be coupled to a computer system bus. In addition, the computers mentioned in the specification may include a single processor or may be an architecture involving multiple processors for increased computing power.
[0253] The processes and displays presented herein inherently do not involve any particular computer or other apparatus. Various general-purpose systems may also be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform one or more method steps. The structures for various such systems are discussed in the following description. Additionally, any specific programming language sufficient to implement the techniques and embodiments of the present application disclosure may be used. Various programming languages may be used to implement the present disclosure, as discussed herein.
[0254] In addition, the language used in this specification has been principally selected for readability and instructional purposes and may not have been selected to delineate or circumscribe the scope of the disclosed subject matter. Accordingly, the present application disclosure is intended to illustrate rather than limit the scope of the concepts discussed herein.< / pad> < / sep> < / cls> < / cls> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / pad> < / pad> < / cls> < / pad> < / pad> < / sep> < / cls> < / sep> < / sep> < / sep> < / cls> < / sep> < / sep> < / cls> < / cls> < / pad>
Claims
1. A semantic parsing method, characterized in that, The method includes: Obtaining the corpus data to be parsed; Calculating the degree of intention relevance between the words included in the corpus data to be parsed and the intention represented by the corpus data to be parsed, and the degree of slot relevance between the words and the slots represented by the corpus data to be parsed; At each time step, based on the semantic information of the word, the semantic information of the context of the word, as well as the degree of intention relevance and slot relevance of the word, determining a plurality of predicted slots; According to the plurality of predicted slots, determining the slot position confidence of each of the predicted slots; Determining the slots of the corpus data to be parsed for the predicted slots whose slot position confidence exceeds the threshold.
2. The method according to claim 1, characterized in that, It further includes: Predicting a plurality of intentions from the corpus data to be parsed; From the predicted slots, determining the slots corresponding to each intention among the plurality of intentions.
3. The method according to claim 1, characterized in that, The above-mentioned context semantic information includes the semantic information of at least one word located before the word in the corpus data to be parsed.
4. The method according to claim 1, wherein It further includes: Generating the sentence semantic information of the corpus data to be parsed and the semantic information of each word in the corpus data to be parsed.
5. The method according to claim 4, wherein The method is implemented by a neural network model.
6. The method according to claim 5, wherein The neural network model includes a fully connected layer and a long short-term memory network model.
7. The method according to claim 5 or 6, characterized in that, The sentence semantic information of the corpus data to be parsed, the context semantic information of the word, the degree of intention relevance and slot relevance of the word are represented in the form of vectors in the neural network model.
8. A human-machine dialogue method, characterized in that, It includes: Receiving a user voice command; Converting the user voice command into a corpus to be parsed in text form; Parsing the intention and the slots corresponding to each intention in the corpus to be parsed by the semantic parsing method described in any one of claims 1 to 6; Based on the parsed intention and the slots corresponding to each intention, performing the operation corresponding to the user voice command or generating a response voice.
9. The method according to claim 8, wherein The operations include one or more of sending commands to smart home devices, opening application software, searching the web, making calls, and sending and receiving text messages.
10. A human-machine dialogue system, characterized in that, The system includes: A speech recognition module for converting a user voice command into corpus data in text form; A semantic parsing module for performing the semantic parsing method described in any one of claims 1 to 6; A problem-solving module for finding a solution for the result parsed by the semantic parsing module; A language generation module for generating a natural language sentence corresponding to the solution; A speech synthesis module for synthesizing the natural language sentence into a response voice; A dialogue management module for scheduling the speech recognition module, the semantic parsing module, the problem-solving module, the language generation module, and the speech synthesis module to cooperate with each other to achieve human-computer dialogue.
11. A readable medium, characterized in that, Instructions are stored on the readable medium, and when executed on an electronic device, the instructions cause the electronic device to execute the method described in any one of claims 1-6 and claim 9.
12. An electronic device, characterized in that, It includes: A memory for storing instructions executed by one or more processors of the electronic device, and A processor, which is one of the processors of the electronic device, for executing the method described in any one of claims 1-6 and claim 9.
Citation Information
Patent Citations
Human-computer interaction-based natural language processing method, device, apparatus and medium
CN109101545A
Semantic parsing method and device, computer-readable storage medium, and electronic device
CN109241524A