Intention recognition method, device, equipment and storage medium in multi-scenario application

By introducing preset scene identification and hidden layer nonlinear operations in text classifiers, the problems of high memory usage and large tasks in traditional text classifiers in multi-scene applications are solved, and more efficient intention recognition and system management are achieved.

CN111460829BActive Publication Date: 2025-05-02PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010156965.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-09
Publication Date
2025-05-02
Estimated Expiration
2040-03-09

AI Technical Summary

Technical Problem

In traditional text classifier applications, multiple text classifiers need to be configured for multiple task scenarios, resulting in high memory usage and large terminal processing tasks, and difficult to effectively manage.

Method used

The feature vectors of different scene texts are extracted through an input layer, and the preset scene identification control feature vectors enter the corresponding hidden layer for nonlinear operations to identify the intent of the target person.

Benefits of technology

It reduces the terminal's memory space, reduces the amount of tasks that the terminal needs to process, and improves the efficiency and manageability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111460829B_ABST
    Figure CN111460829B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses an intention recognition method, device, equipment and storage medium under multi-scenario application, which are used for controlling a scene feature vector to enter a corresponding scene hidden layer for nonlinear operation according to a scene identifier, thereby obtaining the intention of a target person, reducing the occupation of terminal memory space, and reducing the amount of tasks that the terminal needs to process. The method of the present invention comprises: obtaining scene text data from a plurality of voice scene information of the target person; inputting the scene text data into a preset scene model to obtain a predicted scene label; matching the predicted scene label with a plurality of preset scene identifiers to obtain a target scene identifier; transmitting the scene text data to a target scene hidden layer according to the target scene identifier to obtain a target scene vector; obtaining a target label according to the target scene vector and a preset output layer, and obtaining an intention scene label according to the target label; and determining the intention of the target person according to the intention scene label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method, device, equipment and storage medium for intention recognition in multi-scenario applications. Background Art

[0002] With the rapid development of computer technology, the application fields of artificial intelligence technology are becoming more and more extensive. Artificial intelligence technology can enable machines to interpret user intentions and help users complete specific tasks, bringing users humanized service experience and service convenience.

[0003] In artificial intelligence technology, text classifiers are usually required in natural language processing. Compared with classification algorithms based on neural networks, text classifiers have the advantages of maintaining high precision, improving training speed, improving testing speed, and having an accuracy rate comparable to that of deep learning models.

[0004] However, in traditional text classifier applications, a single task often requires the configuration of a corresponding classifier. If there are N task scenarios in the project requirements, then N text classifiers need to be configured. When the model capacity is small, the impact of the number of text classifiers on the terminal may not be obvious. When the model capacity is large, loading multiple text classifiers at the same time will take up a lot of memory space and cause a large load on the terminal's memory space, resulting in unpredictable results. Summary of the invention

[0005] The embodiments of the present invention provide a method, apparatus, device and storage medium for intention recognition in multi-scenario applications, which is used to extract feature vectors of texts in different scenarios through an input layer, use preset scene identifiers to control the feature vectors to enter the corresponding hidden layer for nonlinear operations, and obtain the intention of the target person, thereby reducing the memory space occupied and reducing the amount of tasks that the terminal needs to process.

[0006] The first aspect of an embodiment of the present invention provides an intention recognition method under multi-scenario applications, including: obtaining scene text data from multiple voice scene information of a target person; inputting the scene text data into a preset scene model for label prediction to obtain a predicted scene label; matching the predicted scene label with multiple preset scene identifiers to obtain a target scene identifier, and the target scene identifier belongs to the multiple preset scene identifiers; transmitting the scene text data to a corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, and the target scene hidden layer belongs to the multiple preset scene hidden layers; obtaining a target label according to the target scene vector and a preset output layer, and obtaining an intention scene label according to the target label; determining the intention of the target person according to the intention scene label.

[0007] Optionally, in a first implementation method of the first aspect of the embodiment of the present invention, the scene text data is transmitted to the corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, and the target scene hidden layer belongs to the multiple preset scene hidden layers, including: inputting the scene text data into a preset input layer to obtain a scene feature vector; obtaining a target scene hidden layer from the multiple preset scene hidden layers according to the target scene identifier, and transmitting the scene feature vector to the target scene hidden layer; performing nonlinear calculation on the scene feature vector in the target scene hidden layer to obtain multiple document vectors; and splicing the multiple document vectors to obtain a target scene vector.

[0008] Optionally, in a second implementation method of the first aspect of the embodiment of the present invention, obtaining a target scene hidden layer from the multiple preset scene hidden layers according to the target scene identifier, and transmitting the scene feature vector to the target scene hidden layer includes: matching the target scene identifier with the multiple preset scene hidden layers to obtain a hidden layer probability sequence; searching for a maximum probability in the hidden layer probability sequence to obtain a maximum hidden layer probability; obtaining a target scene hidden layer according to the maximum hidden layer probability, and inputting the scene feature vector into the target scene hidden layer, the target scene hidden layer being a preset scene hidden layer corresponding to the maximum hidden layer probability.

[0009] Optionally, in a third implementation method of the first aspect of the embodiment of the present invention, the step of inputting the scene text data into a preset input layer to obtain a scene feature vector includes: performing word segmentation on the scene text data to obtain word segmentation data; converting the word segmentation data into a word vector based on a preset vector model; and adding the word vectors to obtain a scene feature vector.

[0010] Optionally, in a fourth implementation method of the first aspect of the embodiment of the present invention, the target label is obtained according to the target scene vector and the preset output layer, and the intention scene label is obtained according to the target label, including: based on the preset probability function in the preset output layer, the label probability of the target scene vector is calculated to obtain a target label probability sequence; searching for the two smallest label probabilities in the target label probability sequence, and adding the two label probabilities to construct a binary tree to obtain a node probability; judging whether the node probability is a probability threshold; if the node probability is not a probability threshold, searching for the two smallest label probabilities in the remaining target label probability sequence to construct a binary tree, and stopping the search until the obtained node probability is the probability threshold to obtain a target binary tree; performing Huffman encoding according to the target binary tree to obtain a target label encoding; obtaining a target label based on the target label encoding, and obtaining the intention scene label according to the target label.

[0011] Optionally, in a fifth implementation manner of the first aspect of the embodiment of the present invention, after acquiring scene text data from multiple voice scene information of the target person, and before inputting the scene text data into a preset scene model for label prediction to obtain a predicted scene label, the intention recognition method under multi-scene application also includes: inputting the scene text data into a preset initial scene model to obtain input layer weights, hidden layer weights, output layer weights and output labels; updating the preset initial scene model according to the input layer weights, the hidden layer weights, the output layer weights and the output labels to obtain a preset scene model.

[0012] Optionally, in a sixth implementation manner of the first aspect of the embodiment of the present invention, the preset initial scene model is updated according to the input layer weight, the hidden layer weight, the output layer weight and the output label to obtain the preset scene model, including: judging whether the error value between the output label and the preset true label is greater than an error threshold; if the error value between the output label and the preset true label is greater than the error threshold, back-propagating the error value using the gradient descent method, propagating the error value to the preset input layer, and adjusting the input layer weight, the hidden layer weight and the output layer weight to obtain the preset scene model.

[0013] A second aspect of an embodiment of the present invention provides an intention recognition device for multi-scenario applications, including:

[0014] A first acquisition unit, used to acquire scene text data from a plurality of speech scene information of a target person;

[0015] A prediction unit, used for inputting the scene text data into a preset scene model for label prediction to obtain a predicted scene label;

[0016] A matching unit, configured to match the predicted scene label with a plurality of preset scene identifiers to obtain a target scene identifier, wherein the target scene identifier belongs to the plurality of preset scene identifiers;

[0017] A second acquisition unit is used to transmit the scene text data to a corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, wherein the target scene hidden layer belongs to the plurality of preset scene hidden layers;

[0018] A third acquisition unit, configured to acquire a target label according to the target scene vector and a preset output layer, and obtain an intended scene label according to the target label;

[0019] A determination unit is used to determine the intention of the target person according to the intention scenario label.

[0020] Optionally, in a first implementation method of the second aspect of the embodiment of the present invention, the second acquisition unit specifically includes: an acquisition module, used to input the scene text data into a preset input layer to obtain a scene feature vector; a processing module, used to obtain a target scene hidden layer from the multiple preset scene hidden layers according to the target scene identifier, and transmit the scene feature vector to the target scene hidden layer; a calculation module, used to perform nonlinear calculation on the scene feature vector in the target scene hidden layer to obtain multiple document vectors; and a splicing module, used to splice the multiple document vectors to obtain a target scene vector.

[0021] Optionally, in a second implementation method of the second aspect of the embodiment of the present invention, the processing module is specifically used to: match the target scene identifier with multiple preset scene hidden layers to obtain a hidden layer probability sequence; search for the maximum probability in the hidden layer probability sequence to obtain a maximum hidden layer probability; obtain the target scene hidden layer based on the maximum hidden layer probability, and input the scene feature vector into the target scene hidden layer, the target scene hidden layer being the preset scene hidden layer corresponding to the maximum hidden layer probability.

[0022] Optionally, in a third implementation method of the second aspect of the embodiment of the present invention, the acquisition module is specifically used to: perform word segmentation on the scene text data to obtain word segmentation data; convert the word segmentation data into word vectors based on a preset vector model; and add the word vectors to obtain a scene feature vector.

[0023] Optionally, in a fourth implementation method of the second aspect of the embodiment of the present invention, the third acquisition unit is specifically used to: calculate the label probability of the target scene vector based on a preset probability function in a preset output layer to obtain a target label probability sequence; search for the two smallest label probabilities in the target label probability sequence, add the two label probabilities, construct a binary tree, and obtain a node probability; determine whether the node probability is a probability threshold; if the node probability is not a probability threshold, search for the two smallest label probabilities in the remaining target label probability sequences, construct a binary tree, and stop searching until the obtained node probability is the probability threshold to obtain a target binary tree; perform Huffman encoding according to the target binary tree to obtain a target label code; obtain a target label based on the target label code, and obtain an intended scene label based on the target label.

[0024] Optionally, in a fifth implementation manner of the second aspect of the embodiment of the present invention, the intention recognition device under multi-scenario application also includes: a fourth acquisition unit, used to input the scene text data into a preset initial scene model to obtain input layer weights, hidden layer weights, output layer weights and output labels; an update unit, used to update the preset initial scene model according to the input layer weights, the hidden layer weights, the output layer weights and the output labels to obtain a preset scene model.

[0025] Optionally, in a sixth implementation method of the second aspect of the embodiment of the present invention, the update unit is specifically used to: determine whether the error value between the output label and the preset true label is greater than an error threshold; if the error value between the output label and the preset true label is greater than the error threshold, then using the gradient descent method to back-propagate the error value, propagate the error value to the preset input layer, and adjust the input layer weight, the hidden layer weight and the output layer weight to obtain a preset scene model.

[0026] A third aspect of an embodiment of the present invention provides an intent recognition device for multi-scenario applications, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the intent recognition method for multi-scenario applications described in any one of the above-mentioned embodiments when executing the computer program.

[0027] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the method described in the first aspect above.

[0028] It can be seen from the above technical solutions that the embodiments of the present invention have the following advantages:

[0029] The present invention provides an intention recognition method, device, equipment and storage medium under multi-scenario application, which obtains scene text data from multiple voice scene information of a target person; inputs the scene text data into a preset scene model for label prediction to obtain a predicted scene label; matches the predicted scene label with multiple preset scene identifiers to obtain a target scene identifier, and the target scene identifier belongs to the multiple preset scene identifiers; transmits the scene text data to a corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, and the target scene hidden layer belongs to the multiple preset scene hidden layers; obtains a target label according to the target scene vector and a preset output layer, and obtains an intention scene label according to the target label; determines the intention of the target person according to the intention scene label. The embodiment of the present invention controls the scene feature vector to enter the corresponding scene hidden layer for nonlinear operation according to the scene identifier to obtain the intention of the target person, thereby reducing the occupation of the terminal memory space and reducing the amount of tasks that the terminal needs to process. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of an embodiment of the method for intention recognition in multiple scenarios of the present invention;

[0031] Figure 2 It is a schematic diagram of another embodiment of the method for intention recognition in multi-scenario applications of the present invention;

[0032] Figure 3 A schematic diagram of an embodiment of an intention recognition device for multi-scenario applications in the present invention;

[0033] Figure 4 It is a schematic diagram of another embodiment of the intention recognition device in multi-scenario applications of the present invention;

[0034] Figure 5 A schematic diagram of an embodiment of an intention recognition device for multi-scenario applications in the present invention. DETAILED DESCRIPTION

[0035] The embodiments of the present invention provide a method, apparatus, device and storage medium for intention recognition in multi-scenario applications, which is used to extract feature vectors of texts in different scenarios through an input layer, use preset scene identifiers to control the feature vectors to enter the corresponding hidden layer for nonlinear operations, and obtain the intention of the target person, thereby reducing the memory space occupied and reducing the amount of tasks that the terminal needs to process.

[0036] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0037] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] See also Figure 1 In one embodiment of the present invention, an intention recognition method for multi-scenario applications includes:

[0039] 101. Obtain scene text data from multiple voice scene information of the target person.

[0040] The terminal obtains scene text data from various voice scene information of the target person.

[0041] In this embodiment, the terminal obtains scene text data from a variety of task scenarios. For example, the terminal may obtain scene text data A from dialogue scenario A and obtain scene text data B from dialogue scenario B.

[0042] For ease of understanding, the following is an explanation based on specific scenarios:

[0043] The target person says the dialogue information of "Excuse me, is this Mr. Zhang?", the terminal obtains the voice scene information of "Excuse me, is this Mr. Zhang?", and then converts the voice scene information of "Excuse me, is this Mr. Zhang?" into corresponding scene text data, and the terminal obtains the scene text data of "Excuse me, is this Mr. Zhang?"

[0044] 102. Input the scene text data into the preset scene model for label prediction to obtain a predicted scene label.

[0045] The terminal inputs the scene text data into the preset scene model for label prediction to obtain the predicted scene label.

[0046] Before formally identifying the target label, the terminal first predicts the scene text data to obtain a predicted scene label. For example, if the scene text data obtained is "Excuse me, is this Mr. Zhang?", the terminal uses the scene text data of "Excuse me, is this Mr. Zhang?" as a training data, and sequentially inputs it into the preset input layer, preset scene hidden layer, and preset output layer in the preset scene model, and finally obtains a predicted label "Confirm Identity".

[0047] 103. Match the predicted scene label with multiple preset scene identifiers to obtain a target scene identifier, where the target scene identifier belongs to the multiple preset scene identifiers.

[0048] The terminal matches the predicted scene label with multiple preset scene identifiers to obtain a target scene identifier, and the target scene identifier belongs to the multiple preset scene identifiers.

[0049] The terminal matches the predicted scene label with multiple preset scene identifiers, and selects a preset scene identifier that is most similar to the predicted scene label from the multiple preset scene identifiers as the target scene identifier.

[0050] 104. According to the target scene identifier, the scene text data is transmitted to the corresponding target scene hidden layer to obtain a target scene vector, and the target scene hidden layer belongs to a plurality of preset scene hidden layers.

[0051] The terminal transmits the scene text data to the corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, and the target scene hidden layer belongs to a plurality of preset scene hidden layers.

[0052] The terminal first inputs the scene text data into a preset input layer to obtain a scene feature vector. Then, the terminal selects from multiple preset scene hidden layers according to the target scene identifier to obtain a target scene hidden layer. The terminal inputs the scene feature vector into the target scene hidden layer corresponding to the target scene identifier to obtain multiple document vectors. Finally, the terminal concatenates the multiple document vectors to obtain a target scene vector.

[0053] 105. According to the target scene vector and the preset output layer, the target label is obtained, and the intention scene label is obtained according to the target label.

[0054] The terminal obtains the target label according to the target scene vector and the preset output layer, and obtains the intended scene label according to the target label.

[0055] Since the scene model cannot process textual intents such as intent scene labels, the terminal can only process the textual intents such as scene text data into target labels through the preset scene model, and then the terminal obtains the intent scene label based on the target label.

[0056] For example, the intent of the scene text data "Excuse me, are you Mr. Zhang?" is to "confirm identity". We can define the category of the intent scene label of the target person "confirm identity" as "label 1", input the target scene vector into the output layer, and obtain the target label of "label 1". Based on the target label of "label 1", we can obtain the intent scene label of "confirm identity".

[0057] 106. Determine the target person’s intention based on the intention scenario label.

[0058] The terminal determines the target person's intention based on the intention scenario label.

[0059] For example, the intent of the scene text data "Excuse me, are you Mr. Zhang?" is "confirm identity". We can define the category of the intent of "confirm identity" as "label 1", input the target scene vector into the output layer, and obtain the target label of "label 1". Based on the target label of "label 1", we can obtain the intent scene label of "confirm identity". Based on the intent scene label of "confirm identity", we can determine that the target person's intent is "confirm identity".

[0060] The embodiment of the present invention controls the scene feature vector to enter the corresponding scene hidden layer to perform nonlinear operations according to the scene identification, obtains the intention of the target person, reduces the terminal memory space occupation, and reduces the amount of tasks that the terminal needs to process.

[0061] See also Figure 2 Another embodiment of the method for identifying intentions in multiple scenarios in the embodiment of the present invention includes:

[0062] 201. Obtain scene text data from multiple voice scene information of a target person.

[0063] The terminal obtains scene text data from various voice scene information of the target person.

[0064] In this embodiment, the terminal obtains scene text data from a variety of task scenarios. For example, the terminal may obtain scene text data A from dialogue scenario A and obtain scene text data B from dialogue scenario B.

[0065] 202. Input the scene text data into the preset initial scene model to obtain input layer weights, hidden layer weights, output layer weights and output labels.

[0066] The terminal inputs the scene text data into the preset initial scene model to obtain the input layer weight, hidden layer weight, output layer weight and output label.

[0067] The terminal inputs the scene text data as training data into the initial preset model. The training data is transmitted to the hidden layer through the input layer, and then to the output layer through the hidden layer. During this propagation process, the terminal will obtain the input layer weight, hidden layer weight, output layer weight and output label.

[0068] 203. Update the preset initial scene model according to the input layer weight, the hidden layer weight, the output layer weight and the output label to obtain the preset scene model.

[0069] The terminal updates the preset initial scene model according to the input layer weight, the hidden layer weight, the output layer weight and the output label to obtain the preset scene model.

[0070] Specifically, the terminal determines whether the error value between the output label and the preset true label is greater than the error threshold; if the error value between the output label and the preset true label is greater than the error threshold, the terminal uses the gradient descent method to back-propagate the error value, propagates the error value to the preset input layer, and adjusts the input layer weight, hidden layer weight and output layer weight to obtain the preset scene model.

[0071] The terminal calculates the error value based on the output label and the preset true label. If the error value is greater than the error threshold, the terminal will back propagate the error value. During the back propagation, the error value is calculated using the gradient descent method. Back propagation is performed from the hidden layer to the input layer. During this process, the terminal divides the error value into each layer so that the input layer, hidden layer and output layer obtain their own error signals, and corrects the input layer weight, hidden layer weight and output layer weight through the error signal. The terminal adjusts the weights of each layer to minimize the error value, thereby achieving the purpose of updating the initial preset model. The terminal finally obtains the preset model in the preset text classifier.

[0072] 204. Input the scene text data into a preset scene model for label prediction to obtain a predicted scene label.

[0073] The terminal inputs the scene text data into the preset scene model for label prediction to obtain the predicted scene label.

[0074] Before formally identifying the target label, the terminal first predicts the scene text data to obtain a predicted scene label. For example, if the scene text data obtained is "Excuse me, is this Mr. Zhang?", the terminal uses the scene text data of "Excuse me, is this Mr. Zhang?" as a training data, and sequentially inputs it into the preset input layer, preset scene hidden layer, and preset output layer in the preset scene model, and finally obtains a predicted label "Confirm Identity".

[0075] 205. Match the predicted scene label with multiple preset scene identifiers to obtain a target scene identifier, where the target scene identifier belongs to the multiple preset scene identifiers.

[0076] The terminal matches the predicted scene label with multiple preset scene identifiers to obtain a target scene identifier, and the target scene identifier belongs to the multiple preset scene identifiers.

[0077] The terminal matches the predicted scene label with multiple preset scene identifiers, and selects a preset scene identifier that is most similar to the predicted scene label from the multiple preset scene identifiers as the target scene identifier.

[0078] 206. Transmit the scene text data to the corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, where the target scene hidden layer belongs to a plurality of preset scene hidden layers.

[0079] The terminal transmits the scene text data to the corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, and the target scene hidden layer belongs to a plurality of preset scene hidden layers.

[0080] The terminal first inputs the scene text data into a preset input layer to obtain a scene feature vector. Then, the terminal selects from multiple preset scene hidden layers according to the target scene identifier to obtain a target scene hidden layer. The terminal inputs the scene feature vector into the target scene hidden layer corresponding to the target scene identifier to obtain multiple document vectors. Finally, the terminal concatenates the multiple document vectors to obtain a target scene vector.

[0081] Specifically, the terminal inputs the scene text data into a preset input layer to obtain a scene feature vector; the terminal obtains a target scene hidden layer from multiple preset scene hidden layers according to the target scene identifier, and transmits the scene feature vector to the target scene hidden layer; the terminal performs nonlinear calculations on the scene feature vector in the target scene hidden layer to obtain multiple document vectors; the terminal concatenates multiple document vectors to obtain a target scene vector.

[0082] The terminal inputs the scene text data into the preset input layer for vectorization to obtain a scene feature vector. The terminal controls the scene feature vector to enter the corresponding hidden layer according to the target scene identifier. After the scene feature vector is input into the corresponding hidden layer through the target scene identifier, the scene feature vector no longer enters other preset scene hidden layers. In the target scene hidden layer, the scene feature vector is nonlinearly calculated to obtain multiple document vectors. Since the scene feature vector that enters the target scene hidden layer no longer enters other preset scene hidden layers, the document vectors of other preset scene hidden layers are 0, so the multiple document vectors are spliced ​​to obtain the target scene vector.

[0083] When performing nonlinear calculations on scene feature vectors, the terminal uses the following formula:

[0084]

[0085] Where h is the document vector in the hidden layer of the target scene, c is the number of words, and W is the preset weight matrix from the preset input layer to the hidden layer of the target scene. i is the scene feature vector.

[0086] The terminal obtains one document vector among the multiple document vectors according to the above formula, and calculates other document vectors among the multiple document vectors in the same way to obtain multiple document vectors.

[0087] After obtaining multiple document vectors, the terminal concatenates the multiple document vectors. When concatenating multiple document vectors, the terminal uses the following formula:

[0088] H=H 1 +0+0+0

[0089] In the formula, H is the target feature vector, H 1 and 0 are document vectors.

[0090] Optionally, according to the target scene identifier, a target scene hidden layer is obtained from multiple preset scene hidden layers, and the scene feature vector is transmitted to the target scene hidden layer, which specifically also includes: matching the target scene identifier with multiple preset scene hidden layers to obtain a hidden layer probability sequence; searching for the maximum probability in the hidden layer probability sequence to obtain a maximum hidden layer probability; obtaining the target scene hidden layer according to the maximum hidden layer probability, and inputting the scene feature vector into the target scene hidden layer, the target scene hidden layer being the preset scene hidden layer corresponding to the maximum hidden layer probability.

[0091] The terminal uses the softmax function to calculate the matching results between the target scene identification and multiple preset scene hidden layers to obtain a hidden layer probability sequence, searches for the maximum probability in the hidden layer probability sequence, and obtains the maximum hidden layer probability. The preset scene hidden layer corresponding to the maximum hidden layer probability is the target scene hidden layer.

[0092] Optionally, inputting the scene text data into a preset input layer to obtain the scene feature vector specifically also includes: segmenting the scene text data to obtain segmentation data; converting the segmentation data into word vectors based on a preset vector model; and adding the word vectors to obtain the scene feature vector.

[0093] Text vectorization refers to representing text data as a series of vectors that can express the text's semantics. Before text vectorization, the text data must first be segmented. A word is the smallest meaningful language component that can act independently. Spaces are used as natural delimiters between English words, while Chinese is written in characters, with no obvious distinguishing marks between words. Therefore, Chinese word analysis is the basis and key to Chinese information processing.

[0094] In this embodiment, the forward maximum matching method is used to segment the scene text data. The terminal first takes n characters of the scene text data to be segmented from left to right as the matching field, where n is the number of characters in the longest entry in the large machine dictionary. Then, it searches the large machine dictionary for a match. If a match is successful, this matching field is segmented as a word segmentation data. If the match is unsuccessful, the last character of this matching field is removed, and the remaining string is used as the new word segmentation data for re-matching. The above process is repeated until all word segmentation data is segmented out.

[0095] For example, when the terminal segments the scene text data "Excuse me, is this Mr. Zhang?", the terminal first takes the first three characters "Excuse me" from the sentence and matches these 3 characters with the large machine dictionary. The result shows that the match fails, indicating that there is no such word. Then, the number of characters taken is reduced, and the first two characters "Excuse me" are taken. The result can be successfully matched with the large machine dictionary, indicating that there is such a word. Then, the word "Excuse me" is segmented out to obtain a word segmentation data. Then, the forward maximum matching method is used to segment the remaining five characters. Finally, the terminal obtains the word segmentation data of "Excuse me", "is", "Mr. Zhang", and "right".

[0096] In this embodiment, the preset vector model is an N-gram model. The word segmentation data is input into the n-gram model to obtain word vectors.

[0097] For example, the Chinese words that the word segmentation data of "Excuse me" may be recognized by the N-gram model are "Excuse me" or "Qingwen". The terminal needs to calculate the probability of each word segmentation data appearing through the N-gram model.

[0098] The process of calculating the probability is as follows:

[0099] P 1 = P(Qing| <bos>)×P(ask|please)

[0100] P 2 =P(Sunny| <bos>)×P(闻|晴)

[0101] Determine P through the maximum likelihood value 1 The corresponding word segmentation data is the word segmentation data that conforms to the subsequent semantics. BOS is the large machine dictionary. The probability of "请" appearing in the large machine dictionary and the probability of "问" in the large machine dictionary and the probability matching "请" form a word segmentation vector. The terminal finally adds up the obtained word segmentation vectors to get the feature vector of "请问是张先生吗", that is, the scenario feature vector.

[0102] 207. Obtain the target label based on the target scenario vector and the preset output layer, and obtain the intent scenario label based on the target label.

[0103] The terminal obtains the target label based on the target scenario vector and the preset output layer, and obtains the intent scenario label based on the target label.

[0104] Since the scenario model cannot process the text intent of the intent scenario label internally, the terminal can only process the text intent of the scenario text data into the target label through the preset scenario model, and the terminal then obtains the intent scenario label based on the target label.

[0105] For example, the intent of the scenario text data group "请问您是张先生吗?" is "confirm identity". We can define the category of the intent scenario label of this target person's intent of "confirm identity" as "标签1". Input the target scenario vector into the output layer to get the target label of "标签1", and obtain the intent scenario label of "confirm identity" based on the target label of "标签1".

[0106] Specifically, the terminal calculates the label probability of the target scenario vector based on the preset probability function in the preset output layer to obtain the target label probability sequence; the terminal searches for the two smallest label probabilities in the target label probability sequence, adds the two label probabilities, constructs a binary tree, and obtains the node probability; the terminal determines whether the node probability is the probability threshold; if the node probability is not the probability threshold, the terminal searches for the two smallest label probabilities in the remaining target label probability sequence, constructs a binary tree, and stops searching until the obtained node probability is the probability threshold to obtain the target binary tree; the terminal performs Huffman coding based on the target binary tree to obtain the target label coding; the terminal obtains the target label based on the target label coding, and obtains the intent scenario label based on the target label.

[0107] Huffman coding is a coding method. Huffman coding is a type of variable word length coding. Huffman coding uses a variable length coding table to encode source symbols (such as a letter in text data). The variable length coding table is obtained by a method of evaluating the probability of occurrence of source symbols. Letters with high probability of occurrence use shorter codes, and letters with low probability of occurrence use longer codes. This reduces the average length and expected value of the encoded string, thereby achieving the purpose of lossless data compression.

[0108] In this embodiment, the preset probability function is a softmax function. The terminal first uses the softmax function to calculate the target scene vector to obtain a target label probability sequence, and then adds the two smallest label probabilities to obtain a node probability. The minimum two label probabilities at this time are the left and right subtrees of the first binary tree. If the node probability is not 1 at this time, the minimum two label probabilities are searched in the remaining label probabilities (except the minimum two label probabilities obtained in the last search) to obtain the first node probability. The minimum two label probabilities at this time are the left and right subtrees of the second binary tree. Repeat the steps of constructing a binary tree until the node probability obtained is 1, and the terminal stops searching for the minimum label probability. At this time, the terminal obtains the target binary tree. The terminal specifies the left subtree of each binary tree with a node probability of 1 as 0 in binary, and the right subtree as 1 in binary, and then the terminal goes down along the top of the topmost binary tree to obtain the encoding of each left and right subtree to obtain the target label encoding. The terminal decodes according to the obtained encoding to obtain the target label, and the terminal finally maps the target label to the intended scene label.

[0109] 208. Determine the target person’s intention based on the intention scenario label.

[0110] The terminal determines the target person's intention based on the intention scenario label.

[0111] For example, the intent of the scene text data "Excuse me, are you Mr. Zhang?" is "confirm identity". We can define the category of the intent of "confirm identity" as "label 1", input the target scene vector into the output layer, and obtain the target label of "label 1". Based on the target label of "label 1", we obtain the intent scene label of "confirm identity". Based on the intent scene label of "confirm identity", we determine that the target person's intent is "confirm identity".

[0112] The embodiment of the present invention controls the scene feature vector to enter the corresponding scene hidden layer to perform nonlinear operations according to the scene identification, thereby obtaining the intention of the target person, reducing the terminal memory space occupation, and reducing the amount of tasks that the terminal needs to process.

[0113] The above describes the intention recognition method in multiple scenarios of the present invention. The following describes the intention recognition device in multiple scenarios of the present invention. Figure 3 , an embodiment of the intention recognition device in multi-scenario application in an embodiment of the present invention includes:

[0114] A first acquisition unit 301 is used to acquire scene text data from a plurality of speech scene information of a target person;

[0115] A prediction unit 302, configured to input scene text data into a preset scene model for label prediction to obtain a predicted scene label;

[0116] A matching unit 303, configured to match the predicted scene label with a plurality of preset scene identifiers to obtain a target scene identifier, where the target scene identifier belongs to the plurality of preset scene identifiers;

[0117] A second acquisition unit 304 is used to transmit the scene text data to the corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, where the target scene hidden layer belongs to a plurality of preset scene hidden layers;

[0118] The third acquisition unit 305 is used to acquire a target label according to the target scene vector and the preset output layer, and obtain an intended scene label according to the target label;

[0119] The determination unit 306 is used to determine the intention of the target person according to the intention scene label.

[0120] The embodiment of the present invention controls the scene feature vector to enter the corresponding scene hidden layer to perform nonlinear operations according to the scene identification, obtains the intention of the target person, reduces the terminal memory space occupation, and reduces the amount of tasks that the terminal needs to process.

[0121] See also Figure 4 Another embodiment of the device for identifying intentions in multiple scenarios in the embodiment of the present invention includes:

[0122] A first acquisition unit 301 is used to acquire scene text data from a plurality of speech scene information of a target person;

[0123] A prediction unit 302, configured to input scene text data into a preset scene model for label prediction to obtain a predicted scene label;

[0124] A matching unit 303, configured to match the predicted scene label with a plurality of preset scene identifiers to obtain a target scene identifier, where the target scene identifier belongs to the plurality of preset scene identifiers;

[0125] A second acquisition unit 304 is used to transmit the scene text data to the corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, where the target scene hidden layer belongs to a plurality of preset scene hidden layers;

[0126] The third acquisition unit 305 is used to acquire a target label according to the target scene vector and the preset output layer, and obtain an intended scene label according to the target label;

[0127] The determination unit 306 is used to determine the intention of the target person according to the intention scene label.

[0128] Optionally, the second acquiring unit 304 specifically includes:

[0129] An acquisition module 3041 is used to input the scene text data into a preset input layer to obtain a scene feature vector;

[0130] The processing module 3042 is used to obtain a target scene hidden layer from a plurality of preset scene hidden layers according to the target scene identifier, and transmit the scene feature vector to the target scene hidden layer;

[0131] A calculation module 3043, used for performing nonlinear calculation on the scene feature vector in the target scene hidden layer to obtain multiple document vectors;

[0132] The concatenation module 3044 is used to concatenate multiple document vectors to obtain a target scene vector.

[0133] Optionally, the processing module 3042 is specifically used for:

[0134] Matching the target scene identifier with multiple preset scene hidden layers to obtain a hidden layer probability sequence;

[0135] Search for the maximum probability in the hidden layer probability sequence to obtain the maximum hidden layer probability;

[0136] According to the maximum hidden layer probability, the target scene hidden layer is obtained, and the scene feature vector is input into the target scene hidden layer. The target scene hidden layer is the preset scene hidden layer corresponding to the maximum hidden layer probability.

[0137] Optionally, the acquisition module 3041 is specifically used for:

[0138] Segment the scene text data to obtain segmentation data;

[0139] Based on the preset vector model, the word segmentation data is converted into word vectors;

[0140] Add the word vectors together to get the scene feature vector.

[0141] Optionally, the third acquisition unit 305 is specifically used to: calculate the label probability of the target scene vector based on a preset probability function in a preset output layer to obtain a target label probability sequence;

[0142] Search for the two smallest label probabilities in the target label probability sequence, add the two label probabilities, build a binary tree, and obtain the node probability;

[0143] Determine whether the node probability is a probability threshold;

[0144] If the node probability is not the probability threshold, search for the two smallest label probabilities in the remaining target label probability sequence and construct a binary tree until the node probability is the probability threshold and stop searching to obtain the target binary tree.

[0145] Perform Huffman coding according to the target binary tree to obtain the target label code;

[0146] Based on the target label encoding, the target label is obtained, and the intention scene label is obtained according to the target label.

[0147] Optionally, the intention recognition device for multi-scenario applications further includes:

[0148] A fourth acquisition unit 307, used for inputting the scene text data into a preset initial scene model to obtain input layer weights, hidden layer weights, output layer weights and output labels;

[0149] The updating unit 308 is used to update the preset initial scene model according to the input layer weight, the hidden layer weight, the output layer weight and the output label to obtain the preset scene model.

[0150] Optionally, the optimization unit 308 is specifically configured to:

[0151] Determine whether the error value between the output label and the preset true label is greater than the error threshold;

[0152] If the error value between the output label and the preset true label is greater than the error threshold, the gradient descent method is used to back-propagate the error value, propagate the error value to the preset input layer, and adjust the input layer weight, hidden layer weight and output layer weight to obtain the preset scene model.

[0153] The embodiment of the present invention controls the scene feature vector to enter the corresponding scene hidden layer to perform nonlinear operations according to the scene identification, thereby obtaining the intention of the target person, reducing the terminal memory space occupation, and reducing the amount of tasks that the terminal needs to process.

[0154] above Figure 3 to Figure 4 The intention recognition device for multi-scenario applications in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the intention recognition device for multi-scenario applications in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0155] Combine the following Figure 5 A detailed introduction to the various components of the intent recognition device in multiple scenarios:

[0156] Figure 5 : is a structural diagram of an intent recognition device for multi-scenario applications provided by an embodiment of the present invention. The device 500 for intent recognition for multi-scenario applications may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 501 (for example, one or more processors) and a memory 509, and one or more storage media 508 (for example, one or more massive storage devices) storing application programs 507 or data 506. Among them, the memory 509 and the storage medium 508 may be short-term storage or persistent storage. The program stored in the storage medium 508 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the intent recognition device for multi-scenario applications. Furthermore, the processor 501 may be configured to communicate with the storage medium 508 to execute a series of instruction operations in the storage medium 508 on the intent recognition device 500 for multi-scenario applications.

[0157] The intention recognition device 500 for multi-scenario applications may also include one or more power supplies 502, one or more wired or wireless network interfaces 503, one or more input and output interfaces 504, and / or one or more operating systems 505, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 5 The structure of the intent recognition device for multi-scenario applications shown in the figure does not constitute a limitation on the intent recognition device for multi-scenario applications, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0158] Combine the following Figure 5 A detailed introduction to the various components of the intent recognition device in multiple scenarios:

[0159] Processor 501 is the control center of the intention recognition device under multi-scenario applications, and can be processed according to the intention recognition method under multi-scenario applications. Processor 501 uses various interfaces and lines to connect the various parts of the intention recognition device under the entire multi-scenario application, and by running or executing the software program and / or module stored in the memory 509, and calling the data stored in the memory 509, the preset scene identification controls the feature vector to enter the corresponding hidden layer for nonlinear operation, and obtains the intention of the target person, reducing the amount of tasks that the terminal needs to process. Storage medium 508 and memory 509 are both carriers for storing data. In the embodiment of the present invention, storage medium 508 can refer to an internal memory with a small storage capacity but a fast speed, and memory 509 can be an external memory with a large storage capacity but a slow storage speed.

[0160] The memory 509 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing of the intent recognition device 500 under multi-scenario applications by running the software programs and modules stored in the memory 509. The memory 509 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the intent recognition device under multi-scenario applications, etc. In addition, the memory 509 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. The intent recognition program under multi-scenario applications provided in the embodiment of the present invention and the received data stream are stored in the memory, and when needed, the processor 501 calls from the memory 509.

[0161] When loading and executing computer program instructions on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, a computer instruction may be transmitted from a website site, a computer, a server, or a data center to another website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, twisted pair) or wireless (e.g., infrared, wireless, microwave, etc.) means. Computer-readable storage media may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. Available media may be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., optical disk), or semiconductor media (e.g., solid state disk (SSD)), etc.

[0162] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0163] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / bos> < / bos>

Claims

1. A method for intention recognition in multi-scenario applications, characterized in that: include: Acquire scene text data from multiple voice scene information of the target person; Inputting the scene text data into a preset scene model for label prediction to obtain a predicted scene label; Matching the predicted scene label with a plurality of preset scene identifiers to obtain a target scene identifier, wherein the target scene identifier belongs to the plurality of preset scene identifiers; The scene text data is transmitted to a corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, wherein the target scene hidden layer belongs to a plurality of preset scene hidden layers; According to the target scene vector and the preset output layer, a target label is obtained, and an intended scene label is obtained according to the target label; Determine the intention of the target person according to the intention scenario label; The step of transmitting the scene text data to a corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, wherein the target scene hidden layer belongs to the plurality of preset scene hidden layers, comprises: Inputting the scene text data into a preset input layer to obtain a scene feature vector; According to the target scene identifier, obtaining a target scene hidden layer from the plurality of preset scene hidden layers, and transmitting the scene feature vector to the target scene hidden layer; Performing nonlinear calculation on the scene feature vector in the target scene hidden layer to obtain multiple document vectors; Concatenate the multiple document vectors to obtain a target scene vector; Among them, the nonlinear calculation formula is: In the formula, is the document vector in the hidden layer of the target scene, is the number of words, is the preset weight matrix from the preset input layer to the target scene hidden layer, is the scene feature vector.

2. The method for identifying intentions in multi-scenario applications according to claim 1, characterized in that: The step of acquiring a target scene hidden layer from the plurality of preset scene hidden layers according to the target scene identifier, and transmitting the scene feature vector to the target scene hidden layer comprises: Matching the target scene identifier with the plurality of preset scene hidden layers to obtain a hidden layer probability sequence; Searching for the maximum probability in the hidden layer probability sequence to obtain the maximum hidden layer probability; According to the maximum hidden layer probability, a target scene hidden layer is obtained, and the scene feature vector is input into the target scene hidden layer, where the target scene hidden layer is a preset scene hidden layer corresponding to the maximum hidden layer probability.

3. The method for identifying intentions in multi-scenario applications according to claim 1, characterized in that: The step of inputting the scene text data into a preset input layer to obtain a scene feature vector comprises: Performing word segmentation on the scene text data to obtain word segmentation data; Based on a preset vector model, converting the word segmentation data into a word vector; The word vectors are added together to obtain a scene feature vector.

4. The method for identifying intentions in multi-scenario applications according to claim 1, characterized in that: The step of acquiring a target label according to the target scene vector and a preset output layer, and obtaining an intended scene label according to the target label includes: Based on a preset probability function in a preset output layer, calculating the label probability of the target scene vector to obtain a target label probability sequence; Searching for the two smallest label probabilities in the target label probability sequence, adding the two label probabilities, constructing a binary tree, and obtaining a node probability; Determining whether the node probability is a probability threshold; If the node probability is not the probability threshold, the two smallest label probabilities are searched in the remaining target label probability sequence to construct a binary tree, and the search is stopped when the obtained node probability is the probability threshold to obtain the target binary tree; Perform Huffman coding according to the target binary tree to obtain a target label code; Based on the target label encoding, a target label is obtained, and an intention scene label is obtained according to the target label.

5. The method for identifying intentions in multi-scenario applications according to any one of claims 1 to 4, characterized in that: After acquiring scene text data from multiple speech scene information of the target person, and before inputting the scene text data into a preset scene model for label prediction to obtain a predicted scene label, the intention recognition method under multi-scene application further includes: Inputting the scene text data into a preset initial scene model to obtain input layer weights, hidden layer weights, output layer weights and output labels; The preset initial scene model is updated according to the input layer weight, the hidden layer weight, the output layer weight and the output label to obtain a preset scene model.

6. The method for identifying intentions in multiple scenarios according to claim 5, characterized in that: The updating of the preset initial scene model according to the input layer weight, the hidden layer weight, the output layer weight and the output label to obtain the preset scene model comprises: Determine whether the error value between the output label and the preset true label is greater than an error threshold; If the error value between the output label and the preset true label is greater than the error threshold, the error value is back-propagated using the gradient descent method, the error value is propagated to the preset input layer, and the input layer weight, the hidden layer weight and the output layer weight are adjusted to obtain a preset scene model.

7. An intention recognition device for multi-scenario applications, characterized in that: include: A first acquisition unit, used to acquire scene text data from a plurality of speech scene information of a target person; A prediction unit, used for inputting the scene text data into a preset scene model for label prediction to obtain a predicted scene label; A matching unit, configured to match the predicted scene label with a plurality of preset scene identifiers to obtain a target scene identifier, wherein the target scene identifier belongs to the plurality of preset scene identifiers; A second acquisition unit is used to transmit the scene text data to a corresponding target scene hidden layer according to the target scene identifier to obtain a target scene vector, wherein the target scene hidden layer belongs to a plurality of preset scene hidden layers; A third acquisition unit, configured to acquire a target label according to the target scene vector and a preset output layer, and obtain an intended scene label according to the target label; A determination unit, used to determine the intention of the target person according to the intention scene label; The second acquisition unit specifically includes: An acquisition module, used for inputting the scene text data into a preset input layer to obtain a scene feature vector; A processing module, configured to obtain a target scene hidden layer from the plurality of preset scene hidden layers according to the target scene identifier, and transmit the scene feature vector to the target scene hidden layer; A calculation module, used for performing nonlinear calculation on the scene feature vector in the target scene hidden layer to obtain multiple document vectors; A splicing module, used for splicing the multiple document vectors to obtain a target scene vector; Among them, the nonlinear calculation formula is: In the formula, is the document vector in the hidden layer of the target scene, is the number of words, is the preset weight matrix from the preset input layer to the target scene hidden layer, is the scene feature vector.

8. According to the intention recognition device for multi-scenario applications according to claim 7, the processing module is specifically used to: match the target scene identifier with multiple preset scene hidden layers to obtain a hidden layer probability sequence; search for the maximum probability in the hidden layer probability sequence to obtain a maximum hidden layer probability; obtain the target scene hidden layer based on the maximum hidden layer probability, and input the scene feature vector into the target scene hidden layer, the target scene hidden layer is the preset scene hidden layer corresponding to the maximum hidden layer probability.

9. An intention recognition device for multi-scenario applications, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the intention recognition method for multi-scenario applications as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: It includes instructions, which, when executed on a computer, enable the computer to execute the steps of the method for identifying intentions in multi-scenario applications as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice-input-based image information extraction analysis method and device

    CN103064936A

  • Tax document hierarchical classification method based on multi-tag classification

    CN104199857A