Intention recognition method and device, electronic equipment and computer readable storage medium

By combining neural network model and entity linking method in intention recognition, the initial recognition results are corrected, and the problem of poor accuracy of complex user input and statement ambiguity entity recognition in the prior art is solved, and the accuracy of intention recognition and word slot filling is improved.

CN120106086APending Publication Date: 2025-06-06BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311651110.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the prior art deals with complex user input and entities with statement ambiguity, the accuracy of intention recognition and word slot filling results is poor, which is prone to deviation from the user's real thoughts.

Method used

By obtaining input statements, intent recognition and word slot filling are performed based on the preset neural network model, and the initial recognition result is determined; then the reference filling result is determined based on the entity linking method, and the initial recognition result is corrected based on the reference filling result to improve the accuracy of intent recognition.

Benefits of technology

The intent and slot filling results are determined through the processing process of two different branches, and the reference filling results with higher accuracy of the entity linking method is corrected, which improves the accuracy of the corresponding intent recognition and slot filling results of the user input statement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106086A_ABST
    Figure CN120106086A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an intention recognition method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining an input statement, carrying out the intention recognition and word slot filling of the input statement based on a preset neural network model, and determining an initial recognition result; determining a reference filling result corresponding to the input statement based on an entity linking method; and correcting the initial recognition result according to the reference filling result to determine an intention recognition result. Therefore, the intention and the word slot filling result corresponding to the input statement are determined through the processing processes of two different branches, and the initial recognition result is corrected based on the reference filling result with higher accuracy determined by the entity linking method, so that the corrected intention recognition result can be more fit with the input statement, and the recognition accuracy of the input statement is improved. And the accuracy of the corresponding intention recognition and word slot filling result of the user input statement is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an intent recognition method, device, electronic device and computer-readable storage medium. Background Art

[0002] In recent years, with the rapid development of intelligent dialogue systems, natural language understanding (NLU) plays a vital role in task-oriented dialogue systems. NLU usually includes two tasks: intent classification and slot filling, which aims to make semantic analysis for user expressions. Among them, intent classification is to predict the intention and demand of the user's query based on the user's input information, such as whether to listen to music, ask about the weather, or other. Slot recognition focuses on extracting semantic concepts and identifying the basic information required to execute user intentions. For example, the slot corresponding to the intention of listening to music is the singer's name and the song name. By completing the intent classification and slot filling of the user's input information, the corresponding dialogue strategy can be generated.

[0003] At the same time, the prior art often independently models and predicts intent recognition and slot filling, trains intent classification and slot filling models separately, and uses the two models to predict intent categories and slot filling results respectively during prediction; or uses an end-to-end multi-task joint model to simultaneously train intent classification and slot filling tasks, and uses a single model to simultaneously predict intent and slot recognition results during prediction. However, when dealing with complex user input and ambiguous entities in user input sentences, the intent recognition and slot filling results determined by the existing intent recognition and slot filling methods are more likely to deviate from the user's true thoughts and have poor accuracy. Summary of the invention

[0004] In view of this, an object of an embodiment of the present invention is to provide an intent recognition method, device, electronic device and computer-readable storage medium to improve the accuracy of intent recognition and word slot filling results corresponding to user input sentences.

[0005] In a first aspect, an embodiment of the present invention is to provide a method for identifying intent, the method comprising:

[0006] Get input sentence;

[0007] Performing intent recognition and word slot filling on the input sentence based on a preset neural network model to determine an initial recognition result, wherein the initial recognition result includes an initial intent and an initial filling result;

[0008] Determine a reference completion result corresponding to the input sentence based on an entity linking method;

[0009] The initial recognition result is modified according to the reference filling result to determine the intention recognition result.

[0010] Further, the determining of the reference completion result corresponding to the input sentence based on the entity linking method includes:

[0011] Perform entity extraction on the input sentence to determine a corresponding entity set;

[0012] Recalling each entity in the entity set to determine a candidate entity set;

[0013] A reference filling result is determined according to each candidate entity in the candidate entity set.

[0014] Furthermore, the extracting entities from the input sentence to determine the corresponding entity set includes:

[0015] Entity extraction is performed on the input sentence based on at least one of a business dictionary, a coarse-grained word segmentation, and an entity extraction model to determine a corresponding entity set.

[0016] Further, the recalling of each entity in the entity set to determine the candidate entity set includes:

[0017] The entities in the entity set are recalled based on a dictionary recall method and a vector recall method respectively to determine a candidate entity set.

[0018] Further, determining a reference filling result according to each candidate entity in the candidate entity set includes:

[0019] Generate corresponding entity description information according to each candidate entity in the candidate entity set;

[0020] Determine a matching score between a corresponding candidate entity and the input sentence according to each entity description information;

[0021] In response to the matching score satisfying a preset condition, the corresponding candidate entity is determined as a reference filling result.

[0022] Further, the determining the matching score between the corresponding candidate entity and the input sentence according to each entity description information includes:

[0023] Determine an input vector corresponding to the entity description information;

[0024] The input vector is fused with feature information of a candidate entity corresponding to the entity description information to determine a matching score between the corresponding candidate entity and the input sentence.

[0025] Furthermore, the entity description information is determined based on a combination of the input sentence and one or more of the following: entity type, entity intent, entity summary, and entity attribute relationship.

[0026] Furthermore, the feature information includes a combination of one or more of the following: popularity of the candidate entity, category of the candidate entity, and vector similarity between the candidate entity and the corresponding entity.

[0027] Further, the modifying the initial recognition result according to the reference filling result to determine the intention recognition result includes:

[0028] In response to the reference filling result being inconsistent with the initial filling result, the initial filling result is corrected to the reference filling result, so as to determine the corrected initial recognition result as the intended recognition result; or,

[0029] The content in the initial recognition result that is inconsistent with the reference filling result is corrected according to the reference filling result, so as to determine the corrected initial recognition result as the intended recognition result.

[0030] Furthermore, the preset neural network model is a multi-task joint learning model or the preset neural network model includes an intent recognition model and a word slot filling model.

[0031] In a second aspect, an embodiment of the present invention is intended to provide an intention recognition device, the device comprising:

[0032] An acquisition unit, used for acquiring an input sentence;

[0033] A first recognition unit, configured to perform intent recognition and word slot filling on the input sentence based on a preset neural network model, and determine an initial recognition result, wherein the initial recognition result includes an initial intent and an initial filling result;

[0034] A second recognition unit, used to determine a reference filling result based on an entity linking method;

[0035] A correction unit is used to correct the initial recognition result according to the reference filling result to determine the intended recognition result.

[0036] In a third aspect, an embodiment of the present invention aims to provide a computer program product, wherein the computer program product comprises a computer program / instructions, and when the computer program / instructions are executed by a processor, the method as described in any one of the above items is implemented.

[0037] In a fourth aspect, an embodiment of the present invention aims to provide an electronic device, comprising a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement a method as described in any one of the above items.

[0038] In a fifth aspect, an embodiment of the present invention aims to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps as described in any one of the above items are implemented.

[0039] The technical solution of the embodiment of the present invention obtains an input sentence, performs intent recognition and word slot filling on the input sentence based on a preset neural network model, and determines an initial recognition result; determines a reference filling result corresponding to the input sentence based on an entity linking method; and corrects the initial recognition result based on the reference filling result to determine the intent recognition result. Thus, by determining the intent and word slot filling result corresponding to the input sentence through the processing of two different branches, and correcting the initial recognition result based on a more accurate reference filling result determined by the entity linking method, the corrected intent recognition result can be made to fit the input sentence more closely, thereby improving the accuracy of the intent recognition and word slot filling results corresponding to the user's input sentence. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0041] Figure 1 is a schematic diagram of an intention recognition method according to an embodiment of the present invention;

[0042] Figure 2 is a flow chart of an intention recognition method according to an embodiment of the present invention;

[0043] Figure 3 is a processing flow chart of the entity linking method according to an embodiment of the present invention;

[0044] Figure 4 is a flow chart of determining a reference filling result according to an embodiment of the present invention;

[0045] Figure 5 is a schematic diagram of a vector recall model processing process according to an embodiment of the present invention;

[0046] Figure 6 is a flow chart of determining a reference filling result according to a candidate entity set according to an embodiment of the present invention;

[0047] Figure 7 is a schematic diagram of an entity sorting model according to an embodiment of the present invention;

[0048] Figure 8 is a flow chart of determining a matching score between a candidate entity and an input sentence according to an embodiment of the present invention;

[0049] Fig. 9 is a schematic diagram of an intention recognition device according to an embodiment of the present invention;

[0050] Fig.10 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.

[0052] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.

[0053] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".

[0054] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0055] The solutions described in this specification and in the examples, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.

[0056] In the prior art, for user input sentences, two models are often used to predict the intent category and the slot filling results respectively, or an end-to-end multi-task joint model is used to simultaneously predict the intent and slot recognition results. However, for complex input sentences (e.g., there are a large number of entities in the sentence) or when there are ambiguous entities in the input sentence (e.g., the entities in the sentence correspond to multiple types of information), the intent recognition and slot filling results determined according to the above method are more likely to deviate from the user's true thoughts and have poor accuracy. In view of this, an embodiment of the present invention aims to provide an intent recognition method to improve the accuracy of intent recognition and slot filling results corresponding to user input sentences.

[0057] Figure 1 FIG. 1 is a schematic diagram of the intention recognition method according to an embodiment of the present invention. Figure 1As shown, in this embodiment, when performing intent recognition on the user's input sentence, on the basis of the original intent recognition processing of the user input sentence only through the preset neural network model, an entity linking module is newly added, and the processing result obtained after the entity linking module performs intent recognition on the user input sentence based on the entity linking method is used as a new intent recognition result. Since the entity linking method can make the linked entity more accurate through rich knowledge, therefore, in this embodiment, by fusing the intent recognition result obtained based on the preset neural network model with the intent recognition result determined based on the entity linking, the intent and word slot filling results in the biased intent recognition results can be corrected, thereby solving the problem that complex user input sentences or word slot recognition are prone to errors, and improving the accuracy of the intent recognition and word slot filling results corresponding to the user input sentence.

[0058] Figure 2 is a flow chart of the intention recognition method of an embodiment of the present invention. Figure 2 As shown, the intention recognition method of this embodiment includes the following steps.

[0059] In step S100, an input sentence is obtained.

[0060] In this embodiment, the input sentence is a sentence to be subjected to intent recognition. Taking intent recognition in a dialogue system as an example, the input sentence is usually a sentence input through a user terminal (ie, a query), including an intent and a word slot. For example, when the input sentence is "Please play work x by person A", the corresponding intent is "play the work", and the word slot has two items, "person" and "work", and one word slot corresponds to one slot. The word slot filling result corresponding to the "person" slot is "A", and the word slot filling result corresponding to the "work" slot is "x". By performing intent recognition on the input sentence, it is necessary to determine that the intent in the corresponding result to be recognized is "play the work", and the slot filling results are "person" = "A", "work" = "x", and "=" means the assignment.

[0061] In step S200, the input sentence is subjected to intent recognition and word slot filling based on a preset neural network model to determine an initial recognition result, wherein the initial recognition result includes an initial intent and an initial filling result.

[0062] In this embodiment, by using a preset neural network model to perform intent recognition and word slot filling on the input sentence, the intent recognition result (that is, the initial recognition result) corresponding to one input sentence can be determined. The preset neural network model is a multi-task joint learning model, or a preset neural network model intent recognition model and word slot filling model.

[0063] In an optional implementation, a multi-task joint learning model is used in this embodiment to combine the two tasks of intent recognition and word slot filling, and learning and prediction are performed simultaneously, and the intent in the initial recognition result (i.e., the initial intent) and the word slot filling result (i.e., the initial filling result) are determined. In the intent recognition task, the goal of the model is to determine the intent or purpose of the input sentence, for example, to play a work, to query a person, to book a restaurant, etc. In the word slot filling task, the goal of the model is to extract specific information items from the input sentence, such as time, place, name, etc., by judging the word slots that exist in the input sentence and filling the corresponding content of the word slots into the corresponding slots.

[0064] Optionally, the multi-task joint learning model in this embodiment can adopt a model based on a shared encoder, a model based on a joint model, a model based on an attention mechanism, a model based on a multi-task learning framework, or other models. Among them, the model based on a shared encoder uses a shared encoder network to learn the representation of the input sentence, and builds two branches on the representation for intent recognition and word slot filling. The shared encoder can be a convolutional neural network (CNN), a recurrent neural network (RNN) or a Transformer model, etc. The model based on the joint model solves the intent recognition and word slot filling tasks as a joint sequence labeling problem. The model can use structures such as a recurrent neural network (RNN) or a Transformer model, and simultaneously predicts intent and word slot labels through a shared encoder. The model based on the attention mechanism uses the attention mechanism to model the input sentence, and on this basis, performs intent recognition and word slot filling prediction. The model can use the Transformer model to capture key information in the input sentence through a self-attention mechanism. The model based on the multi-task learning framework uses multi-task learning frameworks such as multi-task learning network (MTL) and joint encoder-decoder (UED) to achieve information sharing between tasks by sharing parameters and encoders, and simultaneously learn and predict intent recognition and word slot filling tasks.

[0065] In another optional implementation, the present embodiment uses an intent recognition model and a slot filling model to respectively perform intent recognition and slot filling on the input sentence, and then determine the initial recognition result corresponding to the input sentence. Among them, the intent recognition model can use a convolutional neural network (CNN), a recurrent neural network (RNN) or a Transformer model, etc., to perform intent classification and determine the intent corresponding to the input sentence, that is, the initial intent. The slot filling model can use a conditional random field (CRF), a sequence to sequence model (Seq2Seq) or a transfer learning model (such as a pre-trained language model BERT, GPT, etc.) to fill the slot and determine the slot filling result corresponding to the input sentence, that is, the initial filling result.

[0066] In step S300, a reference completion result corresponding to the input sentence is determined based on an entity linking method.

[0067] In this embodiment, the intent recognition and word slot filling of the input sentence are performed based on the entity linking method, and the intent recognition result corresponding to another input sentence (that is, the reference filling result) can be determined. At the same time, since the entity linking method determines the recognition result through rich knowledge, the linked entity can be made more accurate. Therefore, the intent recognition of the input sentence based on the entity linking method can improve the accuracy of the intent recognition and word slot filling results.

[0068] It should be understood that in this embodiment, the intent and word slot filling results of the input sentence can be determined only by a preset neural network model, or only by an entity linking method. It is also possible to simultaneously use a preset neural network model and an entity linking method as two branches to respectively determine the initial recognition result and reference filling result corresponding to the input sentence, and combine the initial recognition results and reference filling results determined by the two branches to finally determine the intent and word slot filling result corresponding to the input sentence. Compared to using a single method for intent recognition, using a preset neural network model and an entity linking method for intent recognition at the same time can improve the accuracy and reliability of the intent recognition and word slot filling results corresponding to the input sentence, especially when the input sentence is a complex sentence or contains ambiguous entities and other sentences that are prone to recognition errors, the accuracy of the intent recognition result corresponding to the input sentence is more significantly improved.

[0069] Optionally, in this embodiment, by judging the attributes of the input sentence, it can be determined whether to use both the preset neural network model and the entity linking method to identify the intent of the input sentence. For example, when the number of entities in the input sentence is greater than a predetermined value, the corresponding information type of the entity in the input sentence is greater than a predetermined value, or other situations that are prone to inaccurate recognition, it is determined to use a combination of the preset neural network model and the entity linking method to identify the intent of the input sentence, and then determine the final intent recognition result. This can avoid the problem of low efficiency in determining the intent recognition result due to excessive model calculation and long calculation time when identifying simple input sentences.

[0070] Furthermore, in this embodiment, when a combination of a preset neural network model and an entity linking method is used to identify the intent of an input sentence, a parallel processing calling method is adopted. When the preset neural network model is used to start intent recognition, intent recognition is also performed based on the entity linking method to improve the efficiency of determining the intent recognition results without causing other additional system losses.

[0071] Further, in this embodiment, the processing flow of determining the reference filling result corresponding to the input sentence based on the entity linking method is as follows: Figure 3 As shown, it includes three stages: entity extraction, entity recall and entity sorting to determine the reference filling result. Correspondingly, the method for determining the reference filling result through the above process includes the following: Figure 4 The following steps are shown.

[0072] In step S210, entity extraction is performed on the input sentence to determine the corresponding entity set.

[0073] Optionally, in order to ensure the recall rate of entities, this embodiment extracts entities from input sentences based on at least one of the business dictionary, coarse-grained word segmentation, and entity extraction model to determine the corresponding entity set. In the business dictionary, the input sentence is matched with the entities in the business dictionary based on a preset rule matching method to extract the corresponding entity words. Coarse-grained word segmentation uses an open source word segmentation tool or model for word segmentation. The entity extraction model extracts entities by using a BERT+BILSTM+CRF model or a span-based pointer network model.

[0074] Furthermore, in this embodiment, a method combining business dictionaries, coarse-grained word segmentation and entity extraction models is used to extract entities from input sentences, and entities extracted by various methods are integrated to determine an entity set corresponding to the input sentence.

[0075] In step S220, recall is performed based on each entity in the entity set to determine a candidate entity set.

[0076] In this embodiment, accurate identification of entities in the entity set has an important impact on improving the accuracy of intent recognition corresponding to the input sentence, and can reduce or avoid errors in intent recognition results caused by missing entities, diversity of entity information, etc. Therefore, in this embodiment, a candidate entity set is determined by recall, and then the intent recognition result corresponding to the input sentence is determined by the candidate entity set to improve the overall accuracy and reliability of the intent recognition result. Among them, each candidate entity in the candidate entity set is an entity that is consistent with the corresponding entity in the input sentence or has a semantic relationship such as contextual connection. Therefore, by determining the candidate entity set through the above method, more entity-related information can be referenced during the intent recognition process, the accuracy of entity recognition can be improved, and the overall accuracy and reliability of the intent recognition result can be improved, especially for intent recognition in the case of a large number of entities or rich entity information in the input sentence, the optimization effect is more obvious.

[0077] Optionally, in this embodiment, entities in the entity set are recalled based on dictionary recall and vector recall methods respectively to determine a candidate entity set. In dictionary recall, candidate entities corresponding to the entity are determined by a pre-established recall dictionary. In vector recall, entity vectors similar to the vector corresponding to the entity are recalled from an online server through a vector recall model to determine candidate entities corresponding to the entity. Thus, after the candidate entities are determined by the above method, each candidate entity is added to the candidate entity set to determine the candidate entity set.

[0078] Furthermore, in this embodiment, when determining candidate entities through dictionary recall, a recall dictionary of key-value structure is first established offline, where the key is the normalized entity name and the value is all candidate entity IDs under the normalized entity name. Then, the entities in the entity set are normalized and recalled online based on the normalized entity name to determine the candidate entities.

[0079] Optionally, to ensure the real-time and effectiveness of the recall dictionary, the recall dictionary in this embodiment will be updated regularly or irregularly to make the knowledge in the recall dictionary more comprehensive and timely, so that the candidate entities determined based on the dictionary recall can be more realistic, thereby improving the accuracy and timeliness of the entity recall results, which is conducive to further improving the accuracy and reliability of the intent recognition results.

[0080] Alternatively, since an entity in an input sentence may be recalled to determine that there are multiple candidate entities, in this embodiment, the multiple candidate entities determined by the dictionary recall method will be screened using popularity as a screening condition, and the candidate entities whose popularity meets the preset conditions will be screened out as candidate entities in the candidate entity set. The preset condition can be that the popularity is greater than a predetermined threshold, or that after sorting by popularity, the popularity ranking is before a predetermined position.

[0081] Specifically, in this embodiment, the recalled entities are sorted according to popularity, and the topM entities with the highest popularity are determined as candidate entities for the slots corresponding to the entities in the input sentence. For example, assuming that the entity in the input sentence is work A, the recalled candidate entities and the popularity of each candidate entity are: the original author's work A1-popularity a1, the first copy author's work A2-popularity a2, and the second copy author's work A3-popularity a3, and the popularity corresponding to each candidate entity is sorted as a1>a2>a3. Therefore, when a candidate entity with the highest popularity is selected as a candidate entity in the candidate entity set, the original author's work A1 is used as a candidate entity for the entity work A. Therefore, by determining the candidate entity corresponding to the entity in the input sentence through the dictionary recall method, the accuracy of the recall can be guaranteed, thereby improving the accuracy of the subsequent intent recognition determined based on the candidate entity. At the same time, by screening the candidate entities, the overall data processing volume of the system can be reduced, which is conducive to improving the overall intent recognition efficiency.

[0082] It should be noted that in this embodiment, only popularity is used as an example of a screening condition, and other non-popularity screening conditions can also be selected to screen the candidate entities obtained by recall. At the same time, the predetermined threshold of popularity mentioned therein and the specific values ​​of the number of candidate entities M determined by dictionary recall can be determined by experience and set according to the actual usage scenario, and their setting method and specific values ​​are not limited here.

[0083] Furthermore, in this embodiment, when determining candidate entities through vector recall, the vector recall model is first trained offline, and then the candidate entities corresponding to the entities in the input sentence are recalled from the vector library of the online server according to the trained vector recall model. The vector library stores various entities and the associated information of each entity (label, summary, attribute relationship, etc.).

[0084] Optionally, in this embodiment, by updating the knowledge in the vector library in a regular or irregular manner, the knowledge in the vector library can be made more comprehensive and more timely, thereby making the candidate entities determined based on the vector library more practical, improving the accuracy and timeliness of the entity recall results, and helping to further improve the accuracy and reliability of the intent recognition results.

[0085] Optionally, the vector recall model in this embodiment adopts a dual-tower model of a bi-encode framework. The processing process of the vector recall model is as follows: Figure 5As shown. In this embodiment, one side of the vector recall model is input by the user, and the other side is the summary and attribute relationship data of the entity to be selected in the vector library. After the input data on both sides pass through the encoding layer (BERT), the pooling layer (pooling) and the fully connected layer (MLP), the similarity between the vector representation (query embedding) of the corresponding input sentence and the vector representation (entity embedding) of the entity to be selected is obtained. Among them, the vector dimension corresponding to the input sentence and the vector corresponding to the entity to be selected are consistent, and the default dimension is 256. Afterwards, each entity to be selected is screened according to the similarity between the corresponding vector of the input sentence and the corresponding vector of each entity to be selected, and the candidate entity is determined from each entity to be selected. Therefore, in this embodiment, the candidate entity corresponding to the entity in the input sentence is determined by the vector recall method, and it can be recalled based on the semantic similarity, which is conducive to improving the generalization rate and recall rate of the recall.

[0086] Furthermore, the vector recall model in this embodiment is determined by pre-training. In an offline state, the entity encode of the bi-encode dual-tower model is used to perform vector estimation on all candidate entities, and the corresponding vectors of the candidate entities are written into the online vector server, and the user-side encode model is deployed online. When performing vector recall, the vector recall model is called online, and the user-side encode model is used to predict the vector of the input sentence. After obtaining the vector of the input sentence, the vector similarity retrieval and recall of the entity is performed based on the vector retrieval tool (such as faiss) to determine the candidate entity. Optionally, when determining the candidate entity by the vector recall method in this embodiment, the candidate entity with a similarity greater than a predetermined threshold can be selected as the candidate entity, or the recalled candidate entity can be screened according to the similarity, and the topN candidate entities with high similarity are selected as candidate entities. Therefore, in this embodiment, after determining the candidate entity by the above-mentioned vector recall method, each selected entity is added to the candidate entity set to determine the candidate entity set.

[0087] It should be noted that the predetermined threshold of similarity mentioned in this embodiment and the specific value of the number N of candidate entities determined by vector recall can be determined empirically and set according to actual usage scenarios. The setting method and specific values ​​are not limited here.

[0088] In step S230, a reference filling result is determined according to each candidate entity in the candidate entity set.

[0089] In this embodiment, after the candidate entity set is determined, the candidate entities in the candidate entity set are screened and a reference filling result corresponding to the input sentence is determined.

[0090] Alternatively, if Figure 6As shown, in this embodiment, the following steps are performed to determine the reference filling result according to each candidate entity in the candidate entity set.

[0091] In step S310, corresponding entity description information is generated according to each candidate entity in the candidate entity set.

[0092] In this embodiment, the entity description information is determined based on the input sentence and a combination of one or more of the following: entity type, entity intent, entity summary, and entity attribute relationship. Among them, entity type refers to the category to which the entity belongs, and each entity may belong to one or more types, such as people, places, actors, singers, etc. Entity summary refers to a brief description or overview information of the entity, which may include a text description, key attributes or other relevant information, such as definition, background, characteristics, etc. Entity attribute information refers to data items that describe entity characteristics or attributes, and provides detailed information about the entity, such as name, publication date, etc.

[0093] Optionally, in this embodiment, after determining the entity information of the candidate entity, such as the entity type, entity intention, entity summary, and entity attribute relationship, the user's input sentence and the entity information of the candidate entity are spliced ​​in the form of triples to generate entity description information corresponding to the candidate entity. The input sentence and the entity information of the candidate entity can be distinguished by markers (such as CLS and SEP markers) so that the candidate entities can be screened according to the processing results of the entity description information, and the reference filling result corresponding to the input sentence can be determined.

[0094] In step S320, the matching score between the corresponding candidate entity and the input sentence is determined according to the description information of each entity.

[0095] In this embodiment, since the entity description information includes both input sentence information and entity information of the candidate entity, the degree of match (i.e., matching score) between the corresponding candidate entity and the input sentence can be determined by determining each entity description information, and then each candidate entity in the candidate entity set is screened according to the degree of match to determine the reference filling result corresponding to the input sentence.

[0096] Optionally, in this embodiment, the matching score between each candidate entity and the input sentence is determined by processing each entity description information using an entity ranking model. The entity ranking model uses a cross-encode ranking model to determine the matching score between each candidate entity and the input sentence by processing the entity description information, that is, the input data based on the input sentence and the entity description information that integrates multiple features of the entity.

[0097] Figure 7 Schematic diagram of an entity sorting model according to an embodiment of the present invention. Figure 7As shown, the entity ranking model in this embodiment includes an input layer, an encoding layer (BERT / Robert), a first fully connected layer (MLP-1), a fusion layer (concat), and a second fully connected layer (MLP-2). When determining the matching score between the candidate entity and the input sentence, the entity ranking model in this embodiment performs the following steps: Figure 8 The method shown in determines the matching score between each candidate entity and the input sentence, which specifically includes the following steps.

[0098] In step S410, an input vector corresponding to the entity description information is determined.

[0099] In this embodiment, when determining the matching score between a candidate entity and an input sentence, the entity description information of the candidate entity is first transmitted to the encoding layer through the input layer, and then the entity description information is encoded through the encoding layer and converted into data in a predetermined vector format, and output through the first fully connected layer. The output result is the input vector corresponding to the entity description information.

[0100] In step S420, the input vector is fused with feature information of the candidate entity corresponding to the entity description information to determine a matching score between the candidate entity and the input sentence.

[0101] In this embodiment, the feature information includes a combination of one or more of the following: the popularity of the candidate entity, the category of the candidate entity, and the vector similarity between the candidate entity and the corresponding entity (i.e., the entity in the input sentence). Optionally, in this embodiment, the features of the candidate entity are discretized to obtain the corresponding embedding vector (i.e., embedding vector), and the input vector and the embedding vector of the candidate entity feature information are fused through the fusion layer to match the input sentence and the candidate entity through fusion, thereby determining whether the candidate entity matches the semantics, context and other features of the input sentence. Finally, the matching score of the corresponding candidate entity and the input sentence is output through the second fully connected layer and the activation function.

[0102] In step S330, in response to the matching score satisfying a preset condition, the corresponding candidate entity is determined as a reference filling result.

[0103] In this embodiment, the reference filling result is an intent recognition result determined based on the entity linking method, and is used to replace the initial recognition result for answer indexing when the initial recognition result determined based on the traditional intent recognition method cannot well express the intent and word slot filling result in the user's input sentence, so as to improve the accuracy of intent recognition, and push matching content based on accurate recognition results to enhance the user's conversation experience.

[0104] Furthermore, in this embodiment, after determining the matching score between each candidate entity in the candidate entity set and the input sentence by the above method, it is determined whether to use the corresponding candidate entity as a reference filling result according to the matching score. Optionally, in this embodiment, the confidence of the corresponding candidate entity is characterized by the matching score between the candidate entity and the input sentence. When the confidence meets the preset conditions, such as the confidence is greater than the preset threshold (such as 0.9), or the confidence is ranked highest compared to other candidate entities, it indicates that the corresponding candidate entity is completely confident, and the candidate entity can be determined as the reference filling result of the corresponding input sentence. Therefore, in this embodiment, by determining the matching score between each candidate entity and the input sentence and comparing each matching score with the preset conditions to determine the reference filling result, the confidence of the reference filling result can be made higher, which is conducive to improving the accuracy of the intention recognition result of the input sentence.

[0105] In step S400, the initial recognition result is modified according to the reference filling result to determine the intention recognition result.

[0106] In this embodiment, after the initial recognition result and the reference filling result corresponding to the input sentence are determined respectively by two different methods, the final intent recognition result is determined according to the reference filling result and the initial recognition result. Optionally, since the reference filling result is determined based on the entity linking method, compared with the determination of the initial recognition result, the knowledge and information used in the process are more comprehensive and diverse, and the result accuracy is higher. Therefore, the initial recognition result can be corrected or determined by the reference filling result, so that the final intention recognition result is more accurate and more in line with the query requirements of the user input sentence, thereby improving the user experience.

[0107] Optionally, in this embodiment, whether to correct the initial recognition result is determined by judging whether the reference filling result and the initial recognition result are consistent. Furthermore, when the reference filling result is inconsistent with the initial filling result in the initial recognition result, for example, the entity name in the word slot is recognized as a category, two entities that are close to each other are recognized as one entity, the entity type is recognized incorrectly, etc., it indicates that the initial recognition result is inaccurate (including inaccurate initial filling result and / or inaccurate initial intent), and it is necessary to correct the initial filling result and / or initial intent in the initial recognition result through the reference filling result to determine the final intent recognition result or directly determine the final intent recognition result based on the reference filling result.

[0108] Since the word slot filling result and the intention corresponding to the input sentence have a strong correlation, the word slot filling result can reflect the intention of the corresponding input sentence. Therefore, in this embodiment, the word slot filling result (i.e., the initial filling result) in the initial recognition result is corrected first, and then the initial intention is corrected based on the correction of the initial filling result. In addition, the operation of correcting the initial recognition result can be to replace the initial filling result with a reference filling result, or to use the content in the reference filling result to replace the content in the initial recognition result that is inconsistent with the corresponding position of the reference filling result, so as to achieve the effect of correcting the initial filling result and the initial intention in the initial recognition result. Correspondingly, the method for correcting the initial recognition result in this embodiment includes: in response to the inconsistency between the reference filling result and the initial filling result, the initial filling result is corrected to the reference filling result, so as to determine the corrected initial recognition result as the intention recognition result; or, according to the reference filling result, the content in the initial recognition result that is inconsistent with the reference filling result is corrected, so as to determine the corrected initial recognition result as the intention recognition result.

[0109] In addition, when the reference filling result is consistent with the initial recognition result, it indicates that the initial recognition result is accurate, and the initial filling result or the reference filling result in the initial recognition result can be used as the word slot filling result corresponding to the input sentence, and then the initial intention in the initial recognition result and the initial filling result can be used as the final intention recognition result, or the reference filling result and the corresponding intention can be used as the final intention recognition result. Correspondingly, the method for correcting the initial result includes: in response to the reference filling result being consistent with the initial filling result, determining the final intention recognition result according to the initial recognition result or the reference filling result.

[0110] The technical solution of the embodiment of the present invention obtains an input sentence, performs intent recognition and word slot filling on the input sentence based on a preset neural network model, and determines an initial recognition result; determines a reference filling result corresponding to the input sentence based on an entity linking method; and corrects the initial recognition result according to the reference filling result to determine the intent recognition result. Thus, the intent and word slot filling result corresponding to the input sentence are determined through the processing of two different branches, and a more accurate reference filling result is determined by injecting rich entity knowledge when performing intent recognition based on the entity linking method. The initial recognition result is corrected by fusing the reference filling result with the initial recognition result, so that the corrected intent recognition result can be more consistent with the input sentence, solve the problem that the intent of the input sentence and the word slot filling result are prone to errors, and improve the accuracy of the intent recognition and word slot filling results corresponding to the user's input sentence, thereby improving the user experience.

[0111] Fig. 9 Schematic diagram of an intention recognition device according to an embodiment of the present invention. Fig. 9As shown, the intention recognition device in this embodiment includes an acquisition unit 1, a first recognition unit 2, a second recognition unit 3 and a correction unit 4. Among them, the acquisition unit 1 is used to obtain an input sentence. The first recognition unit 2 is used to perform intention recognition and word slot filling on the input sentence based on a preset neural network model, and determine an initial recognition result, and the initial recognition result includes an initial intention and an initial filling result. The second recognition unit 3 is used to determine a reference filling result based on an entity linking method. The correction unit 4 is used to correct the initial recognition result according to the reference filling result to determine the intention recognition result.

[0112] Optionally, when the first recognition unit 2 in this embodiment determines the initial recognition result based on a preset neural network model, it is specifically used to perform intent recognition and word slot filling on the input sentence according to a multi-task joint learning model or a neural network model including an intent recognition model and a word slot filling model to determine the initial recognition result.

[0113] Optionally, the second recognition unit 3 in this embodiment is also used to perform entity extraction on the input sentence to determine the corresponding entity set when determining the reference filling result; recall each entity in the entity set to determine the candidate entity set; and determine the reference filling result according to each candidate entity in the candidate entity set. When performing entity extraction, the second recognition unit 3 is specifically used to perform entity extraction on the input sentence based on at least one of a business dictionary, a coarse-grained word segmentation, and an entity extraction model to determine the corresponding entity set. When performing entity recall, the second recognition unit 3 is specifically used to recall the entities in the entity set based on the dictionary recall and vector recall methods, respectively, to determine the candidate entity set. When determining the reference filling result according to the candidate entity set, the second recognition unit 3 is specifically used to generate corresponding entity description information according to each candidate entity in the candidate entity set; determine the matching score between the corresponding candidate entity and the input sentence according to each entity description information; in response to the matching score satisfying the preset condition, the corresponding candidate entity is determined as the reference filling result.

[0114] Furthermore, the second recognition unit 3 in this embodiment is also used to determine the input vector corresponding to the entity description information; the input vector is merged with the feature information of the candidate entity corresponding to the entity description information to determine the matching score between the corresponding candidate entity and the input sentence. The entity description information is determined based on the input sentence and a combination of one or more of the following: entity type, entity intent, entity summary, and entity attribute relationship. The feature information includes a combination of one or more of the following: the popularity of the candidate entity, the category of the candidate entity, and the vector similarity between the candidate entity and the corresponding entity.

[0115] Optionally, when the correction unit 4 in this embodiment corrects the initial recognition result according to the reference filling result to determine the intention recognition result, it is specifically used to correct the initial filling result to the reference filling result in response to the inconsistency between the reference filling result and the initial filling result, so as to determine the corrected initial recognition result as the intention recognition result; or, correct the content in the initial recognition result that is inconsistent with the reference filling result according to the reference filling result, so as to determine the corrected initial recognition result as the intention recognition result.

[0116] Fig.10 Schematic diagram of an electronic device according to an embodiment of the present invention. Fig.10 The electronic device shown is a general address query device, which includes a general computer hardware structure, which at least includes a processor 101 and a memory 102. The processor 101 and the memory 102 are connected via a bus 103. The memory 102 is suitable for storing instructions or programs executable by the processor 101. The processor 101 can be an independent microprocessor or a collection of one or more microprocessors. Thus, the processor 101 executes the instructions stored in the memory 102, thereby executing the method flow of the embodiment of the present invention as described above to realize the processing of data and the control of other devices. The bus 103 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to the display controller 104 and the display device and the input / output (I / O) device 105. The input / output (I / O) device 105 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices known in the art. Typically, the input / output device 105 is connected to the system through an input / output (I / O) controller 106.

[0117] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices (equipment) or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may adopt a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] The present application is described with reference to flowcharts of methods, apparatuses (devices) and computer program products according to embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0119] These computer program instructions may be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device that implements the process Figure 1 A function specified in a process or multiple processes.

[0120] These computer program instructions may also be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce the instructions for implementing the process Figure 1 A device that specifies functions in a process or multiple processes.

[0121] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used for a computer to execute part or all of the above method embodiments.

[0122] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by specifying relevant hardware through a program, and the program is stored in a storage medium, including several instructions for a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0123] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for identifying intent, It is characterized in that The method comprises: Get input sentence; Performing intent recognition and word slot filling on the input sentence based on a preset neural network model to determine an initial recognition result, wherein the initial recognition result includes an initial intent and an initial filling result; Determine a reference completion result corresponding to the input sentence based on an entity linking method; The initial recognition result is modified according to the reference filling result to determine the intention recognition result.

2. The method according to claim 1, It is characterized in that Determining the reference completion result corresponding to the input sentence based on the entity linking method includes: Perform entity extraction on the input sentence to determine a corresponding entity set; Recalling each entity in the entity set to determine a candidate entity set; A reference filling result is determined according to each candidate entity in the candidate entity set.

3. The method according to claim 2, It is characterized in that The extracting entities from the input sentence to determine the corresponding entity set comprises: Entity extraction is performed on the input sentence based on at least one of a business dictionary, a coarse-grained word segmentation, and an entity extraction model to determine a corresponding entity set.

4. The method according to claim 2, It is characterized in that The recalling of each entity in the entity set to determine the candidate entity set includes: The entities in the entity set are recalled based on a dictionary recall method and a vector recall method respectively to determine a candidate entity set.

5. The method according to claim 2, It is characterized in that Determining a reference filling result according to each candidate entity in the candidate entity set includes: Generate corresponding entity description information according to each candidate entity in the candidate entity set; Determine a matching score between a corresponding candidate entity and the input sentence according to each entity description information; In response to the matching score satisfying a preset condition, the corresponding candidate entity is determined as a reference filling result.

6. The method according to claim 5, It is characterized in that The determining the matching score between the corresponding candidate entity and the input sentence according to each entity description information comprises: Determine an input vector corresponding to the entity description information; The input vector is fused with feature information of a candidate entity corresponding to the entity description information to determine a matching score between the corresponding candidate entity and the input sentence.

7. The method according to claim 5, It is characterized in that The entity description information is determined according to a combination of the input sentence and one or more of the following: entity type, entity intent, entity summary, and entity attribute relationship.

8. The method according to claim 6, It is characterized in that The feature information includes a combination of one or more of the following: popularity of the candidate entity, category of the candidate entity, and vector similarity between the candidate entity and the corresponding entity.

9. The method according to claim 1, It is characterized in that The modifying the initial recognition result according to the reference filling result to determine the intention recognition result includes: In response to the reference filling result being inconsistent with the initial filling result, the initial filling result is corrected to the reference filling result, so as to determine the corrected initial recognition result as the intended recognition result; or, The content in the initial recognition result that is inconsistent with the reference filling result is corrected according to the reference filling result, so as to determine the corrected initial recognition result as the intended recognition result.

10. The method according to claim 1, It is characterized in that The preset neural network model is a multi-task joint learning model or the preset neural network model includes an intent recognition model and a word slot filling model.

11. An intention recognition device, It is characterized in that The device comprises: An acquisition unit, used for acquiring an input sentence; A first recognition unit, configured to perform intent recognition and word slot filling on the input sentence based on a preset neural network model, and determine an initial recognition result, wherein the initial recognition result includes an initial intent and an initial filling result; A second recognition unit, used to determine a reference filling result based on an entity linking method; A correction unit is used to correct the initial recognition result according to the reference filling result to determine the intended recognition result.

12. A computer program product, It is characterized in that The computer program product comprises a computer program / instructions, and when the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

13. An electronic device comprising a memory and a processor, It is characterized in that The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 10 are implemented.