A method, apparatus and electronic device for rapid recognition of large-scale intents

Through the cluster index-deep matching recall-sorting method, combined with the BERT pre-trained model for semantic representation, the problem of extremely large-scale intention category recognition is solved, fast and accurate intention recognition is achieved, and labor costs are reduced.

CN112035626BActive Publication Date: 2025-05-27BEIHAI QIANG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010642447.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-06
Publication Date
2025-05-27
Estimated Expiration
2040-07-06

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify extremely large-scale intention categories, especially in open dialogue scenarios. Traditional methods such as softmax-based multi-classifiers and one-vs-other binary classifiers fail when the number of intentions reaches tens of thousands or even hundreds of thousands, and ignore the correlation information between intentions.

Method used

The cluster index-deep matching recall-sorting scheme is adopted. By semantic vector transformation and clustering of the intent category information input by historical user dialogue, the index is established for rapid retrieval, combined with the BERT pre-trained model for semantic representation, and the intent recognition result is determined through the sorting model.

Benefits of technology

It realizes rapid identification of extremely large-scale intention categories, reduces labor costs, improves the accuracy and efficiency of intention recognition, and can effectively capture the related information between intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112035626B_ABST
    Figure CN112035626B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus and electronic device for rapid recognition of large-scale intents. The method includes: performing semantic vector conversion on the intent category information of the dialogue inputs of historical users; performing semantic clustering on each intent after semantic vector conversion; establishing an index for the results after semantic clustering, where the index is used to search for the intent category corresponding to the dialogue input in a preset intent database; obtaining the dialogue input of the current user in real time and using the index for search and matching; inputting the semantic vector of the dialogue input and the search and matching results into a sorting model for sorting to determine the intent recognition result. The method of the present invention adopts a clustering index-depth matching recall-sorting scheme, which greatly saves manpower, realizes extremely large-scale intent recognition at the ten-thousand level, improves the intent recognition efficiency, and also improves the accuracy of user intent recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer information processing, and particularly to a method, device and electronic device for quickly identifying large-scale intents. Background Art

[0002] With the development of Internet technology, dialogue systems have been widely used in e-commerce, intelligent devices and other aspects, and have attracted more and more attention. Common dialogue systems include Siri, Echo, Bixby, Microsoft Xiaoice, Alibaba Xiaomi, intelligent speakers, etc. Intent recognition is the primary and important task in dialogue systems. Especially in open dialogue scenarios, limited by the capabilities of classifiers, conversations are classified into dozens or hundreds of coarse-grained intents. Such coarse granularity makes chatbots unable to accurately capture users' intents, thus affecting the effect of human-computer interaction.

[0003] Most of the existing intent recognizers limit the types of intents to dozens to hundreds. A multi-classifier based on softmax, or a combination of multiple binary classifiers based on one-vs-other can solve such intent recognition. However, when the number of intent categories reaches tens of thousands or even hundreds of thousands, the multi-classifier based on softmax will fail because the probability assigned to each intent after softmax processing becomes very small and has no distinguishability. While the method of combining multiple binary classifiers based on one-vs-other can be extended as the number of intents increases, it includes tens of thousands of binary classifiers in its construction, and the labor cost is basically unacceptable. In addition, both of the above two methods ignore the correlation information between intents.

[0004] Therefore, it is necessary to provide a fast recognition method applicable to large-scale intent recognition. Summary of the Invention

[0005] To solve the above problems, the present invention provides a method for quickly identifying large-scale intents for human-computer interaction, including: performing semantic vector conversion on the intent category information of the historical user's dialogue input; performing semantic clustering on each intent after semantic vector conversion; establishing an index for the result of semantic clustering, where the index is used to search for the intent category corresponding to the dialogue input in a preset intent database; obtaining the current user's dialogue input in real time and using the index for search and matching; inputting the semantic vector of the dialogue input and the search and matching result into a sorting model for sorting to determine the intent recognition result.

[0006] Preferably, it further includes: the index includes establishing the correspondence between the user ID and the clustered intention categories; calculating the semantic vector similarity, and based on the similarity judgment, clustering multiple intention categories together and representing them with one search ID to form multiple intention category sets, where each intention category includes key text features.

[0007] Preferably, the real-time acquisition of the current user's conversation input and the use of the index for search and matching include: when receiving the conversation text input of the current user, inputting the conversation text and the key text features in the category samples into the vector conversion model to obtain the semantic vector of the conversation text, and combining it with the semantic vectors of the preset intention categories, and inputting them into the deep matching network model together to output the matching value between the conversation text and the intention category.

[0008] Preferably, it further includes: comparing the output matching value with a preset threshold; in the case where the matching value is higher than the preset threshold, determining that the conversation text of the current user is relevant to the current intention category, and determining that the conversation text is relevant to other intention categories in the intention category set where the current intention category is located.

[0009] Preferably, it further includes: through search and matching, recalling the intention category set related to the conversation text input of the current user.

[0010] Preferably, it further includes: inputting the semantic vector of the conversation text of the current user and the semantic vectors of the recalled intention categories into the ranking model to output ranking scores; selecting the intention category with the highest ranking score as the intention recognition result of the conversation text.

[0011] Preferably, it further includes: based on the semantic vector of the conversation text of the current user, the number of times of executing the output of the ranking score using the ranking model is equal to the number of the recalled intention categories, and when ranking scores are obtained for each recalled intention category, score ranking is performed.

[0012] Preferably, the vector conversion model includes the BERT model and the RoBERTa model.

[0013] In addition, the present invention also provides a device for quickly identifying large-scale intents for human-computer interaction, including: a conversion module for performing semantic vector conversion on the intent category information of the historical user's conversation input; a clustering module for semantically clustering each intent after semantic vector conversion; a building module for building an index for the results after semantic clustering, and the index is used to search for the intent category corresponding to the conversation input in a preset intent database; a search and matching module for obtaining the current user's conversation input in real time and using the index for search and matching; a determination module for inputting the semantic vector of the conversation input and the search and matching result into a sorting model for sorting to determine the intent recognition result.

[0014] Preferably, it further includes: the index includes establishing the correspondence between the user ID and the clustered intent categories; calculating the semantic vector similarity, and based on the similarity judgment, clustering multiple intent categories together and representing them with one search ID to form multiple intent category sets, where each intent category includes key text features.

[0015] Preferably, it further includes a processing module, and the processing module is used to, when receiving the conversation text input of the current user, input the conversation text and the key text features in the category samples into a vector conversion model to obtain the semantic vector of the conversation text, and combine it with the semantic vectors of the preset intent categories and input them into a deep matching network model together to output the matching value between the conversation text and the intent category.

[0016] Preferably, it further includes a comparison module, and the comparison module is used to compare the output matching value with a preset threshold; in the case where the matching value is higher than the preset threshold, it is determined that the conversation text of the current user is related to the current intent category, and it is determined that the conversation text is related to other intent categories in the intent category set where the current intent category is located.

[0017] Preferably, it further includes: through search and matching, an intent category set related to the conversation text input of the current user is recalled.

[0018] Preferably, it further includes: inputting the semantic vector of the conversation text of the current user and the semantic vectors of the recalled intent categories into the sorting model to output sorting scores; selecting the intent category with the highest sorting score as the intent recognition result of the conversation text.

[0019] Preferably, it further includes: based on the semantic vector of the conversation text of the current user, the number of executions of using the sorting model to output sorting scores is equal to the number of the recalled intent categories, and when sorting scores are obtained for each recalled intent category, score sorting is performed.

[0020] Preferably, the vector conversion model includes a BERT model and a RoBERTa model.

[0021] In addition, the present invention also provides an electronic device, which includes: a processor; and a memory storing computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the method for quickly identifying large-scale intents of the present invention.

[0022] In addition, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and the one or more programs, when executed by a processor, implement the method for quickly identifying large-scale intents of the present invention.

[0023] Beneficial effects

[0024] Compared with the prior art, the present invention adopts a clustering index-depth matching recall-ranking scheme, which solves the problem that traditional softmax-based methods fail in the scenario of extremely large-scale intent category classification. At the same time, it is not necessary to design a large number of binary classifiers, which greatly saves manpower and realizes extremely large-scale intent recognition at the ten-thousand level; based on the BERT pre-training model for semantic representation, compared with traditional bag-of-words features or tf-idf features, it realizes the transfer and fusion of information and can well capture the inherent semantic information; through semantic-level intent category clustering, not only the association between intents is better discovered, but also the number of intent categories is greatly reduced after clustering. Through the index, fast retrieval is realized, and the intent recognition efficiency is further improved; by calculating all recalled intent categories through a ranking model and selecting the one with the highest score as the intent category of the current dialogue text input, the accuracy of user intent recognition is further improved. Brief description of the drawings

[0025] In order to make the technical problems solved by the present invention, the technical means adopted, and the technical effects obtained more clear, the specific embodiments of the present invention will be described in detail below with reference to the drawings. However, it should be noted that the drawings described below are only the drawings of the exemplary embodiments of the present invention, and those skilled in the art can obtain the drawings of other embodiments without creative efforts based on these drawings.

[0026] Figure 1 is a flowchart of an example of the method for quickly identifying large-scale intents of the present invention.

[0027] Figure 2 is a flowchart of another example of the method for quickly identifying large-scale intents of the present invention.

[0028] Figure 3 is a flowchart of yet another example of the method for quickly identifying large-scale intents of the present invention.

[0029] Figure 4 It is a schematic structural block diagram of an example of a device for quickly identifying large-scale intentions of the present invention.

[0030] Figure 5 It is a schematic structural block diagram of another example of a device for quickly identifying large-scale intentions of the present invention.

[0031] Figure 6 It is a schematic structural block diagram of yet another example of a device for quickly identifying large-scale intentions of the present invention.

[0032] Figure 7 It is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention.

[0033] Figure 8 It is a structural block diagram of an exemplary embodiment of a computer-readable medium according to the present invention. Detailed implementation manners

[0034] Now, the exemplary embodiments of the present invention will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments enables the present invention to be more comprehensive and complete, and more convenient to fully convey the inventive concept to those skilled in the art. Identical reference numerals in the figures denote the same or similar elements, components, or parts, and thus their repeated description will be omitted.

[0035] On the premise of conforming to the technical concept of the present invention, the features, structures, characteristics, or other details described in a specific embodiment do not exclude being combined in a suitable manner in one or more other embodiments.

[0036] In the description of specific embodiments, the features, structures, characteristics, or other details described in the present invention are for enabling those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics, or other details.

[0037] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may be changed according to the actual situation.

[0038] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0039] It should be understood that although the first, second, third, etc. attributives indicating numbers may be used in this text to describe various devices, components, components or parts, this should not be limited by these attributives. These attributives are used to distinguish one from another. For example, the first device can also be called the second device without departing from the essential technical solution of the present invention.

[0040] The term "and / or" or "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0041] In order to more accurately identify the user's intention and solve the problem of extremely large-scale intention recognition, the present invention proposes a method for quickly recognizing extremely large-scale intentions, using a fine-tuned BERT pre-trained model to implement semantic vector representations of conversations and intention categories, and performing intention category clustering, establishing an index for the clustering results, recalling multiple relevant intention categories based on this index, and selecting the most relevant intention category through a sorting model. The method of the present invention reduces the number of intention categories greatly while considering the relevance between intentions through semantic-level intention category clustering, realizes fast retrieval through the index, improves the intention recognition efficiency, and also improves the accuracy of user intention recognition.

[0042] In order to make the purpose, technical solution and advantages of the present invention clearer and more understandable, the following will further describe the present invention in detail with reference to specific embodiments and the accompanying drawings.

[0043] Embodiment 1

[0044] Next, with reference to Figures 1 to 3 Describe an embodiment of the method for quickly recognizing large-scale intentions of the present invention.

[0045] Figure 1 is a flowchart of an example of the method for quickly recognizing large-scale intentions of the present invention.

[0046] As Figure 1 shown, a method for quickly recognizing large-scale intentions, the method includes the following steps.

[0047] Step S101, perform semantic vector conversion on the intention category information of the historical user's conversation input.

[0048] Step S102, perform semantic clustering on each intention after semantic vector conversion;

[0049] Step S103: Index the results after semantic clustering. The index is used to search for the intent category corresponding to the conversation input in a preset intent database.

[0050] Step S104: Real-time obtain the conversation input of the current user, and use the index for search and matching.

[0051] Step S105: Input the semantic vector of the conversation input and the search and matching results into a sorting model for sorting to determine the intent recognition result.

[0052] First, in step S101, perform semantic vector conversion on the intent category information of the conversation input of the historical user.

[0053] In this example, obtain the user information of the historical user, the conversation input information in the historical conversation information, the historical business information corresponding to the historical conversation, etc.

[0054] Specifically, the conversation input information includes intent category information. In this example, "A" is the name of a star. For example, user 1 inputs "I like A", forming an intent Figure 1 , and for another example, user 2 inputs "What's the birthday of artist A?", forming an intent Figure 2 .

[0055] Furthermore, for example, use the BERT pre-trained model to perform semantic representation on the intent category information of the user input to perform semantic vector conversion.

[0056] Preferably, based on the relevant data of the historical business task, fine-tune the BERT pre-trained model, and use the fine-tuned BERT pre-trained model to perform semantic vector conversion on the intent category information.

[0057] Specifically, use the category description information of the intent category and the key text features in the category samples as the input of the fine-tuned BERT model to obtain the semantic vector representation of the intent category. Thus, the transfer and fusion of information are realized, and the internal semantic information can be well captured.

[0058] It should be noted that for semantic vector conversion, in other examples, the RoBERTa model, DistilBERT model, XLNet model, etc. can also be used. The above is only an example for illustration and should not be construed as a limitation to the present invention.

[0059] Further, for example, user information includes age, gender, region, city of residence, constellation, personality, education level, family structure, marital status, hobbies, income, occupation information, etc. For example, the historical business information includes the type of financial business applied for, the usage of financial business, etc. For example, the historical conversation information includes conversation duration, conversation time period, conversation connection rate, conversation tone and voice, reply content, etc.

[0060] Next, in step S102, semantic clustering is performed on each intention after semantic vector conversion.

[0061] In this example, for example, calculate the semantic vector similarity between intention Figure 1 and intention Figure 2 etc. Based on the judgment of similarity, multiple intention categories are grouped together and represented by a single search ID to form multiple intention category sets, where each intention category includes key text features.

[0062] Further, based on existing clustering methods, cluster the intention categories after semantic vector conversion for a very large scale. After clustering, the scale of the intention categories will be greatly reduced. Through semantic-level intention category clustering, not only can the associations between intentions be better discovered, but also the number of intention categories is greatly reduced after clustering.

[0063] It should be noted that existing clustering methods include the K-means clustering method, the HAC hierarchical agglomerative clustering method, etc. The above are only examples for illustration and should not be construed as limitations to the present invention.

[0064] Next, in step S103, an index is established for the result of semantic clustering. The index is used to search for the intention category corresponding to the conversation input in a preset intention database.

[0065] As Figure 2 shown, it also includes step S201 of establishing a preset intention database.

[0066] In step S201, a preset intention database is established to identify the most accurate intention corresponding to the user's conversation text input.

[0067] Specifically, the preset intention database includes preset intention categories, multiple intentions corresponding to each intention category, and semantic vectors corresponding to each intention.

[0068] In this example, for the result of semantic clustering, an index is established, where the index includes establishing the correspondence between user IDs and the clustered intention categories.

[0069] Specifically, for example, the intention Figure 1 identified for user 1 and the intention Figure 2, after vector similarity calculation, when the vector similarity is greater than the set threshold, the Figure 1 and Figure 2 are clustered together and represented by an index ID, that is, clustered into intent category 1. And so on, it also includes intent categories 2, 3... and so on. Further clustering is performed on the intent categories to form an intent category combination, and each intent category set corresponds to an index ID. Thus, through the index, it is possible to more quickly search for the intent category corresponding to the current user's conversation text input in the preset intent database. Thus, through the said index, fast retrieval is realized, and the intent recognition efficiency is further improved.

[0070] Furthermore, establish the corresponding relationship between each user ID and the clustered intent, where the user ID is, for example, the user account, the user's mobile phone number, etc.

[0071] Next, in step S104, the current user's conversation input is obtained in real time, and the said index is used for search and matching.

[0072] In this example, when the current user's conversation text input is received, a search is performed. Further, the semantic vector of the current user's conversation text input and the searched matching result are input into the deep matching network model together to output the matching value between the conversation text and the intent category, where the matching network model is preferably implemented using an existing multi-layer perceptron.

[0073] As Figure 3 described, it also includes step S301 of setting a preset threshold.

[0074] In step S301, a preset threshold is set to be used for comparing the output matching value with the preset threshold. When the matching value is higher than the preset threshold, it is determined that the current user's conversation text is related to the current intent category, and it is determined that the conversation text is related to other intent categories in the intent category set where the current intent category is located.

[0075] Thus, through the said index, the intent category set corresponding to the current user's conversation text input can be quickly retrieved. In other words, through search and matching, all intent category sets related to the current user's conversation text input are recalled. Due to the semantic clustering process in step S102, the retrieval result can be obtained more quickly.

[0076] Next, in step S105, the semantic vector of the conversation input and the search and matching result are input into the sorting model for sorting to determine the intent recognition result.

[0077] In this example, the semantic vector of the current user's conversation text and the semantic vectors of the recalled intent categories are input into the sorting model together to output sorting scores.

[0078] Specifically, based on the semantic vector of the current user's conversation text, the number of times of using the sorting model to output sorting scores is equal to the number of the recalled intent categories, and score sorting is performed until each recalled intent category obtains a sorting score.

[0079] Further, based on the result of the score sorting, the intent category with the highest sorting score (i.e., the intent category most relevant to the current conversation input) is selected as the intent recognition result of the conversation text.

[0080] It should be noted that the sorting model is composed of a binary classifier, preferably implemented by adding a sigmoid layer on top of BERT. This classifier will output a decision score (in this example, the sorting score). The process of outputting the sorting score is repeatedly executed to traverse each recalled intent category one by one. Eventually, each recalled intent category obtains a sorting score, and the one with the highest score is selected as the intent category of the current conversation text input. Thereby, the accuracy of user intent recognition is further improved.

[0081] It should be noted that the above is only an example for illustration and should not be construed as a limitation to the present invention.

[0082] The process of the above method is only used for the illustration of the present invention. Among them, the order and quantity of the steps have no special limitations. In addition, the steps in the above method can also be split into two or three steps, or some steps can also be combined into one step, which is adjusted according to the actual example.

[0083] Compared with the prior art, the present invention adopts a clustering index-depth matching recall-sorting scheme, which solves the problem that the traditional softmax-based method fails in the extremely large-scale intent category classification scenario. At the same time, it does not require designing a large number of binary classifiers, greatly saving manpower and realizing extremely large-scale intent recognition at the ten-thousand level; based on the BERT pre-trained model for semantic representation, compared with the traditional bag-of-words features or tf-idf features, it realizes the transfer and fusion of information and can well capture the inherent semantic information; through semantic-level intent category clustering, it not only better discovers the associations between intents, but also greatly reduces the number of intent categories after clustering. Through the index, fast retrieval is realized, and the intent recognition efficiency is further improved; by calculating all the recalled intent categories through the sorting model and selecting the one with the highest score as the intent category of the current conversation text input, the accuracy of user intent recognition is further improved.

[0084] Those skilled in the art can understand that all or part of the steps of implementing the above embodiments are realized as a program (computer program) executed by a computer data processing device. When the computer program is executed, the above method provided by the present invention can be realized. Moreover, the computer program can be stored in a computer-readable storage medium, which can be a readable storage medium such as a disk, an optical disc, a ROM, a RAM, etc., or a storage array composed of multiple storage media, such as a disk or tape storage array. The storage medium is not limited to centralized storage, and it can also be distributed storage, such as cloud storage based on cloud computing.

[0085] Embodiments of the device of the present invention will be described below. The device can be used to execute the method embodiments of the present invention. For the details described in the device embodiments of the present invention, they should be regarded as a supplement to the above method embodiments; for the details not disclosed in the device embodiments of the present invention, they can be implemented with reference to the above method embodiments.

[0086] Embodiment 2

[0087] Referring to Figure 4 、 Figure 5 and Figure 6 The present invention also provides a fast recognition device 400 for large-scale intents, which is used for human-computer interaction and includes: a conversion module 401 for performing semantic vector conversion on the intent category information of the historical user's dialogue input; a clustering module 402 for performing semantic clustering on each intent after semantic vector conversion; a building module 403 for building an index for the result after semantic clustering, and the index is used to search for the intent category corresponding to the dialogue input in a preset intent database; a search and matching module 404 for obtaining the current user's dialogue input in real time and using the index for search and matching; a determination module 405 for inputting the semantic vector of the dialogue input and the search and matching result into a sorting model for sorting to determine the intent recognition result.

[0088] Preferably, it further includes: the index includes establishing the correspondence between the user ID and the clustered intent categories; calculating the semantic vector similarity, and based on the similarity judgment, clustering multiple intent categories together and representing them with one search ID to form multiple intent category sets, where each intent category includes key text features.

[0089] As Figure 5 shown, it further includes a processing module 501. The processing module 501 is used to input the dialogue text and the key text features in the category sample into a vector conversion model when receiving the dialogue text input of the current user, so as to obtain the semantic vector of the dialogue text, and input it together with the semantic vector of the preset intent category into a deep matching network model to output the matching value between the dialogue text and the intent category.

[0090] As shown Figure 6 in the figure, it further includes a comparison module 601, which is configured to compare the output matching value with a preset threshold; in the case where the matching value is higher than the preset threshold, it is determined that the dialogue text of the current user is relevant to the current intention category, and it is determined that the dialogue text is relevant to other intention categories in the intention category set where the current intention category is located.

[0091] Preferably, it further includes: recalling an intention category set related to the dialogue text input of the current user through search matching.

[0092] Preferably, it further includes: inputting the semantic vector of the dialogue text of the current user and the semantic vectors of the recalled intention categories into the sorting model to output sorting scores; selecting the intention category with the highest sorting score as the intention recognition result of the dialogue text.

[0093] Preferably, it further includes: based on the semantic vector of the dialogue text of the current user, the number of times of using the sorting model to output sorting scores is equal to the number of the recalled intention categories, and score sorting is performed until each recalled intention category obtains a sorting score.

[0094] Preferably, the vector conversion model includes a BERT model and a RoBERTa model.

[0095] It should be noted that in Embodiment 2, the description of the same parts as in Embodiment 1 is omitted.

[0096] Compared with the prior art, the present invention adopts a clustering index-depth matching recall-sorting scheme, which solves the problem that the traditional softmax-based method fails in the extremely large-scale intention category classification scenario. At the same time, it does not require designing a large number of binary classifiers, greatly saving manpower and realizing extremely large-scale intention recognition at the ten-thousand level; based on the BERT pre-training model for semantic representation, compared with traditional bag-of-words features or tf-idf features, it realizes the transfer and fusion of information and can well capture the inherent semantic information; through semantic-level intention category clustering, not only better discovers the association between intentions, but also greatly reduces the number of intention categories after clustering. Through the index, fast retrieval is realized, and the intention recognition efficiency is further improved; by calculating all the recalled intention categories through the sorting model and selecting the one with the highest score as the intention category of the current dialogue text input, the accuracy of user intention recognition is further improved.

[0097] Those skilled in the art can understand that the modules in the above device embodiments can be distributed in the device according to the description, or can be correspondingly changed and distributed in one or more devices different from the above embodiments. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0098] Embodiment 3

[0099] The following describes an embodiment of the electronic device of the present invention, and this electronic device can be regarded as a specific physical implementation manner of the above method and device embodiments of the present invention. For the details described in the embodiment of the electronic device of the present invention, it should be regarded as a supplement to the above method or device embodiments; for the details not disclosed in the embodiment of the electronic device of the present invention, reference can be made to the above method or device embodiments to implement.

[0100] Figure 7 is a structural block diagram of an exemplary embodiment of an electronic device according to the present invention. The following will refer to Figure 7 to describe the electronic device 200 according to the present invention. Figure 7 The electronic device 200 shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0101] As Figure 7 shown, the electronic device 200 is presented in the form of a general computing device. The components of the electronic device 200 may include but are not limited to: at least one processing unit 210, at least one storage unit 220, a bus 230 connecting different system components (including the storage unit 220 and the processing unit 210), a display unit 240, etc.

[0102] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 210, so that the processing unit 210 executes the steps according to various exemplary embodiments of the present invention described in the above electronic device processing method part of this specification. For example, the processing unit 210 can execute the steps as Figure 2 shown.

[0103] The storage unit 220 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 2201 and / or a cache storage unit 2202, and may further include a read-only storage unit (ROM) 2203.

[0104] The storage unit 220 may further include a program / utilities 2204 having a set (at least one) of program modules 2205. Such program modules 2205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0105] The bus 230 can represent one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor unit, or a local bus using any of the various bus architectures.

[0106] The electronic device 200 can also communicate with one or more external devices 300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 200, and / or communicate with any device that enables the electronic device 200 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 250. Moreover, the electronic device 200 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through the network adapter 260. The network adapter 260 can communicate with other modules of the electronic device 200 through the bus 230. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0107] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described in the present invention can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium can implement the above method of the present invention.

[0108] As Figure 8As shown, the computer program can be stored on one or more computer-readable media. The computer-readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0109] The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0110] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0111] In summary, the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that general-purpose data processing devices such as microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.

[0112] The specific embodiments described above have further elaborated on the object, technical solution, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A rapid recognition method for large-scale intents, which is used for human-computer interaction, Characterized in that, Comprising: Inputting the intent category description information of the historical user's dialogue input and the key text features in the category samples into the fine-tuned BERT pre-trained model for semantic vector conversion; Clustering multiple intents after semantic vector conversion together to obtain intent categories, and then clustering the intent categories to form an intent category set; Each intent category set corresponds to an index, and the index is used to search for the intent category corresponding to the dialogue input in the preset intent database; the index includes establishing the corresponding relationship between the user ID and the clustered intent categories; wherein, each intent category includes calculating the semantic vector similarity of the key text features, and based on the similarity judgment, clustering multiple intent categories together and representing them with one search ID to form multiple intent category sets, wherein each intent category includes key text features; When receiving the dialogue text input of the current user, using the index for search and matching, inputting the semantic vector of the current user's dialogue text input and the searched matching result into the deep matching network model together to output the matching value between the dialogue text and the intent category, and comparing the output matching value with a preset threshold; in the case where the matching value is higher than the preset threshold, determining that the dialogue text of the current user is related to the current intent category, and determining that the dialogue text is related to other intent categories in the intent category set where the current intent category is located; recalling all intent category sets related to the dialogue text input of the current user; Inputting the semantic vector of the dialogue text input and the semantic vectors of the recalled intent categories into a sorting model for sorting to determine the intent recognition result.

2. The rapid recognition method for large-scale intents according to claim 1, Characterized in that, Further comprising: Inputting the semantic vector of the dialogue text of the current user and the semantic vectors of the recalled intent categories into the sorting model to output sorting scores; Selecting the intent category with the highest sorting score as the intent recognition result of the dialogue text.

3. The rapid recognition method for large-scale intents according to claim 2, Characterized in that, Further comprising: Based on the semantic vector of the dialogue text of the current user, the number of executions of using the sorting model to output sorting scores is equal to the number of recalled intent categories, and when sorting scores are obtained for each recalled intent category, score sorting is performed.

4. A rapid recognition device for large-scale intents, which is used for human-computer interaction, Characterized in that, Comprising: A conversion module, configured to input the intent category description information of the historical user's dialogue input and the key text features in the category samples into the fine-tuned BERT pre-trained model for semantic vector conversion; A clustering module, configured to cluster multiple intents after semantic vector conversion together to obtain intent categories, and then cluster the intent categories to form an intent category set; A building module, which is used to establish an index for each set of intent categories. The index is used to search for the intent category corresponding to the conversation input in a preset intent database. The index includes establishing the correspondence between the user ID and the clustered intent categories. Calculate the semantic vector similarity. Based on the similarity judgment, multiple intent categories are clustered together and represented by a single search ID to form multiple sets of intent categories. Each intent category includes key text features. A search and matching module, which is used to, when receiving the conversation text input of the current user, perform a search and matching using the index, input the semantic vector of the conversation text input of the current user and the searched matching result into a deep matching network model together, so as to output the matching value between the conversation text and the intent category, and compare the output matching value with a preset threshold. When the matching value is higher than the preset threshold, it is determined that the conversation text of the current user is related to the current intent category, and it is determined that the conversation text is related to other intent categories in the set of intent categories where the current intent category is located. Recall all sets of intent categories related to the conversation text input of the current user. A determination module, which is used to input the semantic vector of the conversation text input and the semantic vectors of the recalled intent categories into a sorting model for sorting to determine the intent recognition result.

5. The rapid intent recognition device for large-scale intents according to claim 4, wherein, it further includes: Input the semantic vector of the conversation text of the current user and the semantic vectors of the recalled intent categories into the sorting model to output sorting scores. Select the intent category with the highest sorting score as the intent recognition result of the conversation text.

6. The rapid intent recognition device for large-scale intents according to claim 5, wherein, it further includes: Based on the semantic vector of the conversation text of the current user, the number of times of executing to output sorting scores using the sorting model is equal to the number of the recalled intent categories, and score sorting is performed until sorting scores are obtained for each recalled intent category.

7. An electronic device, wherein, the electronic device includes: a processor; and, a memory storing computer-executable instructions, and the executable instructions, when executed, cause the processor to execute the rapid intent recognition method for large-scale intents according to any one of claims 1 to 3.

8. A computer-readable storage medium, wherein, the computer-readable storage medium stores one or more programs, and the one or more programs, when executed by a processor, implement the rapid intent recognition method for large-scale intents according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Semantic key index creating method and system

    CN107944027A

  • Method and device for establishing hierarchical intention system

    CN110674287A