Text intention recognition method, device, system, and text classification system

By establishing a database that stores historical text and its intent information, the existing text intent recognition methods have limited coverage and low recognition accuracy have been solved, and fast and accurate intent recognition and timely system repairs have been achieved, reducing the repair cost and risk.

CN112115229BActive Publication Date: 2025-05-27BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910538487.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-20
Publication Date
2025-05-27
Estimated Expiration
2039-06-20

AI Technical Summary

Technical Problem

The existing text intent recognition methods cover a limited range of text, low recognition accuracy, and redevelop code or retrain the model when intent recognition results are incorrect, resulting in a time-consuming and cost-effective repair process.

Method used

By establishing a database that stores historical text and its intent information, the system will be repaired in a timely manner when intent identification errors are identified, and the intent information of the text to be identified is determined using similar texts in the database.

Benefits of technology

It realizes fast and accurate text intention recognition, reduces the time and cost of the repair process, avoids the risk of system launch, and improves the performance of the intention classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112115229B_ABST
    Figure CN112115229B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, system for text intention recognition and a text classification system, relating to the field of computer technology. A specific embodiment of the method includes: obtaining one or more word segments in the text to be recognized; when it is determined that there is a historical text stored in a pre-established database that contains at least one of the word segments and the similarity with the text to be recognized meets a preset condition, determining the intention information of the text to be recognized according to the intention information of the historical text stored in the database. This embodiment can achieve timely repair of the system when the intention recognition is incorrect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a method, device, system for text intention recognition, and a text classification system. Background Art

[0002] Intention recognition is an important working link in dialogue systems such as chatbots. There are mainly three existing text intention recognition methods. The first is the method based on text templates, which relies on manual methods to summarize different intention patterns, and then organizes them into regular expression templates to match with the text to be recognized. The second is the method based on knowledge engineering, which uses human experience to define inference rules for each intention, and when the text to be recognized meets a certain rule, it is determined to have the corresponding intention. The third is the method based on statistical learning, which trains an intention classification model through labeled data, and uses the trained model to predict the intention of the text to be recognized. Common algorithms include decision trees, deep neural networks, etc.

[0003] The text ranges covered by the first two methods above are limited, and the recognition accuracy of the third method is relatively low. At the same time, for an intention recognition system implemented by any one method or a combination of multiple methods, when an incorrect intention recognition result occurs, it often requires re-developing code or re-training the model for emergency repair, and then re-releasing the version and conducting version online and online verification. Since system online has relatively high risks and costs and requires development, testing, and multiple levels of approval, the above repair process consumes a lot of time and high labor costs, and also needs to bear certain risks. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method, device, system for text intention recognition, and a text classification system, which can achieve timely repair of the system when intention recognition is incorrect by establishing a database storing historical texts and their intention information.

[0005] To achieve the above object, according to one aspect of the present invention, a method for text intention recognition is provided.

[0006] The method for text intention recognition according to the embodiments of the present invention includes: obtaining one or more word segments in the text to be recognized; when it is determined that there is a historical text stored in a pre-established database that contains at least one of the word segments and whose similarity with the text to be recognized meets a preset condition, determining the intention information of the text to be recognized according to the intention information of the historical text stored in the database.

[0007] Optionally, the historical texts stored in the database include: texts with incorrect intention recognition results in historical periods and whose intention information is manually marked.

[0008] Optionally, the step of determining whether there is a historical text stored in the database that includes at least one of the word segments and whose similarity to the text to be recognized meets a preset condition includes: when a historical text including at least one of the word segments is queried in the database, removing the historical texts whose relevance to the text to be recognized does not meet the preset rule from the queried historical texts.

[0009] Optionally, the removing the historical texts whose relevance to the text to be recognized does not meet the preset rule from the queried historical texts includes: sorting the queried historical texts in descending order of relevance to the text to be recognized; retaining the top preset number of historical texts and removing the remaining historical texts.

[0010] Optionally, the historical text that includes at least one of the word segments and whose similarity to the text to be recognized meets the preset condition is obtained through the following steps: obtaining the historical text with the maximum similarity to the text to be recognized among the retained historical texts; when the similarity between this historical text and the text to be recognized is greater than a preset first threshold, determining this historical text as the historical text that includes at least one of the word segments and whose similarity to the text to be recognized meets the preset condition.

[0011] Optionally, the historical text that includes at least one of the word segments and whose similarity to the text to be recognized meets the preset condition is obtained through the following steps: dividing the retained historical texts into at least one category according to the intention information of the historical texts; obtaining the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determining the historical texts in this category as the historical text that includes at least one of the word segments and whose similarity to the text to be recognized meets the preset condition.

[0012] Optionally, determining the intention information of the text to be recognized based on the intention information of the historical texts stored in the database includes: determining the intention information of the historical text that includes at least one of the word segments and whose similarity to the text to be recognized meets the preset condition as the intention information of the text to be recognized.

[0013] Optionally, the method further includes: when it is determined that there is no historical text stored in the pre-established database that includes at least one of the word segments and whose similarity to the text to be recognized meets the preset condition, determining the intention information of the text to be recognized by using a pre-established intention template set and / or a pre-trained intention classification model; wherein, the intention template set includes at least one intention template, and each intention template is configured with a rule representing a kind of intention information.

[0014] Optionally, the database is an Elastic Search; the similarity includes one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance; the degree of relevance between the historical text and the text to be recognized is determined by the following factors: the words to be recognized segmented and other segmented words included in the historical text, and the arrangement order of the words to be recognized segmented included in the historical text.

[0015] To achieve the above object, according to another aspect of the present invention, there is provided a text intention recognition device.

[0016] The text intention recognition device according to an embodiment of the present invention may include: a word segmentation unit for obtaining one or more segmented words in the text to be recognized; an intention recognition unit for: when it is determined that there is a historical text stored in a pre-established database that includes at least one of the segmented words and the similarity with the text to be recognized meets a preset condition, determining the intention information of the text to be recognized according to the intention information of the historical text stored in the database.

[0017] To achieve the above object, according to still another aspect of the present invention, there is provided a text intention recognition system.

[0018] The text intention recognition system according to an embodiment of the present invention may include: a pre-established database storing at least one historical text and the intention information of the historical text, and a similarity judgment unit; wherein, the database is configured to: in response to a query request carrying the text to be recognized, output a historical text including at least one segmented word of the text to be recognized; the similarity judgment unit is configured to: obtain a historical text whose similarity with the text to be recognized in the historical text output by the database meets a preset condition, and determine the intention information of the text to be recognized according to the intention information of the historical text.

[0019] Optionally, the historical text stored in the database may include: a text with an incorrect intention recognition result in a historical period and manually marked intention information; the database may further be configured to: arrange the historical texts including at least one segmented word of the text to be recognized in descending order of the degree of relevance with the text to be recognized, and output the top preset number of historical texts.

[0020] Optionally, the similarity judgment unit may further be configured to: obtain the historical text with the highest similarity to the text to be recognized from the historical texts output by the database; when the similarity between the historical text and the text to be recognized is greater than a preset first threshold, determine the intent information of the historical text as the intent information of the text to be recognized; or, divide the historical texts output by the database into at least one category according to the intent information of the historical texts, and obtain the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determine the intent information corresponding to this category as the intent information of the text to be recognized.

[0021] Optionally, the system may further include: a pre-established intent template set and a pre-trained intent classification model; wherein, the intent template set can be used to: when the database does not store historical texts containing at least one tokenized text to be recognized, or there is no historical text in the historical texts output by the database whose similarity to the text to be recognized meets the preset conditions, provide at least one intent template to attempt to match with the text to be recognized, and determine the intent information corresponding to the successfully matched intent template as the intent information of the text to be recognized; the intent classification model can be used to: when none of the intent templates in the intent template set match the text to be recognized successfully, receive the text to be recognized and output the intent information of the text to be recognized.

[0022] Optionally, the database is an Elastic Search; the similarity includes one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance; the degree of relevance between the historical text and the text to be recognized is determined by the following factors: the tokenized text to be recognized and other tokenized texts included in the historical text, and the arrangement order of the tokenized text to be recognized included in the historical text.

[0023] To achieve the above object, according to another aspect of the present invention, a text classification system is provided.

[0024] The text classification system according to the embodiments of the present invention may include: a pre-established database storing at least one historical text and the category information of the historical text in a preset dimension, and a similarity calculation unit; wherein, the database can be used to: in response to a query request carrying the text to be recognized, output historical texts containing at least one tokenized text to be recognized; the similarity calculation unit can be used to: obtain the historical texts in the historical texts output by the database whose similarity to the text to be recognized meets the preset conditions, and determine the category information of the text to be recognized according to the category information of the historical texts.

[0025] Optionally, the historical texts stored in the database may include: texts with incorrect classification results in a historical period and manually marked category information; the database may further be used to: sort the historical texts obtained by segmenting at least one text to be recognized in descending order of relevance to the text to be recognized, and output the top preset number of historical texts; and, the similarity calculation unit may further be used to: obtain the historical text with the highest similarity to the text to be recognized from the historical texts output by the database; when the similarity between this historical text and the text to be recognized is greater than a preset first threshold, determine the category information of this historical text as the category information of the text to be recognized; or, place the historical texts output by the database into at least one text set according to the category information of the historical texts; where the text sets correspond to the category information one by one; obtain the text set with the largest number of historical texts; when the average similarity between the historical texts in this text set and the text to be recognized is greater than a preset second threshold, determine the category information corresponding to this text set as the category information of the text to be recognized.

[0026] Optionally, the system may further include: a pre-established text template set and a pre-trained text classification model; where the text template set may be used to: when the database does not store historical texts obtained by segmenting at least one text to be recognized, or there is no historical text in the historical texts output by the database whose similarity to the text to be recognized meets the preset conditions, provide at least one text template to attempt to match with the text to be recognized, and determine the category information corresponding to the successfully matched text template as the category information of the text to be recognized; the text classification model may be used to: when none of the text templates in the text template set match the text to be recognized successfully, receive the text to be recognized and output the category information of the text to be recognized; and, the database may be an Elastic Search; the similarity may include one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance; the relevance between the historical text and the text to be recognized may be determined by the following factors: the text segments to be recognized and other segments included in the historical text, and the arrangement order of the text segments to be recognized included in the historical text.

[0027] To achieve the above object, according to another aspect of the present invention, an electronic device is provided.

[0028] An electronic device according to the present invention includes: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the text intention recognition method provided by the present invention.

[0029] To achieve the above object, according to another aspect of the present invention, a computer-readable storage medium is provided.

[0030] A computer-readable storage medium of the present invention stores a computer program thereon, and when the program is executed by a processor, the text intention recognition method provided by the present invention is implemented.

[0031] According to the technical solution of the present invention, an embodiment of the above invention has the following advantages or beneficial effects:

[0032] First, by establishing a database storing historical texts and their intention information, and querying similar texts of the text to be recognized in the database for intention recognition of the text to be recognized, an accurate and fast text intention recognition method is realized. On this basis, when an intention recognition error occurs, the correct intention of the corresponding text is manually marked and then the text and its correct intention are stored in the database. Thereafter, if the text or a similar text thereof is encountered again, accurate recognition can be achieved, thereby realizing timely hot repair of the intention recognition system (hot repair means that the repair process does not affect the system operation), avoiding defects such as long repair cycles and high cost risks in the existing repair methods of redevelopment of code or retraining of models, and also not requiring a long version online process. In addition, the intention recognition error cases continuously stored in the database are beneficial to the data analysis work of intention recognition, and can improve the classification performance of the intention classification model and the intention recognition system.

[0033] Second, when determining similar texts of the text to be recognized from the database, first obtain a preset number of historical texts that contain at least one word segmentation of the text to be recognized and have a relatively high degree of relevance to the text to be recognized, and then determine historical texts whose similarity meets the preset conditions for judging the intention of the text to be recognized. Through the above settings, the response speed of the system can be improved on the premise of ensuring the accuracy of intention recognition.

[0034] Third, the above database can be combined with the intention template set and the intention classification model in the prior art to form an intention recognition system. Among them, the database is used for quick recognition when storing similar texts of the text to be recognized and for timely repair of the system when an intention recognition error occurs, and the intention template set and the intention classification model are used for supplementary recognition when the database cannot provide recognition results, thereby realizing an intention recognition system that takes into account the text coverage range, recognition accuracy, and response speed.

[0035] The further effects of the above non-conventional optional methods will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0037] Figure 1 It is a schematic diagram of the main steps of the text intention recognition method in an embodiment of the present invention;

[0038] Figure 2 It is a schematic diagram of the components of the text intention recognition device in an embodiment of the present invention;

[0039] Figure 3 It is a schematic diagram of the components of the text intention recognition system in an embodiment of the present invention;

[0040] Figure 4 It is a schematic diagram of the components of the text classification system in an embodiment of the present invention;

[0041] Figure 5 It is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0042] Figure 6 It is a schematic diagram of the structure of an electronic device for implementing the text intention recognition method in an embodiment of the present invention. Detailed implementation manners

[0043] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding. It should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0044] It should be noted that, without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0045] Figure 1 It is a schematic diagram of the main steps of the text intention recognition method according to an embodiment of the present invention.

[0046] As Figure 1 shown, the text intention recognition method of the embodiments of the present invention can be specifically executed according to the following steps:

[0047] Step S101: Obtain one or more word segments in the text to be recognized.

[0048] In this step, the text to be recognized can be the text information input externally, or the text information converted from the externally input voice information. On the other hand, the text to be recognized can be the text formed in various languages, such as Chinese text, English text. The word segmentation in this step can be the words obtained after performing word segmentation on the text to be recognized, or the words included in the text to be recognized without word segmentation. In practical applications, Chinese text generally needs to go through word segmentation to obtain its word segmentation. English text can obtain its word segmentation after word segmentation, or directly use the gaps between words to determine the word segmentation without word segmentation. It should be noted that this step can be implemented by a separately developed program module, or by using the functions of the database to be introduced later. For example, if the database is Elasticsearch (ES), it has the function of text word segmentation.

[0049] Step S102: When it is determined that there is a historical text stored in the pre-established database that contains at least one of the word segments and whose similarity to the text to be recognized meets the preset conditions, determine the intention information of the text to be recognized according to the intention information of the historical text stored in the database.

[0050] In the embodiment of the present invention, the database is used to store historical texts and the intention information of historical texts, and it can be any applicable database such as ES, Mysql, MongoDB, etc. Taking ES as an example, it stores data in the form of records. In a record, the historical text is the value of the text field of the record, and the intention information of the historical text is the value of the intention field of the record. The text to be recognized and its intention information are shown in the following table.

[0051] Text to be recognized Intention information Play the story of Little Red Riding Hood Online radio Wake me up at 6 o'clock tomorrow morning Reminder Play a lively song Play music Turn off the TV Home control Who is the most beautiful in our family Chat

[0052] The similarity in this step can be one of the following similarities: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance (the greater the edit distance, the smaller the similarity), similarity based on Euclidean distance (the greater the Euclidean distance, the smaller the similarity), similarity based on Manhattan distance (the greater the Manhattan distance, the smaller the similarity), similarity based on Minkowski distance (the greater the Minkowski distance, the smaller the similarity). It can be understood that before calculating the cosine similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance, the text needs to be vectorized.

[0053] In this step, it is necessary to obtain historical texts stored in the database that contain the word segmentation of the text to be recognized and whose similarity to the text to be recognized meets the preset conditions as the similar texts of the text to be recognized, so as to determine the intention of the text to be recognized. There are two ways to obtain the above-mentioned similar texts. In the first way, first obtain historical texts containing the word segmentation of the text to be recognized from the database, and then judge the similarity of these historical texts to obtain similar texts. In the second way, first obtain historical texts whose similarity meets the preset conditions from the database, and then judge whether these historical texts contain the word segmentation of the text to be recognized, so as to obtain similar texts. The following will take the first way as an example to introduce the process of obtaining the above-mentioned similar texts.

[0054] After obtaining the word segmentations of the text to be recognized, one or more word segmentations can be used to query in the database, so as to obtain historical texts containing at least one word segmentation of the text to be recognized. In ES, the above query process can be implemented by using its fuzzy search function. After that, historical texts whose relevance to the text to be recognized does not meet the preset rules can be removed from the queried historical texts. In practical applications, the queried historical texts can be sorted in descending order of relevance to the text to be recognized, and the first preset number of historical texts can be retained, and the remaining historical texts can be removed.

[0055] It can be understood that the relevance is a measure of the correlation between the historical text and the text to be recognized, and can be determined by factors such as the word segmentations of the text to be recognized and other word segmentations contained in the historical text, and the arrangement order of the word segmentations of the text to be recognized contained in the historical text. For example, if the text to be recognized contains three word segmentations, the relevance between the historical text and the text to be recognized can be determined according to the following rules.

[0056] 1. First, set the relevance to a value between 0 and 1. After that, determine the relevance of the historical text containing the three word segmentations of the text to be recognized to be greater than 0.6, determine the relevance of the historical text containing two word segmentations of the text to be recognized to be between 0.3 and 0.6, and determine the relevance of the historical text containing one word segmentation of the text to be recognized to be less than 0.3.

[0057] 2. Among the historical texts with a correlation degree greater than 0.6, the correlation degree of the historical texts with the arrangement order of the three word segments being the same as that of the text to be recognized is determined to be greater than 0.8, and the correlation degree of the historical texts with the arrangement order of the three word segments being different from that of the text to be recognized is determined to be between 0.6 and 0.8; among the historical texts with a correlation degree greater than 0.8, the correlation degree of the historical texts that do not contain other word segments except the word segments of the text to be recognized is determined to be greater than 0.9, and the correlation degree of the historical texts that contain other word segments except the word segments of the text to be recognized is determined to be between 0.8 and 0.9; among the historical texts with a correlation degree between 0.6 and 0.8, the correlation degree of the historical texts that do not contain other word segments except the word segments of the text to be recognized is determined to be between 0.7 and 0.8, and the correlation degree of the historical texts that contain other word segments except the word segments of the text to be recognized is determined to be between 0.6 and 0.7.

[0058] 3. Among the historical texts with a correlation degree between 0.3 and 0.6, the correlation degree of the historical texts with the arrangement order of the two word segments being the same as that of the text to be recognized is determined to be between 0.45 and 0.6, and the correlation degree of the historical texts with the arrangement order of the two word segments being different from that of the text to be recognized is determined to be between 0.3 and 0.45; among the historical texts with a correlation degree between 0.45 and 0.6, the correlation degree of the historical texts that do not contain other word segments except the word segments of the text to be recognized is determined to be between 0.5 and 0.6, and the correlation degree of the historical texts that contain other word segments except the word segments of the text to be recognized is determined to be between 0.45 and 0.5; among the historical texts with a correlation degree between 0.3 and 0.45, the correlation degree of the historical texts that do not contain other word segments except the word segments of the text to be recognized is determined to be between 0.4 and 0.45, and the correlation degree of the historical texts that contain other word segments except the word segments of the text to be recognized is determined to be between 0.3 and 0.4.

[0059] 4. Among the historical texts with a correlation degree less than 0.3, the correlation degree of the historical texts that do not contain other word segments except the word segments of the text to be recognized is determined to be between 0.2 and 0.3, and the correlation degree of the historical texts that contain other word segments except the word segments of the text to be recognized is determined to be less than 0.2.

[0060] It should be noted that the above rules are only for demonstrating the measurement method of the correlation degree and do not limit the actual calculation process of the correlation degree. In specific applications, when calculating the correlation degree, in addition to considering the above factors, other factors such as the weights of the word segments of the text to be recognized included may also be considered.

[0061] After that, the similarity between the retained historical text and the text to be recognized can be calculated, and by determining whether the similarity meets the preset conditions, the similar text of the text to be recognized can be obtained. The following introduces three specific ways to obtain the similar text. It can be understood that the following ways are only examples and do not impose any restrictions on the preset conditions for similarity judgment. In fact, the preset conditions can be flexibly set according to the application environment and actual needs.

[0062] In the first way, first obtain the historical text with the greatest similarity to the text to be recognized in the retained historical text, and then compare the similarity between this historical text and the text to be recognized with the preset first threshold (this threshold is related to the application scenario and can be obtained by experience or through experiments): when the similarity is greater than the first threshold, determine this historical text as the similar text of the text to be recognized; when the similarity is not greater than the first threshold, it is considered that there is no similar text of the text to be recognized in the database.

[0063] In the second way, first divide the retained historical text into at least one category according to the intention information of the historical text, then obtain the category with the largest number of historical texts, and compare the average similarity (such as arithmetic mean, geometric mean, etc.) between the historical texts in this category and the text to be recognized with the preset second threshold (this threshold is related to the application scenario and can be obtained by experience or through experiments): when the average similarity is greater than the second threshold, determine the historical texts in this category as the similar text of the text to be recognized; when the average similarity is not greater than the second threshold, it is considered that there is no similar text of the text to be recognized in the database.

[0064] In the third way, it combines the first two ways, that is, first obtain the historical text with the greatest similarity to the text to be recognized in the retained historical text, and then compare the similarity between this historical text and the text to be recognized with the preset first threshold: when the similarity is greater than the first threshold, determine this historical text as the similar text of the text to be recognized; when the similarity is not greater than the first threshold, divide the retained historical text into at least one category according to the intention information of the historical text, obtain the category with the largest number of historical texts, and compare the average similarity between the historical texts in this category and the text to be recognized with the preset second threshold: when the average similarity is greater than the second threshold, determine the historical texts in this category as the similar text of the text to be recognized; when the average similarity is not greater than the second threshold, it is considered that there is no similar text of the text to be recognized in the database.

[0065] In step S102, after obtaining the similar text of the text to be recognized, the intention information of the text to be recognized can be determined based on the intention information of the similar text stored in the database. After obtaining the similar text of the text to be recognized through the above three ways, the intention information of the similar text can be determined as the intention information of the text to be recognized, thus realizing intention recognition.

[0066] Preferably, in the embodiments of the present invention, if no similar text of the text to be recognized is obtained through step S102, the intent information of the text to be recognized can be determined by using a pre-established intent template set or a pre-trained intent classification model. Specifically, the intent template set includes at least one intent template formed by using regular expressions or knowledge engineering, and each intent template is configured with a rule representing a kind of intent information. In use, the text to be recognized is respectively tried to be matched with each intent template. If the match is successful, the intent information corresponding to the corresponding template is determined as the intent information of the text to be recognized. The intent classification model can adopt algorithms such as decision trees and deep neural networks, and the model is trained by using manually labeled data. In practical applications, the intent template set and the intent classification model can also be combined for intent recognition, that is, first the text to be recognized is input into the intent template set, and when the recognition result is obtained, it is output. When the recognition result is not obtained, the text to be recognized is input into the intent classification model for judgment. Combining the database, the intent template set and the intent classification model can enable the system to meet the requirements of text coverage, recognition accuracy and response speed.

[0067] Particularly, the historical texts stored in the above database may include: texts with incorrect intent recognition results in historical periods and manually marked intent information. That is to say, during the process of performing intent recognition, if an incorrect recognition case is encountered, after the correct intent information of the corresponding text is manually marked, the text and its marked intent information are stored in the database. It can be understood that thereafter, if the text or a similar text of the text is encountered, the text stored in the database before can be located through the above process of obtaining similar texts, so that its correct intent can be directly displayed.

[0068] For example, if the system incorrectly predicts the intent of the text to be recognized "Play a lively song" as "Chat", then after learning about the error, the intent of "Play a lively song" is manually marked as "Play music", and the text and its marked intent are stored in ES. Thereafter, when the system faces "Play a lively song" or its similar texts "Play a happy song", "Play a lively tune", it can take the "Play a lively song" stored in ES as the similar text of the text to be recognized through the above-mentioned fuzzy retrieval and similarity judgment steps, and take the intent information "Play music" of the "Play a lively song" stored in ES as the intent recognition result.

[0069] With the above settings, when the system encounters a situation where the recognition result is incorrect and needs to be urgently repaired, it does not need to re-develop the code, re-train the model, or re-execute the version online process. It only needs to store the corresponding text and its correct intention in the database and perform online verification, thereby realizing the timely hot repair of the system, ensuring the normal operation of the system and the user experience, and avoiding the risks, large time costs, and labor costs brought by the original repair method.

[0070] In the technical solution of the embodiment of the present invention, first, by establishing a database storing historical texts and their intention information, and querying similar texts of the text to be recognized in the database for the intention recognition of the text to be recognized, an accurate and fast text intention recognition method is realized. On this basis, when an intention recognition error occurs, the correct intention of the corresponding text is manually marked and then the text and its correct intention are stored in the database. Thereafter, if the text or a similar text of the text is encountered again, it can be accurately recognized, thereby realizing the timely hot repair of the intention recognition system, avoiding the defects such as long repair cycles, high cost risks, etc. in the existing repair methods, and not needing to execute a long version online process. In addition, the intention recognition error cases continuously stored in the database are beneficial to the data analysis work of intention recognition, and can improve the classification performance of the intention classification model and the intention recognition system. Second, when determining similar texts of the text to be recognized from the database, first obtain a preset number of historical texts that contain at least one word segmentation of the text to be recognized and have a relatively high degree of relevance to the text to be recognized, and then determine historical texts whose similarity meets the preset conditions from them for judging the intention of the text to be recognized. Through the above settings, the response speed of the system can be improved while ensuring the accuracy of intention recognition. Third, the above database can be combined with the intention template set and the intention classification model in the prior art to form an intention recognition system. Among them, the database is used for quick recognition when storing similar texts of the text to be recognized and for the timely repair of the system when an intention recognition error occurs, and the intention template set and the intention classification model are used for supplementary recognition when the database cannot provide a recognition result, thereby realizing an intention recognition system that takes into account the text coverage, recognition accuracy, and response speed.

[0071] It should be noted that for the foregoing method embodiments, for the sake of description, they are expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence. In fact, some steps can be performed in other sequences or simultaneously. In addition, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for implementing the present invention.

[0072] To facilitate better implementation of the above solutions of the embodiments of the present invention, the following also provides related devices for implementing the above solutions.

[0073] Please refer to Figure 3 As shown, the text intention recognition device 200 provided by an embodiment of the present invention may include a word segmentation unit 201 and an intention recognition unit 202.

[0074] Among them, the word segmentation unit 201 can be used to obtain one or more word segments in the text to be recognized. The intention recognition unit 202 can be used to: when it is determined that there is a historical text stored in a pre-established database that contains at least one of the word segments and whose similarity to the text to be recognized meets a preset condition, determine the intention information of the text to be recognized according to the intention information of the historical text stored in the database.

[0075] In an embodiment of the present invention, the historical texts stored in the database may include: texts with incorrect intention recognition results in a historical period and whose intention information is manually marked.

[0076] In practical applications, the intention recognition unit 202 can further be used to: when a historical text containing at least one of the word segments is queried in the database, remove the historical texts whose degree of relevance to the text to be recognized does not meet the preset rules.

[0077] In specific applications, the intention recognition unit 202 can further be used to: sort the queried historical texts in descending order according to their degree of relevance to the text to be recognized; retain the top preset number of historical texts and remove the remaining historical texts.

[0078] Preferably, in an embodiment of the present invention, the intention recognition unit 202 can further be used to: obtain the historical text with the greatest similarity to the text to be recognized among the retained historical texts; when the similarity between this historical text and the text to be recognized is greater than a preset first threshold, determine this historical text as the historical text that contains at least one of the word segments and whose similarity to the text to be recognized meets the preset condition.

[0079] As a preferred solution, the intention recognition unit 202 can further be used to: divide the retained historical texts into at least one category according to the intention information of the historical texts; obtain the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determine the historical texts in this category as the historical texts that contain at least one of the word segments and whose similarity to the text to be recognized meets the preset condition.

[0080] In one embodiment, the intention recognition unit 202 can further be used to: determine the intention information of the historical text that contains at least one of the word segments and whose similarity to the text to be recognized meets the preset condition as the intention information of the text to be recognized.

[0081] In an optional implementation, the text intention recognition device may further include an auxiliary recognition unit, which may be used to: when it is determined that there is no historical text stored in the pre-established database that contains at least one of the word segments and whose similarity to the text to be recognized meets the preset conditions, determine the intention information of the text to be recognized by using the pre-established intention template set and / or the pre-trained intention classification model; wherein, the intention template set includes at least one intention template, and each intention template is configured with a rule representing a kind of intention information.

[0082] In addition, in the technical solution of the embodiments of the present invention, the database is ElasticSearch; the similarity may include one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance; the degree of relevance between the historical text and the text to be recognized may be determined by the following factors: the word segments of the text to be recognized and other word segments included in the historical text, and the arrangement order of the word segments of the text to be recognized included in the historical text.

[0083] In the technical solution of the embodiments of the present invention, first, by establishing a database storing historical texts and their intention information, and querying for similar texts of the text to be recognized in the database for intention recognition of the text to be recognized, an accurate and fast text intention recognition method is realized. On this basis, when an intention recognition error occurs, the correct intention of the corresponding text is manually marked and then the text and its correct intention are stored in the database. Thereafter, if the text or a similar text of the text is encountered again, it can be accurately recognized, thereby realizing the timely hot repair of the intention recognition system, avoiding the defects of long repair cycles and high cost risks in re-developing code or re-training models in the existing repair methods, and not requiring a long version online process. In addition, the intention recognition error cases continuously stored in the database are beneficial to the data analysis work of intention recognition, and can improve the classification performance of the intention classification model and the intention recognition system. Second, when determining the similar text of the text to be recognized from the database, first obtain a preset number of historical texts that contain at least one word segment of the text to be recognized and have a relatively high degree of relevance to the text to be recognized, and then determine the historical text whose similarity meets the preset conditions from them for judging the intention of the text to be recognized. Through the above settings, the response speed of the system can be improved on the premise of ensuring the intention recognition accuracy. Third, the above database can be combined with the intention template set and the intention classification model in the prior art to form an intention recognition system. Among them, the database is used for quick recognition when storing similar texts of the text to be recognized and for timely repair of the system when an intention recognition error occurs, and the intention template set and the intention classification model are used for supplementary recognition when the database cannot provide a recognition result, thereby realizing an intention recognition system that takes into account the text coverage, recognition accuracy, and response speed.

[0084] Figure 3 It is a schematic diagram of the components of the text intention recognition system in an embodiment of the present invention.

[0085] As Figure 3 shown, the text intention recognition system in an embodiment of the present invention may include: a pre-established database storing at least one historical text and the intention information of the historical text, and a similarity judgment unit.

[0086] Among them, the database may be any applicable database such as ES, Mysql, MongoDB, etc., which can be used to: in response to a query request carrying a text to be recognized, output a historical text containing at least one word segmentation of the text to be recognized. The similarity judgment unit can be used to: obtain the historical text in the historical text output by the database whose similarity with the text to be recognized meets a preset condition, and determine the intention information of the text to be recognized according to the intention information of the historical text. In practical applications, the similarity judgment unit can be implemented in the database or independent of the database. It can be understood that the system further includes an input unit for receiving input information and an output unit for displaying the intention recognition result.

[0087] In an embodiment of the present invention, the historical text stored in the database may include: texts with incorrect intention recognition results in a historical period and manually marked intention information; the database can be further used to: sort the historical texts containing at least one word segmentation of the text to be recognized in descending order of relevance to the text to be recognized, and output the top preset number of historical texts.

[0088] In an actual application scenario, the similarity judgment unit can be further used to: obtain the historical text with the greatest similarity to the text to be recognized in the historical text output by the database; when the similarity between the historical text and the text to be recognized is greater than a preset first threshold, determine the intention information of the historical text as the intention information of the text to be recognized; or, divide the historical text output by the database into at least one category according to the intention information of the historical text, and obtain the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determine the intention information corresponding to this category as the intention information of the text to be recognized.

[0089] In one embodiment, the system may further include: a pre-established set of intent templates and a pre-trained intent classification model. Specifically, the set of intent templates includes at least one intent template formed by using regular expressions or knowledge engineering, and each intent template is configured with a rule representing a type of intent information. The set of intent templates can be used to: when the database does not store historical texts containing at least one tokenized text to be recognized, or when there is no historical text in the historical texts output by the database whose similarity with the text to be recognized meets the preset conditions, provide at least one intent template to attempt to match with the text to be recognized, and determine the intent information corresponding to the successfully matched intent template as the intent information of the text to be recognized. The intent classification model can be a single model or a fusion model, and algorithms such as decision trees and deep neural networks can be used, and the model is trained with manually labeled data. The intent classification model can be used to: when none of the intent templates in the set of intent templates match the text to be recognized successfully, receive the text to be recognized and output the intent information of the text to be recognized.

[0090] In addition, in the embodiments of the present invention, the similarity may include one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance; the degree of relevance between the historical text and the text to be recognized can be determined by the following factors: the tokenized text to be recognized and other tokens included in the historical text, and the arrangement order of the tokenized text to be recognized included in the historical text.

[0091] Figure 4 It is a schematic diagram of the components of the text classification system in the embodiments of the present invention.

[0092] As Figure 4 shown, the text classification system in the embodiments of the present invention may include: a pre-established database storing at least one historical text and the category information of the historical text in a preset dimension, and a similarity calculation unit. Among them, the preset dimension can be various dimensions such as an intent dimension, an emotion dimension, etc. The category information of the intent dimension can be chatting, playing music, online radio, etc., and the category information of the emotion dimension can be neutral, angry, contemptuous, disgusted, fearful, happy, sad, surprised, etc.

[0093] The database can be any applicable database such as ES, Mysql, MongoDB, etc., which can be used to: in response to a query request carrying the text to be recognized, output historical texts containing at least one word segmentation of the text to be recognized. The similarity calculation unit can be used to: obtain historical texts in the historical texts output by the database whose similarity with the text to be recognized meets the preset conditions, and determine the category information of the text to be recognized based on the category information of this historical text. In practical applications, the similarity calculation unit can be implemented within the database or independent of the database. It can be understood that the system further includes an input unit for receiving input information and an output unit for displaying the classification result.

[0094] In the embodiment of the present invention, the historical texts stored in the database include: texts with incorrect classification results in the historical period and whose category information is manually marked. The database can be further used to: sort the historical texts containing at least one word segmentation of the text to be recognized in descending order according to the relevance to the text to be recognized, and output the top preset number of historical texts.

[0095] In practical application scenarios, the similarity calculation unit can be further used to: obtain the historical text with the greatest similarity to the text to be recognized in the historical texts output by the database; when the similarity between this historical text and the text to be recognized is greater than a preset first threshold, determine the category information of this historical text as the category information of the text to be recognized; or, place the historical texts output by the database into at least one text set according to the category information of the historical texts; where the text set corresponds to the category information one by one; obtain the text set with the largest number of historical texts; when the average similarity between the historical texts in this text set and the text to be recognized is greater than a preset second threshold, determine the category information corresponding to this text set as the category information of the text to be recognized.

[0096] In one embodiment, the system may further include: a pre-established set of text templates and a pre-trained text classification model. Specifically, the set of text templates includes at least one text template formed by using regular expressions or knowledge engineering, and each text template is configured with a rule representing a category of information. The set of text templates can be used to: when the database does not store historical texts containing at least one token of the text to be recognized, or when there is no historical text in the historical texts output by the database whose similarity with the text to be recognized meets the preset conditions, provide at least one text template to attempt to match with the text to be recognized, and determine the category information corresponding to the successfully matched text template as the category information of the text to be recognized. The text classification model can be a single model or a fusion model, and algorithms such as decision trees and deep neural networks can be used, and the model is trained with manually labeled data. The text classification model can be used to: when none of the text templates in the set of text templates match the text to be recognized successfully, receive the text to be recognized and output the category information of the text to be recognized.

[0097] In addition, in the embodiments of the present invention, the similarity may include one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance; the degree of relevance between the historical text and the text to be recognized can be determined by the following factors: the tokens of the text to be recognized and other tokens included in the historical text, and the arrangement order of the tokens of the text to be recognized included in the historical text.

[0098] Figure 5 An exemplary system architecture 500 to which the text intent recognition method or text intent recognition device according to the embodiments of the present invention can be applied is shown.

[0099] As Figure 5 shown, the system architecture 500 may include terminal devices 501, 502, 503, a network 504, and a server 505 (this architecture is only an example, and the components included in the specific architecture can be adjusted according to the specific situation of the application). The network 504 is used to provide a medium for communication links between the terminal devices 501, 502, 503 and the server 505. The network 504 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0100] Users can use the terminal devices 501, 502, 503 to interact with the server 505 through the network 504 to receive or send messages, etc. Various client applications, such as intent recognition applications (only examples), may be installed on the terminal devices 501, 502, 503.

[0101] The terminal devices 501, 502, and 503 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and so on.

[0102] The server 505 can be a server that provides various services, such as a background server (only an example) that supports the intent recognition application operated by the user using the terminal devices 501, 502, and 503. The background server can process the received intent recognition requests and feedback the processing results (such as the recognized intent information - only an example) to the terminal devices 501, 502, and 503.

[0103] It should be noted that the text intent recognition method provided by the embodiments of the present invention is generally executed by the server 505. Correspondingly, the text intent recognition device is generally arranged in the server 505.

[0104] It should be understood that Figure 5 the numbers of the terminal devices, networks, and servers in

[0105] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0106] The present invention also provides an electronic device. The electronic device according to the embodiments of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the text intent recognition method provided by the present invention.

[0106] Next, refer to Figure 6 which shows a schematic structural diagram of a computer system 600 suitable for implementing the electronic device according to the embodiments of the present invention. Figure 6 The shown electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0107] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0108] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 610 as required so that a computer program read therefrom is installed into the storage section 608 as required.

[0109] Specifically, according to the embodiments disclosed in the present invention, the process described in the above main step diagram can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the main step diagram. In the above embodiment, the computer program can be downloaded and installed from a network through the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit 601, the above functions defined in the system of the present invention are executed.

[0110] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in a block can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0112] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a word segmentation unit and an intent recognition unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the word segmentation unit can also be described as "a unit that provides word segmentation of the text to be recognized to the intent recognition unit".

[0113] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the steps performed by the device include: obtaining one or more word segmentations in the text to be recognized; when it is determined that there is a historical text stored in the pre-established database that contains at least one of the word segmentations and whose similarity to the text to be recognized meets a preset condition, determining the intent information of the text to be recognized according to the intent information of the historical text stored in the database.

[0114] In the technical solution of the embodiment of the present invention, first, by establishing a database storing historical texts and their intention information, and querying for similar texts of the text to be recognized in the database for intention recognition of the text to be recognized, an accurate and fast text intention recognition method is realized. On this basis, when an intention recognition error occurs, the correct intention of the corresponding text is manually marked and the text and its correct intention are stored in the database. Thereafter, if the text or a similar text of the text is encountered again, it can be accurately recognized, thereby realizing the timely hot repair of the intention recognition system, avoiding the defects of long repair cycles and high cost risks in the existing repair methods such as redevelopment of code or retraining of models, and not requiring the execution of a long version online process. In addition, the continuously stored intention recognition error cases are beneficial to the data analysis work of intention recognition, and can improve the classification performance of the intention classification model and the intention recognition system. Secondly, when determining the similar text of the text to be recognized from the database, first obtain a preset number of historical texts that contain at least one word segment of the text to be recognized and have a relatively high degree of relevance to the text to be recognized, and then determine the historical texts whose similarity meets the preset conditions for judging the intention of the text to be recognized. Through the above settings, the response speed of the system can be improved on the premise of ensuring the accuracy rate of intention recognition. Thirdly, the above database can be combined with the intention template set and the intention classification model in the prior art to form an intention recognition system. Among them, the database is used for quick recognition when storing similar texts of the text to be recognized and for timely repair of the system when an intention recognition error occurs, and the intention template set and the intention classification model are used for supplementary recognition when the database cannot provide recognition results, thereby realizing an intention recognition system that takes into account the text coverage range, recognition accuracy rate, and response speed.

[0115] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for text intention recognition, characterized in that, it includes: obtaining one or more word segments in the text to be recognized; when it is determined that there is a historical text stored in a pre-established database that contains at least one of the word segments and whose similarity to the text to be recognized meets a preset condition, determining the intention information of the text to be recognized according to the intention information of the historical text stored in the database; the historical texts stored in the database include: texts with incorrect intention recognition results in a historical period and whose intention information is manually marked; the step of determining whether there is a historical text stored in the database that contains at least one of the word segments and whose similarity to the text to be recognized meets a preset condition includes: when a historical text containing at least one of the word segments is queried in the database, removing the historical texts whose relevance to the text to be recognized does not meet the preset rules; the relevance of the historical text to the text to be recognized is determined by the following factors: the number of word segments of the text to be recognized contained in the historical text, the arrangement order of the word segments of the text to be recognized contained in the historical text, and whether there are other word segments contained in the historical text; dividing the remaining historical texts into at least one category according to the intention information of the historical texts; obtaining the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determining the historical texts in this category as the historical texts that contain at least one of the word segments and whose similarity to the text to be recognized meets the preset condition.

2. The method according to claim 1, characterized in that, the removing of the historical texts whose relevance to the text to be recognized does not meet the preset rules from the queried historical texts includes: sorting the queried historical texts in descending order of relevance to the text to be recognized; retaining the top preset number of historical texts and removing the remaining historical texts.

3. The method according to claim 2, characterized in that, the determining of the intention information of the text to be recognized according to the intention information of the historical text stored in the database includes: determining the intention information of the historical text that contains at least one of the word segments and whose similarity to the text to be recognized meets the preset condition as the intention information of the text to be recognized.

4. The method according to claim 1, characterized in that, the method further includes: when it is determined that there is no historical text stored in a pre-established database that contains at least one of the word segments and whose similarity to the text to be recognized meets a preset condition, determining the intention information of the text to be recognized by using a pre-established intention template set and / or a pre-trained intention classification model; wherein, the intention template set includes at least one intention template, and each intention template is configured with a rule representing a kind of intention information.

5. The method according to any one of claims 1-4, characterized in that, the database is Elastic Search, an elastic search engine; The similarity includes one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance.

6. A text intention recognition device, characterized in that, it includes: a word segmentation unit for obtaining one or more word segments in the text to be recognized; an intention recognition unit for: when it is determined that there is a historical text stored in a pre-established database that contains at least one of the word segments and the similarity with the text to be recognized meets a preset condition, determining the intention information of the text to be recognized according to the intention information of the historical text stored in the database; the historical texts stored in the database include: texts with incorrect intention recognition results in a historical period and with intention information manually marked; the intention recognition unit is further used for: when a historical text containing at least one of the word segments is queried in the database, removing the historical texts in the queried historical texts whose degree of relevance to the text to be recognized does not meet the preset rules; the degree of relevance between the historical text and the text to be recognized is determined by the following factors: the number of word segments of the text to be recognized contained in the historical text, the arrangement order of the word segments of the text to be recognized contained in the historical text, and whether there are other word segments in the historical text; dividing the remaining historical texts into at least one category according to the intention information of the historical texts; obtaining the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determining the historical texts in this category as the historical texts that contain at least one of the word segments and the similarity with the text to be recognized meets the preset condition.

7. A text intention recognition system, characterized in that, it includes: a pre-established database storing at least one historical text and the intention information of the historical text, and a similarity judgment unit; the historical texts stored in the database include: texts with incorrect intention recognition results in a historical period and with intention information manually marked; wherein, the database is used for: in response to a query request carrying the text to be recognized, outputting historical texts containing at least one word segment of the text to be recognized; the similarity judgment unit is used for: obtaining the historical texts whose similarity with the text to be recognized in the historical texts output by the database meets the preset condition, and determining the intention information of the text to be recognized according to the intention information of this historical text; the database is further used for: arranging the historical texts containing at least one word segment of the text to be recognized in descending order of the degree of relevance to the text to be recognized, and outputting the top preset number of historical texts; the degree of relevance between the historical text and the text to be recognized is determined by the following factors: the number of word segments of the text to be recognized contained in the historical text, the arrangement order of the word segments of the text to be recognized contained in the historical text, and whether there are other word segments in the historical text; The similarity judgment unit is further configured to: divide the reserved historical texts into at least one category according to the intention information of the historical texts; obtain the category with the largest number of historical texts; when the average similarity between the historical texts in this category and the text to be recognized is greater than a preset second threshold, determine the historical texts in this category as the historical texts that contain at least one of the word segmentations and whose similarity with the text to be recognized meets the preset conditions.

8. The system according to claim 7, wherein, the system further includes: a pre-established intention template set and a pre-trained intention classification model; wherein, the intention template set is configured to: when the historical texts containing at least one word segmentation of the text to be recognized are not stored in the database, or there are no historical texts in the historical texts output by the database whose similarity with the text to be recognized meets the preset conditions, provide at least one intention template to attempt to match with the text to be recognized, and determine the intention information corresponding to the successfully matched intention template as the intention information of the text to be recognized; the intention classification model is configured to: when none of the intention templates in the intention template set match the text to be recognized successfully, receive the text to be recognized and output the intention information of the text to be recognized.

9. The system according to claim 7, wherein, the database is an Elastic Search; the similarity includes one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance.

10. A text classification system, wherein, it includes: a pre-established database storing at least one historical text and the category information of the historical text in a preset dimension, and a similarity calculation unit; the historical texts stored in the database include: texts with incorrect classification results in a historical period and whose category information is manually marked; wherein, the database is configured to: in response to a query request carrying the text to be recognized, output historical texts containing at least one word segmentation of the text to be recognized; the similarity calculation unit is configured to: obtain the historical texts in the historical texts output by the database whose similarity with the text to be recognized meets the preset conditions, and determine the category information of the text to be recognized according to the category information of this historical text; the database is further configured to: arrange the historical texts containing at least one word segmentation of the text to be recognized in descending order of the degree of relevance to the text to be recognized, and output the top preset number of historical texts; the degree of relevance between the historical text and the text to be recognized is determined by the following factors: the number of word segmentations of the text to be recognized contained in the historical text, the arrangement order of the word segmentations of the text to be recognized contained in the historical text, and whether there are other word segmentations in the historical text; The similarity calculation unit is further configured to: place the historical texts output by the database into at least one text set according to the category information of the historical texts; wherein, the text sets correspond to the category information one by one; obtain the text set with the largest number of historical texts; when the average similarity between the historical texts and the text to be recognized in this text set is greater than a preset second threshold, determine the category information corresponding to this text set as the category information of the text to be recognized.

11. The system according to claim 10, wherein, the system further includes: a pre-established text template set and a pre-trained text classification model; wherein, the text template set is used to: when the database does not store historical texts including the word segmentation of at least one text to be recognized, or there is no historical text in the historical texts output by the database whose similarity with the text to be recognized meets the preset conditions, provide at least one text template to attempt to match with the text to be recognized, and determine the category information corresponding to the text template with successful matching as the category information of the text to be recognized; the text classification model is used to: when none of the text templates in the text template set match the text to be recognized successfully, receive the text to be recognized and output the category information of the text to be recognized; and, the database is an Elastic Search; the similarity includes one of the following: cosine similarity, Jaccard similarity, Pearson correlation coefficient, adjusted cosine similarity, similarity based on edit distance, similarity based on Euclidean distance, similarity based on Manhattan distance, similarity based on Minkowski distance.

12. An electronic device, wherein, comprising: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-5.

13. A computer-readable storage medium, on which a computer program is stored, wherein, the program, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Search intention identification method and device

    CN105095187A

  • Method and device for matching text

    CN107346344A

  • Document classification method and device and electronic equipment

    CN107844559A