Classifying information acquisition, classifying method, device, electronic equipment and storage medium
By extracting the second word related to the first word and their relationship from the query statement, the problem of insufficient classification accuracy and real-time performance in the existing technology is solved, and a wider and more accurate classification of query statements is achieved.
Patent Information
- Application Number
- CN202210399188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-04-15
AI Technical Summary
Existing technologies struggle to effectively utilize real-time query information when classifying query statements, resulting in insufficient classification accuracy and timeliness. This is especially problematic in emerging industries where significant resources are required for word extraction and classification.
By obtaining the first word and extracting its corresponding second word and related relationships from the query statement, relevant relationships are established to enrich the query classification information and improve classification accuracy and real-time performance.
It increases the scope and accuracy of query classification, and can dynamically adjust classification terms based on real-time query statements, thereby improving the real-time nature and accuracy of classification.
Smart Images

Figure CN114706956B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, in particular to the technical field of big data and artificial intelligence, and more particularly to a classification information acquisition method and device, a classification method and device, an electronic device, and a storage medium. BACKGROUND
[0002] In search technology, it is necessary to identify the intent of a query statement and perform a search based on the identified intent, and present the search results to the user.
[0003] The query statement can be classified using pre-generated words to determine the intent of the query statement. SUMMARY
[0004] The present disclosure provides a classification information acquisition method and device, a classification method and device, an electronic device, and a storage medium.
[0005] According to an aspect of the present disclosure, a classification information acquisition method is provided, comprising:
[0006] acquiring a first word;
[0007] In the query statement, a second word corresponding to the first word is determined, and a correlation between the first word and the second word is established;
[0008] The correlation, the first word, and the second word are determined as query classification information for classifying the query statement.
[0009] According to another aspect of the present disclosure, a classification method is provided, comprising:
[0010] acquiring an input statement input by a user;
[0011] In the query classification information, a target word corresponding to the input statement and a word related to the target word are queried, the type of the input statement is determined, and the query classification information is acquired according to the classification information acquisition method according to any embodiment of the present disclosure.
[0012] According to an aspect of the present disclosure, a classification information acquisition device is provided, comprising:
[0013] a first word acquisition module configured to acquire a first word;
[0014] a word and relationship determination module configured to determine, in a query statement, a second word corresponding to the first word, and establish a correlation between the first word and the second word;
[0015] The query classification information generation module is configured to determine the correlation, the first word and the second word as query classification information for classifying a query statement.
[0016] According to another aspect of the present disclosure, a classification device is provided, comprising:
[0017] The input sentence acquisition module is configured to acquire an input sentence input by a user.
[0018] The input sentence classification module is configured to query a target word corresponding to the input sentence and a word related to the target word in the query classification information, and determine a type of the input sentence, wherein the query classification information is acquired according to the classification information acquisition method of any one of the embodiments of the present disclosure.
[0019] According to another aspect of the present disclosure, an electronic device is provided, comprising:
[0020] at least one processor; and
[0021] a memory connected to the at least one processor in communication; wherein
[0022] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the classification information acquisition method of any one of the embodiments of the present disclosure, or the classification method of any one of the embodiments of the present disclosure.
[0023] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the classification information acquisition method of any one of the embodiments of the present disclosure, or the classification method of any one of the embodiments of the present disclosure.
[0024] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the classification information acquisition method of any one of the embodiments of the present disclosure, or the classification method of any one of the embodiments of the present disclosure.
[0025] The embodiments of the present disclosure can increase classification information and improve classification accuracy.
[0026] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0027] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:
[0028] Fig. 1 is a flowchart of a classification information acquisition method according to an embodiment of the present disclosure;
[0029] Fig. 2 is a flowchart of another classification information acquisition method according to an embodiment of the present disclosure;
[0030] Fig. 3 is a flowchart of a classification method according to an embodiment of the present disclosure;
[0031] Fig. 4 is a schematic diagram of another application scenario according to an embodiment of the present disclosure;
[0032] Fig. 5 is a structural diagram of a classification information acquisition apparatus according to an embodiment of the present disclosure;
[0033] Fig. 6 is a structural diagram of a classification apparatus according to an embodiment of the present disclosure;
[0034] Fig. 7 is a block diagram of an electronic device for implementing the classification information acquisition method or the classification method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help the understanding of the present disclosure. These should be considered in the context of the overall description and should not be considered limiting in nature. Thus, one of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0036] Fig. 1 is a flowchart of a classification information acquisition method according to an embodiment of the present disclosure, which can be applied to the case of generating classification information. The method of the present embodiment can be executed by a classification information acquisition apparatus, which can be implemented in software and / or hardware and specifically configured in an electronic device with certain data operation capability, which can be a client device or a server device, such as a mobile phone, a tablet computer, a vehicle terminal, a desktop computer, etc.
[0037] S101, obtaining a first word.
[0038] The first word is used as a reference classification word of query classification information, and expands more subdivided words to classify the query sentence more accurately. The first word can be extracted from the query sentence or input by the user. Optionally, obtaining the first word includes at least one of the following: obtaining interest information and extracting the first word; and obtaining a medium-tail query sentence and extracting the first word. The interest information can refer to information of interest to the user. The medium-tail query sentence refers to a query sentence with less search volume but search volume existing for a long time. The search volume of the medium-tail query sentence is less than that of the hot query sentence. The search volume existing for a long time can refer to the search volume existing in a preset time period. Specifically, the interest information of multiple users can be clustered, and the word representing each type is extracted as the first word. Alternatively, the collected medium-tail query sentence is screened to obtain a medium-tail query sentence with no clicks and large page views (PV), and the first word is obtained by performing word segmentation or entity extraction thereon. The interest information can be information input by an enterprise user. The medium-tail query sentence can be obtained by screening the query sentence of the enterprise user in the search system.
[0039] In S102, in the query sentence, a second word corresponding to the first word is determined, and a correlation between the first word and the second word is established.
[0040] The query sentence (query) refers to a sentence input by a user to be queried. The query sentence is input by a personal user. A large number of query sentences can be collected in advance, and for each query sentence, a second word is determined. It should be noted that the collected query sentences are authorized by the user and comply with relevant laws and regulations and do not violate public order and good customs. The second word refers to a word in the query sentence related to the first word, specifically a word expanded from the first word and having a certain degree of distinction. The second word is used to further classify the classification information represented by the first word. The first word and the second word can be similar but have different semantics. The correlation can refer to the relationship between the first word and the second word. The correlation is used to determine the corresponding another word according to the word, for example, to determine the first word according to the second word, or to further classify the first word to determine the second word. The second word can be understood as more specific classification query classification information associated with the first word. In fact, the second word and the first word are not isolated, and the correlation between the first word and the second word can be represented in a tree structure. At this time, the first word can be understood as a root node, and the multiple second words corresponding to the first word are child nodes of the first word.
[0041] S103, determine the correlation, the first word and the second word as query classification information for classifying the query statement.
[0042] The query classification information includes words and the correlation between the words. By classifying the query statement, the target word corresponding to the query statement can be queried in the query classification information, and the words related to the target word can be determined according to the correlation in the query classification information, and the expansion word corresponding to the query statement can be screened. Thus, the type of the query statement can be determined as the target word and the expansion word, the type of the query statement is enriched, and more accurate classification of the query statement is realized.
[0043] In the prior art, the query statement is usually classified according to the existing classification words. When a new industry needs to be built, a large amount of resources needs to be invested to extract the words in the new industry as classification words, and the query statement is classified according to the new classification words.
[0044] According to the technical solution of the present disclosure, by obtaining the first word, extracting the second word corresponding to the first word in the query statement, and establishing the correlation between the first word and the second word, the words and the correlation between the words are determined as the query classification information to classify the query statement, the classification words of the query classification information can be increased, the classification range of the query statement can be increased, and the classification accuracy of the query statement can be improved. Moreover, the words used for classification can be added according to the real-time obtained query statement, and the real-time performance of the query classification information can be improved.
[0045] Fig. 2 Another flowchart of a classification information acquisition method according to an embodiment of the present disclosure is disclosed. The above technical solution is further optimized and expanded, and can be combined with the above various optional embodiments. The second word corresponding to the first word in the query statement is determined, which is specifically: identifying the first entity in the query statement; obtaining the target keyword corresponding to the first word according to the first entity; and determining the second word according to the target keyword.
[0046] S201, obtaining a first word.
[0047] S202, identifying a first entity in a query statement.
[0048] In the query statement, an entity is identified, and the first entity is determined. The first entity usually refers to a proper noun. Specifically, the first entity can refer to a noun in the query statement. Among them, the entity recognition method can be a dictionary-based method, a statistical-based method, and an understanding-based method, etc. More specifically, the dictionary-based method refers to matching the string in the dictionary with the string in the query statement to obtain the entity. The statistical-based method, for example, is based on Hidden Markov Model (HMM), and multiple words with high adjacent appearance probability are determined as entities. The understanding-based method can be based on semantic information and syntactic information to identify text, for example, based on a pre-trained neural network model, input the query statement, and output the entity in the query statement. The first entity is used to filter out the words associated with the first word and added to the query classification information.
[0049] For example, the query statement is: How is the effect of hypoglycemic drug A? The first entity identified is: hypoglycemic drug and A.
[0050] In addition, some query statements are not related to the first word, and at this time the target keyword cannot be extracted from these query statements. For example, the first word is blood glucose, and the query statement is: Where is the toilet near XX intersection? The query statement is not related to the first word, and the word corresponding to the first word cannot be extracted in the query statement. Optionally, identifying the first entity in the query statement can include: screening a plurality of pre-collected query statements, and identifying the first entity in the screened query statements. Among them, the screening method can be to determine the query statement similar to the first word as the screened query statement. Among them, the query statement similar to the first word can calculate the similarity between the first word and the query statement through a pre-trained deep learning model, and can also calculate the similarity by extracting the text features of the first word and the text features of the query statement. The query statement with a similarity value greater than or equal to a preset similarity threshold is determined as the query statement similar to the first word. Among them, the deep learning model can be a neural network model, for example, it can be a convolutional neural network model, for example, it can be a language model, such as Enhanced Representation through Knowledge Integration (ERNIE) model, or Bidirectional Encoder Representations from Transformers (BERT) model, etc. The similarity threshold can be 0.7, the highest similarity is 1, and the lowest similarity is 0. It should be noted that the similarity threshold cannot be too high. The query statement similar to the first word is usually similar to the first word, but there is a certain degree of differentiation.
[0051] The first entity is obtained by screening the previously collected query sentences and identifying the screened query sentences, so as to screen the target keyword, reduce the detection data amount of the expansion word of the first word, and improve the detection accuracy of the expansion word of the first word.
[0052] S203, according to the first entity, obtaining the target keyword corresponding to the first word.
[0053] The target keyword refers to a word associated with the first word but having a certain degree of distinction among the plurality of first entities. The target keyword can refer to an expansion word of the first word. The target keyword is used to determine the second word. For example, the first word is blood sugar, and the target keyword corresponding to the first word is hypoglycemic drug and A according to the previous query sentence. For another example, the first word is a mobile phone, and the query sentence is: how is the performance of XX brand mobile phone? The first entity includes XX brand and mobile phone, and the target keyword corresponding to the first word is mobile phone.
[0054] The first entity similar to the first word can be screened from the plurality of first entities according to the similarity value between the first word and the first entity, and determined as the target keyword corresponding to the first word. In addition, the target keyword is different from the first word. The same word as the first word can also be excluded from the first entity.
[0055] Optionally, the target keyword corresponding to the first word is obtained according to the first entity, including: expanding the first word to obtain a similar sentence; respectively extracting features of the first word and the similar sentence to form a first feature vector; obtaining an average feature vector according to each first feature vector; extracting features of the first entity to form a second feature vector; according to the average feature vector and each second feature vector, screening the target keyword corresponding to the first word from each first entity.
[0056] The similar sentence refers to a sentence similar to the first word, wherein the sentence includes at least one of the following: word and sentence, etc. The expansion of the first word can be to obtain some sentences, query similar sentences to the first word from the sentences, determine the similar sentences of the first word, or obtain sentences input by the user, determine the similar sentences, or obtain the second word determined in the history, and obtain the query sentences similar to the first word, determine the similar sentences of the first word, etc. In addition, the query sentence similar to the first word can be the query sentence screened as described above.
[0057] The first word is extracted to obtain a first feature vector, the similar sentence is extracted to obtain a first feature vector, and the first entity is extracted to obtain a second feature vector. The average feature vector can be the average of the first feature vector, which is used to describe the features of the first word. The feature vector can be a feature representing the text, specifically a feature describing the semantics and character shape of the word. Feature extraction can be performed on the text by standard character representation. The feature vector extraction can be achieved by a feature extraction model, for example, the feature extraction model can be a support vector machine, a convolutional neural network model, a BERT model or an ERNIE model, etc.
[0058] In fact, the first word has only one word, and the first feature vector extracted is difficult to represent the features of the first word. Using only the first feature vector extracted from the first word and matching with the second feature vector of each first entity results in low accuracy of the matching result. The similar sentence of the first word can be added to extract the first feature vector, and the average feature vector can be calculated to improve the representativeness of the average feature vector for the first word, generalize the semantic information of the first word, and enrich the feature information of the first word.
[0059] According to the average feature vector and the second feature vector, the target keyword can be screened from the first entity, which is to calculate the similarity between the average feature vector and each second feature vector, and determine the first entity of the second feature vector with a similarity greater than or equal to a preset similarity threshold as the target keyword. The similarity between two vectors can be calculated by the distance between the two vectors.
[0060] By obtaining the similar sentence of the first word, and respectively extracting the features of the first word and the similar sentence to obtain the first feature vector, and taking the average to obtain the average feature vector, the feature information of the first word can be enriched, and the representativeness of the average feature vector for the first word can be improved. At the same time, based on the average feature vector with enriched feature information of the first word and the second feature vector extracted from each first entity, the target keyword can be screened from each first entity, which can increase the detection range of the target keyword related to the first word and improve the detection accuracy of the target keyword.
[0061] S204, according to the target keyword, determine a second word, and establish a correlation between the first word and the second word.
[0062] The second word can be determined according to the target keyword, can be the target keyword determined as the second word, or can be the target keyword further processed to obtain the second word. The correlation between the first word and the second word determined according to the query statement is established. In fact, multiple query statements can be collected, different query statements can determine multiple target keywords with the same character or the same semantics, the target keywords determined according to the multiple query statements can be de-duplicated to reduce the number of redundant target keywords, and the second word is determined according to the de-duplicated target keywords.
[0063] Optionally, the second word is determined according to the target keyword, including: extracting a second entity corresponding to the target keyword in the query statement; and determining the second word according to the target keyword and the second entity.
[0064] The second entity generally refers to a proper noun. Specifically, the second entity refers to a product noun. The second entity corresponding to the target keyword can refer to a second entity belonging to the type of the target keyword. The second word is determined according to the target keyword and the second entity, which can be the target keyword and the second entity determined as the second word, or the second entity determined as the second word. For example, the query statement is: Can A reduce blood sugar? The target keyword is a hypoglycemic drug, and the second entity corresponding to the target keyword is A. For another example, the query statement includes: How does a certain blood glucose meter work? The target keyword is a blood glucose meter, and the extracted second entity includes a certain blood glucose meter. In addition, the second entity can also include a photoelectric blood glucose meter, a photochemical blood glucose meter, or a blood glucose meter of XX brand, etc.
[0065] For example, the second entity corresponding to the target keyword can be extracted in the query statement based on a pre-trained neural network model. The input of the model is the query statement and the target keyword, and the output of the model is the second entity in the query statement. For example, the neural network model includes a convolutional neural network, a generative adversarial network, and an image neural network, etc. More specifically, the neural network module is a hybrid density network model (Mixture Density Networks).
[0066] The second entity corresponding to the target keyword is extracted in the query statement, which is actually a further expansion of the target keyword to determine the second entity that needs to be queried, and to establish a correlation between the target keyword and the first word and the first word to further enrich the query classification information.
[0067] Each query statement determines at least one target keyword. The target keywords can be summarized, and each keyword can be used to extract entities from each query statement to obtain the second entity.
[0068] By identifying the second entity in the query statement based on the target keyword, further expanding the target keyword, and adding the target keyword as a second word related to the first word to the query classification information, the classification information related to the first word can be increased, the range and accuracy of the query classification information can be increased, and the classification accuracy of the query statement can be improved.
[0069] Optionally, the establishing the correlation between the first word and the second word comprises: establishing a first-level correlation between the first word and the target keyword; and establishing a second-level correlation between the target keyword and the corresponding second entity.
[0070] The correlation can include multiple levels of correlation. The first-level correlation is used to represent that the target keyword is an expanded classification word of the first word, and the second-level correlation is used to represent that the second entity is an expanded classification word of the target keyword. In fact, the first word is divided into more specific multiple target keywords, and each target keyword can be divided into more specific multiple second entities. The first word can be understood as a parent node, the target keyword is a child node of the first word, the second entity is a child node of the target keyword, and the target keyword is a parent node of the second entity.
[0071] For example, in a case where it is determined that the type of the query statement is a target second entity, the type of the query statement can further include a target keyword related to the target second entity, and a target first word related to the target keyword, so that the classification information of the query statement can be increased. In addition, in a case where it is determined that the type of the query statement is a target keyword, the second entity of the query statement can be further detected for the second entity related to the target keyword, so that the classification information of the query statement can be increased, and the query statement can be more accurately determined, so that the classification granularity of the query statement can be increased, and the classification of the query statement can be flexibly adjusted.
[0072] By establishing multiple levels of correlation between the first word, the target keyword and the second entity, the correlation between the classification words of the query classification information can be enriched, and the query statement can be accurately classified.
[0073] S205, the correlation, the first word and the second word are determined as query classification information, which is used for classifying the query statement.
[0074] According to the technical solution of the present disclosure, by performing word segmentation on the query statement, the first entity is obtained, and in the first entity, the target keyword corresponding to the first word is screened to obtain the second word. The target keyword related to the first word can be accurately obtained from the real-time query statement, and the second word is determined to be used as a classification word to classify the query statement. The classification word of the first word extension can be accurately obtained, the specific classification branch of the first word is enriched, the classification range is increased, and the classification accuracy of the query statement is improved.
[0075] Fig. 3 is a flowchart of a classification method according to an embodiment of the present disclosure. The present embodiment can be applied to the case of classifying a query statement according to query classification information. The present embodiment method can be executed by a classification device, which can be implemented in software and / or hardware, and is specifically configured in an electronic device with certain data operation capability. The electronic device can be a client device or a server device, such as a mobile phone, a tablet computer, a vehicle terminal, and a desktop computer.
[0076] S301, obtaining an input statement input by a user.
[0077] The input statement is a query statement input by a user. A large number of input statements of users can be obtained. The input statements of users are obtained in accordance with relevant legal regulations and do not violate public order and good customs.
[0078] S302, in the query classification information, querying a target word corresponding to the input statement and a word related to the target word, determining a type of the input statement, and the query classification information is obtained according to the classification information obtaining method of any embodiment of the present disclosure.
[0079] The query classification information includes words and the correlation between words. It should be noted that some words in the query classification information have a correlation. Specifically, the query classification information includes a first word and a second word that has a correlation with the first word. The query classification information can also include a third word that has no correlation with other words. The target word can be a word similar to the input statement, which is used to determine the type of the input statement. The query classification information can be understood as a classification word library. Querying the target word of the input statement according to the query classification information is actually classifying the input statement and determining the type of the input statement as the target word.
[0080] In the query classification information, the target word corresponding to the input sentence is queried, and the related words of the target word are determined according to the related relationship corresponding to the target word. The target word and the related words are determined as the input sentence type. In addition, the related words of the target word can also be queried according to the related relationship corresponding to the related words of the target word, and the related words are also determined as the related words of the target word, which are used to determine the type of the input sentence.
[0081] For example, a rule-based classification method and / or a model-based classification method can be used to query the target word corresponding to the input sentence. The rule-based classification method can be to divide the input sentence into words, match the divided words with the words in the query classification information respectively, determine the words corresponding to the divided words, and determine the corresponding words and related words as the target words of the input sentence. The word corresponding to the divided word refers to the same word as the divided word. The model-based classification method can be to pre-train a machine learning model and input the input sentence into the trained machine learning model to obtain the target word corresponding to the input sentence. Using the two methods to determine the target word corresponding to the input sentence can process the target word corresponding to the input sentence, such as de-duplication to reduce repeated words in the target word, and update the target word corresponding to the input sentence.
[0082] Optionally, the classification method further comprises: classifying the user according to the type of the input sentence.
[0083] The target word corresponding to the input sentence is taken as the target word corresponding to the user inputting the input sentence. The target word can be understood as the label information of the user, and the user is classified according to the target word corresponding to the user. A large number of users are classified to obtain different types of user clusters, and the type of the user cluster can be represented by the same target word corresponding to each user in the user cluster. That is, the target word corresponding to the user in the user cluster can be determined as the label information of the user cluster, which can accurately classify the user and determine the type of the user. The user cluster can represent users interested in the same topic, and the users in the user cluster can be processed according to the application scenario, for example, the user in the user cluster pushes the information associated with the target word corresponding to the user, and for example, the number of users in the user cluster can be used to determine the frequency of pushing. For example, there is information to be pushed, the information to be pushed is determined according to the target word corresponding to the user cluster, the user cluster corresponding to the information to be pushed is queried, and the information to be pushed is sent to the user in the user cluster respectively, to realize accurate information pushing. In addition, there are other application scenarios, which can be processed according to the actual scenario.
[0084] By inputting the type of the sentence, the user providing the input sentence is classified, accurate classification of the user is realized, and the application scene is adapted to process users paying attention to the same topic, and the accuracy of data processing is improved.
[0085] Optionally, the query classification information includes words and the correlation between the words; the target word corresponding to the input sentence and the word related to the target word in the query include: determining the word to be updated from the words included in the query classification information according to the length and semantics of the word; adding the related word in the word to be updated according to the correlation between the words, updating the word to be updated; inputting the input sentence into a pre-trained classification model, and outputting the target word corresponding to the input sentence according to the updated word to be updated.
[0086] In fact, the words suitable for rule-based classification and the words suitable for model-based classification are different. Generally, short phrases and semantically single words are suitable for classification words used in rule-based classification methods. And longer words and polysemous words are suitable for classification words used in model-based classification methods. For example, XX is a brand, but it can also be understood as a kind of food. The classification words suitable for rule-based classification methods can be determined as rule words, and the classification words suitable for model-based classification methods can be determined as model words. Among them, since the model-based classification method usually classifies according to the semantic information of the word, the accuracy of polysemous classification is low, so the model word can be added with constraint information to make the semantic of the model word single and more clear, and improve the classification accuracy of the model. For example, XX adds food to get XX food, so it can be determined that XX food represents the semantic of a kind of food.
[0087] According to the length and semantics of each word in the query classification information, each word is classified to obtain words to be updated and non-words to be updated. Among them, the words to be updated are shorter and / or polysemous words; the non-words to be updated include longer and semantically single words. The words to be updated can be understood as the aforementioned model words, and the non-words to be updated can be understood as the aforementioned rule words.
[0088] For the to-be-updated term, the related terms of the to-be-updated term can be added based on the correlation between the terms in the query classification information, so as to add semantic constraints to the to-be-updated term and make the semantics of the to-be-updated term more accurate. Specifically, the correlation includes first-level correlation and second-level correlation, and the priority of the correlation can be preset. The to-be-added related terms are determined by selecting the correlation with high priority and adding them to the to-be-updated term. For example, the to-be-updated term corresponds to the correlation with high priority, and the related terms of the to-be-updated term are determined. In addition, the related terms of the to-be-updated term can also correspond to the correlation with high priority, and the related terms of the related terms are determined, which are also the related terms of the to-be-updated term. It can be understood that the terms of the parent node of the to-be-updated term are added to the to-be-updated term, rather than the terms of the child node of the to-be-updated term.
[0089] For example, blood glucose has a first-level correlation with blood glucose meter and hypoglycemic drug, blood glucose meter has a second-level correlation with photoelectric blood glucose meter, and hypoglycemic drug has a second-level correlation with A. For example, the to-be-updated term is A, the related term of A is hypoglycemic drug, and hypoglycemic drug is added to A to obtain hypoglycemic drug A. For example, the to-be-updated term is blood glucose meter, the priority of the first-level correlation is higher than that of the second-level correlation, and the related term of the first-level correlation of blood glucose meter is blood glucose. Therefore, blood glucose can be added to blood glucose meter to obtain blood glucose blood glucose meter.
[0090] The pre-classified classification model is used to query the target term corresponding to the input sentence from the plurality of terms, and the target term corresponding to the input sentence is determined as the type of the input sentence. In the embodiment of the present disclosure, the classification model is used to query the target term corresponding to the input sentence from the query classification information and the to-be-updated term after updating. In the case where the target term corresponding to the query sentence is the to-be-updated term after updating, the corresponding term of the to-be-updated term before updating can also be determined according to the correlation of the to-be-updated term in the query classification information, and the corresponding term is also determined as the target term corresponding to the query sentence. The classification model can be a deep learning model, and for example, it can be an ERNIE model.
[0091] By classifying the terms in the query classification information according to the length and semantics of the terms, the to-be-updated term is determined, and the related terms of the to-be-updated term in the query classification information are added to the to-be-updated term, so as to add semantic constraints to the to-be-updated term and make the semantics of the to-be-updated term more accurate. Therefore, the classification of the query sentence based on the to-be-updated term after updating can improve the classification accuracy of the classification model.
[0092] In addition, in the case of querying the words related to the target word, the related relationship with high priority in the target word can be queried to determine the words related to the target word. For example, blood glucose has a first-level related relationship with a blood glucose meter and a hypoglycemic drug, the blood glucose meter has a second-level related relationship with a photoelectric blood glucose meter, and the hypoglycemic drug has a second-level related relationship with A. The priority of the first-level related relationship is higher than that of the second-level related relationship, the target word is A, the words related to A are hypoglycemic drugs, and the words related to the hypoglycemic drugs in the first-level related relationship are blood glucose, so blood glucose, hypoglycemic drugs, and A can be determined as the types corresponding to the input sentence. For another example, the target word is a blood glucose meter, and the words related to the blood glucose meter in the first-level related relationship are blood glucose, so blood glucose is the word related to the blood glucose meter.
[0093] According to the technical solution of the present disclosure, by acquiring the query classification information and determining the target word of the input sentence, the classification accuracy of the input sentence can be improved, and based on the rich words included in the query classification information, the classification accuracy can be improved.
[0094] Fig. 4 is another application scenario disclosed according to an embodiment of the present disclosure. The method can include:
[0095] First, the query classification information is constructed:
[0096] Obtaining a first word: collecting long-tail query sentences that do not generate clicks in recent search query sentences of enterprise users, and screening query sentences with a page view greater than a preset view threshold. In addition, the interest information input by the enterprise user is obtained. From at least one of the screened query sentences and the interest information, the first word is extracted. The following is a detailed description of the first word "blood glucose" and the attention group.
[0097] In fact, simply using the first word "blood glucose" to make a similarity judgment with the query sentence is not enough for the recall of, for example, "insulin", "diabetes", or some hypoglycemic drugs, which will lead to insufficient coverage of the crowd cluster and inaccurate user classification. Therefore, it is necessary to expand the semantics of the first word.
[0098] The first entity is identified in the filtered query sentences. Still taking "blood sugar" as an example: the "blood sugar" is compared with the query sentences of the full amount of a single day for similarity discrimination. Specifically, the task-oriented ERNIE-sim model can be used to calculate the similarity between the first word and the collected query sentences, and the query sentences with a similarity value greater than a similarity threshold (for example, 0.7) are selected as the basic expansion query sentences of the first word for subsequent first entity identification. In order to obtain an expansion word with a certain degree of differentiation from the first word "blood sugar", the similarity threshold should not be too high. After obtaining the basic expansion query sentences, the first entity is extracted therefrom. For example, the three query sentences are "current blood sugar standard", "high blood sugar hypoglycemic drug", and "electronic blood sugar meter price". The two first entities "hypoglycemic drug" and "blood sugar meter" can be identified.
[0099] According to the first entity, a target keyword corresponding to the first word is obtained. Specifically, the first word is expanded to obtain a similar sentence; the first word and the similar sentence are respectively subjected to feature extraction to form a first feature vector; an average feature vector is obtained according to each first feature vector; a second feature vector is formed by feature extraction of the first entity; and the average feature vector and each second feature vector are used to screen a target keyword corresponding to the first word from the first entities.
[0100] The first entity is compared with the first word "blood sugar" for similarity discrimination, which is to extract other entities related to the first word "blood sugar". The feature vector of the first word "blood sugar" alone cannot express all the information contained in the first word "blood sugar", so the feature vector of the filtered query sentence similar to the first word "blood sugar" is extracted as supplementary information of the first word "blood sugar". In order to ensure the consistency of the dimension of the feature vector in the subsequent similarity calculation, the average feature vector is obtained by summing and averaging all the feature vectors, which is used as the feature vector of the first word "blood sugar". The average feature vector and the second feature vector extracted from the target keyword are compared for similarity discrimination. The similarity discrimination method and the similarity threshold can use the aforementioned ERNIE-sim model and similarity threshold.
[0101] For example, the first term "blood sugar" is not expanded, and the first feature vector of the first term "blood sugar" is directly compared with the second feature vector of the target keyword to determine the target keyword corresponding to the first term "blood sugar" and the target keyword before transformation. In the case of expanding similar sentences, the average feature vector of the first term "blood sugar" is compared with the second feature vector of the target keyword to determine the target keyword corresponding to the first term "blood sugar" and the target keyword after transformation. The following shows the comparison between the target keyword before transformation and the target keyword after transformation of the first term "blood sugar", as shown in Table 1:
[0102] Table 1
[0103] Before transformation After transformation Hypoglycemia Hypoglycemia Hypoglycemic agent Hypoglycemic agent Hyperglycemia Hyperglycemia Symptom of hyperglycemia Symptom of hyperglycemia Blood glucose measurement Blood glucose measurement Fasting blood glucose Fasting blood glucose Diabetes Diabetes Treatment of diabetes Treatment of diabetes Glucose Glucose meter
[0104] As can be seen from Table 1, the target keyword before transformation and the target keyword after transformation are different, and as can be seen from the last row, the target keyword after transformation is more similar to blood sugar.
[0105] In the query statement, the second entity corresponding to the target keyword is extracted; and the second term is determined according to the target keyword and the second entity. This step mainly identifies the product entity corresponding to the target keyword in the filtered query statement through entity extraction, and determines the second entity. For example, the target keyword "hypoglycemic drug" can be extracted to the corresponding specific product name, such as A and B, etc. For example, the NLPC-MONET operator in the natural language processing task (Natural Language Processing C, NLPC) based on C language can be used to realize custom entity recognition, and the filtered query statement and the target keyword are input, and the NLPC-MONET operator identifies the second entity in the filtered query statement according to the target keyword. For example, the second entity "A" can be extracted from the query statement "A can lower blood sugar". Table 2 lists the second entity identified in the query statement according to the target keyword:
[0106] Table 2
[0107] Entity word Brand word Glucose meter Three type glucose meter Glucose meter Three glucose meter Glucose meter Glucose meter Glucose meter Photochemical glucose meter Photochemical glucose meter Photoelectric glucose meter Photoelectric glucose meter Photoelectric glucose meter Hypoglycemic agent Degludec Hypoglycemic agent Pioglitazone Hypoglycemic agent Andatang
[0108] The embodiments of the present disclosure can automatically detect whether there is a certain communication tool on the website, and monitor the corresponding communication behavior data.
[0109] The first level of the relationship between the first word and the target keyword is established; the second level of the relationship between the target keyword and the corresponding second entity is established; the relationship, the first word and the second word are determined as query classification information, which is used for classifying the query sentence. According to the first word, a batch of target keywords and second entities corresponding to the first word are obtained. In addition, a plurality of first words can be obtained, and for each first word, a batch of corresponding target keywords and second entities can be expanded.
[0110] After the construction of the query classification information, the user can be classified according to the input sentence input by the user. In the classification process, the rule-based classification method and the model-based classification method can be used, the target word corresponding to the input sentence and the word related to the target word are determined as the type of the input sentence, and the user is classified according to the type of the input sentence. The same type of users is determined as a user cluster, and the type or label information of the user cluster is determined according to the type of the input sentence input by the user.
[0111] In the query classification information, the word suitable for rule classification is different from the word suitable for model classification. Generally, the word suitable for model classification contains the word suitable for rule classification.
[0112] For rule words: in the example of the first word "blood sugar", the target keywords can basically be used for rule discrimination, and some hypoglycemic drugs or blood glucose meters in the second entity can also be used as rule discrimination words.
[0113] Among them, the rule-based classification method cannot simply use the "contains" (in) method for discrimination. For example, when building the "hair loss concern group", a second entity "hair growth" is expanded. If only the "contains" (in) is used as the judgment method, it is easy to expand the "student development" and other bad cases, which affects the actual effect. Therefore, when classifying the query sentence by rule, the query sentence needs to be segmented first. "Student development" here will be divided into "student" and "development", and then matched with the rule word. In this way, the wrong words matched due to improper segmentation can be filtered out. That is, the input sentence is segmented, and according to the segmented words, the words in the query classification information are queried to determine the words corresponding to the input sentence.
[0114] For model words: target keywords can all be model words, but attention should be paid to the fact that, due to the large number of second entities identified, some bad cases are prone to occur in this step. Therefore, when the second entity is used, the prefix of the extracted target keyword and the second entity itself are spliced to serve as the basis for similarity discrimination, so as to reduce the influence of bad cases. This is equivalent to determining the to-be-updated word in the constructed query classification information, that is, the model word, and determining the word related to the to-be-updated word according to the relevant relationship in the query classification information, and adding it to the to-be-updated word. The classification model is based on the updated to-be-updated word to classify the input sentence.
[0115] Based on the model-based classification method, the target word corresponding to the input sentence can be determined based on the updated to-be-updated word and the query classification information, and the first relevant word can be determined according to the relevant relationship of the target word in the query classification information, and the type of the input sentence is also determined.
[0116] Through the above operations, the type of the user can be determined according to the query sentence input by the user, and it can be detected whether the user belongs to a user cluster of a specific focus topic (word).
[0117] According to the technical solution of the present disclosure, the classification information can be increased, the divided user cluster can cover many subfields, and the field corresponding to the hot spot can be automatically generated in real time based on the hot spot information, and the user of the type corresponding to the hot spot can be quickly determined, so as to generate a user cluster focusing on the hot spot, improve the accuracy and real-time performance of user classification, and increase the flexibility of user classification.
[0118] According to an embodiment of the present disclosure, Fig. 5 is a structure diagram of a classification information acquisition device in an embodiment of the present disclosure, and the present embodiment is applicable to the case of generating classification information. The device is implemented by software and / or hardware, and is specifically configured in an electronic device with certain data operation capability.
[0119] As shown in Fig. 5 A classification information acquisition device 500, comprising: a first word acquisition module 501, a word and relationship determination module 502, and a query classification information generation module 503; wherein,
[0120] The first word acquisition module 501 is configured to acquire a first word;
[0121] The word and relationship determination module 502 is configured to determine a second word corresponding to the first word in the query sentence, and establish a relevant relationship between the first word and the second word;
[0122] The query classification information generation module 503 is configured to determine the correlation, the first word and the second word as query classification information for classifying the query statement.
[0123] According to the technical scheme of the present disclosure, the first word is obtained, the second word corresponding to the first word is extracted in the query statement, and the correlation between the first word and the second word is established. The word and the correlation between the words are determined as query classification information to classify the query statement. The classification of the query classification information can be increased, the classification range of the query statement can be increased, the classification accuracy of the query statement can be improved, and the real-time of the query classification information can be improved according to the query statement obtained in real time.
[0124] Further, the word and relationship determination module comprises: a first entity acquisition unit configured to identify a first entity in the query statement; a keyword screening unit configured to obtain a target keyword corresponding to the first word according to the first entity; and a second word determination unit configured to determine a second word according to the target keyword.
[0125] Further, the second word determination unit comprises: a second entity acquisition unit configured to extract a second entity corresponding to the target keyword in the query statement; and a second word generation subunit configured to determine a second word according to the target keyword and the second entity.
[0126] Further, the word and relationship determination module comprises: a first-level correlation establishment unit configured to establish a first-level correlation between the first word and the target keyword; and a second-level correlation establishment unit configured to establish a second-level correlation between the target keyword and the corresponding second entity.
[0127] Further, the keyword screening unit comprises: a first word expansion subunit configured to expand the first word to obtain a similar statement; a first feature extraction subunit configured to extract features of the first word and the similar statement respectively to form a first feature vector; an average vector calculation subunit configured to obtain an average feature vector according to each first feature vector; a second feature extraction subunit configured to extract features of the first entity to form a second feature vector; and a target keyword determination subunit configured to screen a target keyword corresponding to the first word from each first entity according to the average feature vector and each second feature vector.
[0128] The above classification information acquisition device can execute the classification information acquisition method provided by any embodiment of the present disclosure, and has the corresponding function modules and beneficial effects of executing the classification information acquisition method.
[0129] According to an embodiment of the present disclosure, Fig. 6 is a structural diagram of a classification device in an embodiment of the present disclosure, and the embodiment of the present disclosure is applicable to the case of classifying an input sentence. The device is implemented by software and / or hardware, and is specifically configured in an electronic device with certain data operation capability.
[0130] As shown in a classification device 600 shown in Fig. 6 , the classification device 600 comprises an input sentence acquisition module 601 and an input sentence classification module 602; wherein,
[0131] The input sentence acquisition module 601 is configured to acquire an input sentence input by a user;
[0132] The input sentence classification module 602 is configured to query a target word corresponding to the input sentence and a word related to the target word in the query classification information, determine the type of the input sentence, and acquire the query classification information according to the classification information acquisition method as described in any embodiment of the present disclosure.
[0133] According to the technical solution of the present disclosure, the target word of the input sentence is determined by the acquired query classification information, and the type of the input sentence is determined, which can improve the classification accuracy of the input sentence. At the same time, based on the words included in the rich query classification information, the input sentence is classified, which can improve the accuracy of classification.
[0134] Further, the classification device further comprises a user classification module configured to classify the user according to the type of the input sentence.
[0135] Further, the query classification information comprises words and the correlation between the words; the input sentence classification module 602 comprises: a to-be-updated word acquisition unit configured to determine to-be-updated words from the words included in the query classification information according to the length and semantics of the words; a classification information updating unit configured to add related words to the to-be-updated words according to the correlation between the words, and update the to-be-updated words; and a sentence classification unit configured to input the input sentence into a pre-trained classification model, and output the target word corresponding to the input sentence according to the updated to-be-updated words.
[0136] The above classification device can execute the classification method provided by any embodiment of the present disclosure, and has the corresponding function modules and beneficial effects of executing the classification method.
[0137] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution comply with the relevant legal regulations and do not violate public order and good customs.
[0138] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0139] Fig. 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0140] As shown in Fig. 7 The device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded into a random access memory (RAM) 703 from a storage unit 708. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0141] Various components in the device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; the storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0142] The computing unit 701 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as the classification information acquisition method or the classification method. For example, in some embodiments, the classification information acquisition method or the classification method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RAM 703 and executed by the computing unit 701, one or more steps of the classification information acquisition method or the classification method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the classification information acquisition method or the classification method by any other appropriate means, such as by means of firmware.
[0143] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0144] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be embodied on the machine, partially on the machine, partially on the machine and partially on a remote machine or a server, or completely on a remote machine or server.
[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0147] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0148] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0149] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0150] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for obtaining classification information, comprising: obtaining a first word; the first word is used as a reference classification word for querying classification information; in a query sentence, determining a second word corresponding to the first word, and establishing a correlation between the first word and the second word; the second word is a word that is expanded from the first word and has a distinguishing degree; the second word is used to classify the classification information represented by the first word; determining the correlation, the first word and the second word as query classification information for classifying the query sentence; the determination of the second word corresponding to the first word in the query sentence comprises: identifying a first entity in the query sentence; obtaining a target keyword corresponding to the first word according to the first entity; determining a second word according to the target keyword; the determination of the second word according to the target keyword comprises: extracting a second entity corresponding to the target keyword in the query sentence; determining a second word according to the target keyword and the second entity.
2. The method of claim 1, wherein, the establishment of the correlation between the first word and the second word comprises: establishing a first-level correlation between the first word and the target keyword; establishing a second-level correlation between the target keyword and the corresponding second entity.
3. The method of claim 1, wherein, the obtaining of the target keyword corresponding to the first word according to the first entity comprises: expanding the first word to obtain a similar sentence; performing feature extraction on the first word and the similar sentence respectively to form a first feature vector; obtaining an average feature vector according to each first feature vector; performing feature extraction on the first entity to form a second feature vector; according to the average feature vector and each second feature vector, filtering the target keyword corresponding to the first word from each first entity. 4.A classification method, comprising: obtaining an input sentence input by a user; in query classification information, querying a target word corresponding to the input sentence and words related to the target word, and determining the type of the input sentence, wherein the query classification information is obtained according to the method for obtaining classification information according to any one of claims 1-3. 5.The method of claim 4, further comprising: classifying the user according to the type of the input sentence.
6. The method of claim 4, wherein, the query classification information comprises words and correlations between the words; the query of the target word corresponding to the input sentence and the words related to the target word comprises: determining a to-be-updated word from the words included in the query classification information according to the length and semantics of the words; adding related words to the to-be-updated word according to the correlations between the words, and updating the to-be-updated word; inputting the input sentence into a pre-trained classification model, and outputting the target word corresponding to the input sentence according to the updated to-be-updated word. 7.An apparatus for obtaining classification information, comprising: a first word obtaining module, configured to obtain a first word; the first word is used as a reference classification word for querying classification information; The word and relationship determining module is configured to determine a second word corresponding to the first word in the query statement and establish a correlation between the first word and the second word; the second word is a word that is expanded from the first word and has a distinguishing degree; and the second word is used to classify the classification information represented by the first word. The query classification information generating module is configured to determine the correlation, the first word, and the second word as query classification information used to classify the query statement. The word and relationship determining module includes: The first entity obtaining unit is configured to identify a first entity in the query statement. The keyword screening unit is configured to obtain a target keyword corresponding to the first word according to the first entity. The second word determining unit is configured to determine a second word according to the target keyword. The second word determining unit includes: The second entity obtaining unit is configured to extract a second entity corresponding to the target keyword in the query statement. The second word generating subunit is configured to determine a second word according to the target keyword and the second entity.
8. The apparatus of claim 7, wherein, The word and relationship determining module includes: The first-level correlation establishing unit is configured to establish a first-level correlation between the first word and the target keyword. The second-level correlation establishing unit is configured to establish a second-level correlation between the target keyword and the corresponding second entity.
9. The apparatus of claim 7, wherein, The keyword screening unit includes: The first word expanding subunit is configured to expand the first word to obtain a similar statement. The first feature extraction subunit is configured to extract features of the first word and the similar statement respectively to form a first feature vector. The average vector calculation subunit is configured to obtain an average feature vector according to the first feature vectors. The second feature extraction subunit is configured to extract features of the first entity to form a second feature vector. The target keyword determining subunit is configured to screen a target keyword corresponding to the first word from the first entities according to the average feature vector and the second feature vectors.
10. A classification device, comprising: An input statement obtaining module configured to obtain an input statement input by a user. An input statement classification module configured to query a target word corresponding to the input statement and words related to the target word in query classification information to determine a type of the input statement, wherein the query classification information is obtained according to the classification information obtaining method of any one of claims 1-3.
11. The device of claim 10, further comprising: A user classification module configured to classify the user according to the type of the input statement.
12. The apparatus of claim 10, wherein, The query classification information includes words and correlations between the words. The input statement classification module includes: A to-be-updated word obtaining unit configured to determine to-be-updated words from words included in the query classification information according to lengths and semantics of the words. The classification information updating unit is configured to add related words in the to-be-updated words according to the correlation between the words, and update the to-be-updated words. The sentence classification unit is configured to input the input sentence into a pre-trained classification model, and output target words corresponding to the input sentence according to the updated to-be-updated words. 13.An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein 14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the classification information obtaining method of any one of claims 1-3, or the classification method of any one of claims 4-6. The computer instructions are used to enable the computer to perform the classification information obtaining method of any one of claims 1-3, or the classification method of any one of claims 4-6. 15.A computer program product comprising a computer program which, when executed by a processor, implements the classification information obtaining method of any one of claims 1-3, or the classification method of any one of claims 4-6.
Citation Information
Patent Citations
Intelligent question answering method, apparatus, computer device and storage medium
CN109522393A
Semantic analysis method and device, electronic equipment and storage medium
CN110659366A