Method and apparatus for extracting keywords, electronic device, computer storage medium

By performing part-of-speech and syntactic analysis of questions in chat sessions, selecting subject, predicate and object vocabulary, and generating keyword collections, the problem that entity dictionary cannot extract non-noun vocabulary, and achieving more accurate keyword extraction and search results.

CN113392305BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011342868.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-25
Publication Date
2025-07-18
Estimated Expiration
2040-11-25

AI Technical Summary

Technical Problem

The prior art cannot effectively extract keywords that meet user needs, especially in the question scenarios in chat sessions, the entity dictionary cannot extract non-noun words, resulting in the search results that are inconsistent with user needs.

Method used

By obtaining the search text and judging it as the type of question, perform part-of-speech analysis and dependent syntax analysis, select the vocabulary of subject, predicate and object, generate a keyword set, use the convolutional neural network model to perform feature processing and syntactic relationship recognition, eliminate invalid vocabulary, and optimize keywords based on historical search records.

Benefits of technology

It realizes the flexibility to extract keywords that meet user needs from questions, improving the accuracy and user experience of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113392305B_ABST
    Figure CN113392305B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for extracting keywords, an electronic device, and a computer storage medium. Among them, the method includes: obtaining a search text; if the search text is in the form of a question, performing part-of-speech analysis on each word in the search text to obtain the part of speech of each word; using a dependency parsing algorithm to perform syntactic analysis on each word to obtain the syntactic relationship between every two words having a syntactic relationship; based on the syntactic relationship between every two words having a syntactic relationship, selecting the words corresponding to the subject, predicate, and object respectively from the words in the search text; using the words corresponding to the subject, predicate, and object respectively to generate a keyword set; the keyword set includes: the words corresponding to the subject, predicate, and object respectively, and the sentences composed of the words corresponding to the subject, predicate, and object respectively. Thus, keywords that meet the user's needs can be extracted from the question.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of information search, and particularly relates to a method and device for extracting keywords, an electronic device, and a computer storage medium. Background Art

[0002] In the field of artificial intelligence, with the development of Internet search engines, search demands are everywhere. Since there are also search demands in the scenario of chat conversations, many social software has gradually started to support session fingertip search, that is, users can select conversations in the chat and use the conversation as a query string to perform a search.

[0003] In order to further give the search initiative to users and provide users with a better experience, like the previous application scenario where users actively input words and sentences to perform a search, during the current search, multiple candidate keywords will also be provided to users for selection, so as to use the selected candidate keyword as the search keyword to further search for information. Similarly to the previous application scenario, in fingertip search, it is also based on a pre-constructed entity dictionary, and words included in the entity dictionary are selected from the conversations selected by the user as candidate keywords for the user to choose.

[0004] However, this method of extracting candidate keywords can only extract words in the entity dictionary and cannot extract words not included in the entity dictionary. Moreover, the words in the entity dictionary are usually nouns. In the previous application scenarios, users usually input nouns or declarative sentences for search, so extracting nouns can meet the information search needs of users. However, chat conversations often contain interrogative sentences. If only noun words are extracted, it cannot meet the information search needs of users. For example, Figure 1 as shown, what the user selects is "Does Application A need to add a graphics tablet function?". At this time, only "Application A" and "graphics tablet" are extracted as candidate keywords through the entity dictionary. Then, the information obtained by searching for "Application A" and "graphics tablet" later is obviously not very consistent with the user's initial search needs. Therefore, it can be seen that the existing method of extracting keywords cannot well extract keywords that meet the user's subsequent search needs for interrogative sentences in fingertip search. Summary of the Invention

[0005] Based on the above deficiencies of the prior art, this application provides a method and device for extracting keywords, an electronic device, and a computer storage medium to solve the problem that the prior art cannot effectively extract keywords that meet user needs.

[0006] To achieve the above object, this application provides the following technical solutions:

[0007] The first aspect of this application provides a method for extracting keywords, which is characterized by including:

[0008] Obtain the search text;

[0009] Determine whether the search text is of the question type;

[0010] If it is determined that the search text is of the question type, then perform part-of-speech analysis on each word of the search text to obtain the part of speech of each word;

[0011] Use the dependency syntax algorithm to perform syntax analysis on each word to obtain the syntactic relationship between every two words with a syntactic relationship;

[0012] Based on the syntactic relationship between every two words with a syntactic relationship, select the words corresponding to the subject, predicate, and object respectively from the words of the search text;

[0013] Use the words corresponding to the subject, predicate, and object respectively to generate a keyword set; wherein, the keyword set includes: the words corresponding to the subject, predicate, and object respectively, and the sentences composed of the words corresponding to the subject, predicate, and object respectively.

[0014] Optionally, in the above keyword extraction method, the determination of whether the search text is of the question type includes:

[0015] Perform word segmentation on the search text to obtain each word of the search text;

[0016] Perform feature processing on each word respectively to obtain the feature vector of each word;

[0017] Call the pre-trained convolutional neural network model to process the feature vector of each word to determine whether the search text is of the question type.

[0018] Optionally, in the above keyword extraction method, the training method of the convolutional neural network model includes:

[0019] Obtain a plurality of question titles and a plurality of news titles;

[0020] Use each of the question titles and each of the news titles as training sample data; wherein, each of the question titles is used as positive training sample data, and each of the news titles is used as negative training sample data;

[0021] Perform word segmentation on each of the training sample data respectively to obtain the sample words corresponding to each of the training sample data;

[0022] Perform feature processing on each of the sample words respectively to obtain the feature vectors of each of the sample words;

[0023] Input the feature vectors of each of the sample words corresponding to the training sample data into a convolutional neural network model, and perform calculations through the convolutional neural network model to obtain the classification result of the training sample data;

[0024] If the error between the classification result of the training sample data and the label of the training sample data is greater than a preset threshold, then adjust the parameters of the convolutional neural network model, and return to execute the step of inputting the feature vectors of each of the sample words corresponding to the training sample data into the convolutional neural network model; wherein, the label of the positive training sample data is 1, and the label of the negative training sample data is 0;

[0025] If the error between the classification result of the training sample data and the label of the training sample data is not greater than the preset threshold, then determine that the training of the convolutional neural network model is completed.

[0026] Optionally, in the above keyword extraction method, after selecting the words corresponding to the subject, predicate, and object respectively from each word in the search text, it further includes:

[0027] If there are words among the words corresponding to the subject that meet the merging criteria, then merge the words that meet the merging criteria into one word; wherein, the merging criteria are: there are multiple words corresponding to the subject, and their positions in the search text are continuous.

[0028] Optionally, in the above keyword extraction method, after selecting the words corresponding to the subject, predicate, and object respectively from each word in the search text, it further includes:

[0029] Compare the word corresponding to the predicate with a preset target predicate word;

[0030] Remove the word corresponding to the predicate that matches the target predicate word.

[0031] Optionally, in the above keyword extraction method, after generating a keyword set using the words corresponding to the subject, predicate, and object respectively, it further includes:

[0032] Retrieve the user's historical search records;

[0033] Use the historical search records to determine the historical search times of each word in the keyword set;

[0034] Remove the words with historical search times less than the preset number of times from the keyword set.

[0035] The second aspect of this application provides a keyword extraction device, including:

[0036] A first acquisition unit for acquiring a search text;

[0037] A judgment unit for judging whether the search text is of an interrogative sentence type;

[0038] A part-of-speech analysis unit for, when the search text is of an interrogative sentence type, performing part-of-speech analysis on each word in the search text to obtain the part of speech of each word;

[0039] A syntactic analysis unit for performing syntactic analysis on each word by using a dependency syntax algorithm to obtain the syntactic relationship between every two words having a syntactic relationship;

[0040] An extraction unit for selecting the words corresponding to the subject, predicate, and object respectively from each word in the search text based on the syntactic relationship between every two words having a syntactic relationship;

[0041] A generation unit for generating a keyword set by using the words corresponding to the subject, predicate, and object respectively; wherein, the keyword set includes: the words corresponding to the subject, predicate, and object respectively, and the sentences composed of the words corresponding to the subject, predicate, and object respectively.

[0042] Optionally, in the above keyword extraction device, the judgment unit includes:

[0043] A first word segmentation unit for segmenting the search text to obtain each word in the search text;

[0044] A first feature processing unit for respectively performing feature processing on each word to obtain the feature vector of each word;

[0045] A classification unit for calling a pre-trained convolutional neural network model to process the feature vector of each word to determine whether the search text is of an interrogative sentence type.

[0046] Optionally, in the above keyword extraction device, there is also a model training unit, wherein the model training unit includes:

[0047] A second acquisition unit for acquiring a plurality of question titles and a plurality of news titles;

[0048] A sample unit for using each of the question titles and each of the news titles as training sample data; wherein, each of the question titles is used as positive training sample data, and each of the news titles is used as negative training sample data;

[0049] A second word segmentation unit for respectively segmenting each of the training sample data to obtain the sample words corresponding to each of the training sample data;

[0050] A second feature processing unit, configured to perform feature processing on each of the sample words respectively to obtain the feature vectors of each of the sample words;

[0051] An input unit, configured to input the feature vectors of each of the sample words corresponding to the training sample data into a convolutional neural network model, and perform calculations through the convolutional neural network model to obtain the classification result of the training sample data;

[0052] A parameter adjustment unit, configured to adjust the parameters of the convolutional neural network model when the error between the classification result of the training sample data and the label of the training sample data is greater than a preset threshold, and return to execute inputting the feature vectors of each of the sample words corresponding to the training sample data into the convolutional neural network model; wherein, the label of the positive training sample data is 1, and the label of the negative training sample data is 0;

[0053] A first determination unit, configured to determine that the training of the convolutional neural network model is completed when the error between the classification result of the training sample data and the label of the training sample data is not greater than a preset threshold.

[0054] Optionally, in the above keyword extraction device, it further includes:

[0055] A merging unit, configured to merge the words that meet the merging criteria into one word when there are words that meet the merging criteria among the words corresponding to the selected subject; wherein, the merging criteria are: there are multiple words corresponding to the subject, and their positions in the search text are continuous.

[0056] Optionally, in the above keyword extraction device, it further includes:

[0057] A comparison unit, configured to compare the words corresponding to the predicate with the preset target predicate words;

[0058] A first elimination unit, configured to eliminate the words corresponding to the predicate that match the target predicate word.

[0059] Optionally, in the above keyword extraction device, it further includes:

[0060] An extraction unit, configured to extract the historical search records of the user;

[0061] A second determination unit, configured to use the historical search records to determine the historical search times of each of the words in the keyword set;

[0062] A second elimination unit, configured to eliminate the words with historical search times less than a preset number of times from the keyword set.

[0063] A third aspect of the present application provides a computer storage medium for storing a computer program, which when executed is used to implement the keyword extraction method described in any one of the above.

[0064] A fourth aspect of the present application provides an electronic device, including:

[0065] a memory and a processor;

[0066] wherein, the memory is used to store a program;

[0067] the processor is used to execute the program, which when executed is specifically used to implement the keyword extraction method described in any one of the above.

[0068] A keyword method provided by an embodiment of the present application, by obtaining a search text, and when determining that the search text is of an interrogative sentence type, performing part-of-speech analysis on each word in the search text to obtain the part-of-speech of each word, so that based on the part-of-speech of each word, using the dependency parsing algorithm to perform syntactic analysis on each word to obtain the syntactic relationship between every two words with a syntactic relationship, and finally based on the syntactic relationship between every two words with a syntactic relationship, selecting the words corresponding to the subject, predicate, and object respectively from each word in the search text, and using the words corresponding to the subject, predicate, and object respectively to generate a keyword set including each selected word and the sentences composed of the selected words. Thus, based on the dependency parsing algorithm, the extraction of keywords from interrogative sentences is realized, rather than using an entity dictionary to extract keywords, making the extraction of keywords more flexible and better applicable to diverse chat conversations, and further being able to accurately extract keywords that meet the user's needs from interrogative sentences. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0070] Figure 1 It is a schematic diagram of the operation interface for fingertip search;

[0071] Figure 2 It is a flowchart of a keyword extraction method provided by an embodiment of the present application;

[0072] Figure 3 It is a flowchart of a method for determining whether a search text is of an interrogative sentence type provided by another embodiment of the present application;

[0073] Figure 4 Flow chart of a training method for a convolutional neural network model provided by another embodiment of the present application;

[0074] Figure 5 Schematic diagram of the structure of a convolutional neural network model provided by another embodiment of the present application;

[0075] Figure 6 Schematic diagram of an example of lexical part-of-speech analysis provided by another embodiment of the present application;

[0076] Figure 7 Schematic diagram of an example of syntactic analysis provided by another embodiment of the present application;

[0077] Figure 8 Schematic diagram of the structure of a keyword extraction device provided by another embodiment of the present application;

[0078] Figure 9 Schematic diagram of the structure of a model training unit provided by another embodiment of the present application;

[0079] Figure 10 Schematic diagram of the structure of an electronic device provided by another embodiment of the present application. Detailed implementation manners

[0080] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0081] In the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0082] This application provides a method for extracting keywords, which is mainly applied in the field of artificial intelligence. Among them, artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, techniques, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0083] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0084] More specifically, the method for extracting keywords provided by this application belongs to the direction of natural language processing (NLP) in the field of artificial intelligence. Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0085] The embodiment of this application provides a method for extracting keywords, as Figure 2 shown, which specifically includes the following steps:

[0086] S201. Obtain the search text.

[0087] First of all, it should be noted that the embodiment of this application mainly takes fingertip search as an example for illustration, but the method provided by this application is not limited to being used in fingertip search.

[0088] Among them, the search text refers to the text content input or selected by the user for searching. For fingertip search, it can be triggered by performing a specified operation on the chat session. For example, as Figure 1As shown, by long - pressing the current chat session "Does A application need to add a handwriting board function", the user can select the text in the chat session. At this time, the system will correspondingly pop up virtual buttons such as the search button "Search" and copy and cut. After the user clicks the search button "Search", as Figure 1 shown in the right figure of, the system will pop up a search interface, use the selected chat session as the search text, perform information search and feedback the search results to the user. Therefore, when the user selects a session for search, the search text can be obtained. While the system conducts this search for the search term, keyword extraction can be performed on the obtained search text for the user to select for subsequent searches.

[0089] S202. Determine whether the search text is of the question - type.

[0090] It should be noted that the keyword extraction method provided in the embodiments of the present application is for search text of the question - type, that is, for interrogative sentences.

[0091] Therefore, if it is determined that the search text is of the question - type, then step S203 is executed.

[0092] Optionally, referring to Figure 3 , a method for determining whether the search text is of the question - type after obtaining the search text is shown, which specifically includes the following steps:

[0093] S301. Segment the search text to obtain each word of the search text.

[0094] Among them, it should be noted that each word obtained refers to a text unit, and the text unit can be a word or just a single character, etc., and is not limited to words composed of two or more characters.

[0095] Optionally, an existing word - segmentation tool, such as Jieba word - segmentation, can be used to segment the search text, and at the same time, meaningless words such as "de, de, ne, yao" and other stop words can be removed.

[0096] S302. Perform feature processing on each word to obtain the feature vector of each word.

[0097] Optionally, existing tools can also be used to perform feature processing on the words, such as the word2vec model.

[0098] Specifically, several search texts with relatively high occurrence frequencies can be collected in advance, each of these search texts is segmented to obtain multiple words, and the set of these words is called a word bag.

[0099] Subsequently, use this bag of words for training to obtain a word vector model (i.e., the word2vec model). Based on this trained word vector model, when performing step S302, the word vectors of each vocabulary obtained by word segmentation in step S301 can be found from the word vector model.

[0100] For example, the search text can be "May I ask if the function of adding a graphics tablet needs to be considered in Tencent Meeting?", and the word segmentation of this search text can obtain: May I ask, Tencent, Meeting, in, whether, consider, add, graphics tablet, function, and then find the word vectors of each vocabulary from the word vector model.

[0101] S303: Call a pre-trained convolutional neural network model to process the feature vectors of each vocabulary, and determine whether the search text is of the question type.

[0102] Specifically, the feature vectors of each vocabulary obtained by word segmentation of the search text are used as a whole to be input into a pre-trained convolutional neural network model. The convolutional neural network model processes the feature vectors of the input vocabularies to obtain the probability that the search text belongs to the question type. If the probability that the search text output by the convolutional neural network model belongs to the question type is greater than the preset probability value, it is determined that the search text is of the question type.

[0103] Optionally, in the embodiment of the present application, a convolutional neural network model is used as a binary classifier for determining whether the search text is of the text type. Of course, it can also be trained by using other models, and a binary classifier for determining whether the search text is of the text type can be obtained.

[0104] Optionally, in another embodiment of the present application, a training method for the convolutional neural network model in step S303 is provided, as Figure 4 shown, which specifically includes the following steps:

[0105] S401: Obtain a plurality of question titles and a plurality of news titles.

[0106] It should be noted that since chat conversations involve user privacy issues, it is not appropriate to obtain users' chat conversations and screen out training samples from them. Therefore, in the embodiment of the present application, question titles and news titles are used as training samples. Of course, if authorized by the user, users' chat conversations can also be used as training samples for model training.

[0107] Among them, the question titles mainly refer to the titles of questions raised by users on Q&A websites or clients. For example, the question titles on Baidu Knows. These question titles all belong to the interrogative sentence type, that is, the questions asked by users, and this data can reflect the types of questions raised by most users, the ways of asking questions, etc., which is more conducive to the training of the model. News titles, on the other hand, are usually of the declarative sentence type, so they can be used as negative training samples.

[0108] S402. Use each question title and each news title as training sample data.

[0109] Among them, each question title serves as positive training sample data, and each news title serves as negative training sample data. At the same time, it is also necessary to label each training sample data. The label of the positive training sample data is 1, and the label of the negative training sample data is 0.

[0110] S403. Segment each training sample data respectively to obtain the sample vocabulary corresponding to each training sample data.

[0111] Similarly, existing word segmentation tools can be used to segment the training samples and perform subsequent processing such as removing stop words, so as to obtain the sample vocabulary corresponding to each training sample data.

[0112] S404. Perform feature processing on each sample vocabulary respectively to obtain the feature vectors of each sample vocabulary.

[0113] Similarly, existing feature processing models can be used to perform feature processing on each sample vocabulary, so as to obtain the feature vectors of each abridged vocabulary.

[0114] S405. Input the feature vectors of each sample vocabulary corresponding to the training sample data into the convolutional neural network model, and perform calculations through the convolutional neural network model to obtain the classification results of the training sample data.

[0115] Specifically, as Figure 5 shown, the convolutional neural network model mainly includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer. Input the feature vectors of each sample vocabulary of the training sample data into the convolutional neural network model through the input layer, and after passing through the convolutional layer, pooling layer, and fully connected layer in sequence, input the classification results of the training sample data.

[0116] S406. Determine whether the error between the classification result of the training sample data and the label of the training sample data is greater than the preset threshold.

[0117] Among them, if the error between the classification result of the training sample data and the label of the training sample data is greater than the preset threshold, step S407 is executed. If the error between the classification result of the training sample data and the label of the training sample data is not greater than the preset threshold, step S408 is executed.

[0118] S407. Adjust the parameters of the convolutional neural network model.

[0119] Specifically, the way to adjust the parameters can be based on the gradient descent method. Specifically, determine the loss function for the parameters of the convolutional neural network model, and then take the partial derivatives of each parameter in the loss function to obtain the corresponding gradient vectors. Then, along the opposite direction of each parameter's corresponding gradient vector, adjust each corresponding parameter by a preset step size, and after adjusting the parameters, return to step S405 to continue training the model.

[0120] S408. Determine that the training of the convolutional neural network model is completed.

[0121] S203. Perform part-of-speech analysis on each word in the search text to obtain the part of speech of each word.

[0122] Optionally, an existing publicly available part-of-speech tagging tool postag can be used to tag the part of speech of each word. For example, as Figure 6 shown, for the search text "May I ask if the function of adding a graphics tablet is considered in Tencent Meeting", the words obtained after word segmentation are "May I ask, Tencent, Meeting, in, whether, consider, add, graphics tablet, function", and then perform part-of-speech analysis on each word to obtain the analysis result as Figure 6 shown. Among them, n refers to the part of speech of a noun, v refers to the part of speech of a verb, nz and nd respectively refer to proper nouns and other nouns.

[0123] S204. Use the dependency parsing algorithm to perform syntactic analysis on each word to obtain the syntactic relationship between every two words with a syntactic relationship.

[0124] Among them, the basic task of syntactic analysis is to determine the syntactic structure of a sentence or the dependency relationship between words in a sentence. It mainly includes two aspects. One is to determine the grammar system of the language, that is, to give a formal definition of the grammar structure of legal sentences in the language; the other is the syntactic analysis technology, that is, according to the given grammar system, automatically deduce the syntactic structure of the sentence and analyze the syntactic units included in the sentence and the relationship between these syntactic units.

[0125] Among them, some common syntactic relationships are shown in Table 1 below.

[0126] Table 1

[0127]

[0128]

[0129] Therefore, based on the part-of-speech of each word, the dependency parsing algorithm is used to perform syntactic analysis on each word to obtain the syntactic relationship between every two words with a syntactic relationship. Among them, the specific process of using the dependency parsing algorithm for analysis is the same as the existing one, which will not be elaborated here. For example, as Figure 7 shown, for the words "Excuse me, Tencent, meeting, in, whether, consider, add, tablet PC, function" obtained after word segmentation, after syntactic analysis, the syntactic relationship as Figure 7 shown is obtained.

[0130] S205. Based on the syntactic relationship between every two words with a syntactic relationship, select the words corresponding to the subject, predicate, and object respectively from the words in the search text.

[0131] After obtaining the syntactic relationship between every two words with a syntactic relationship, then according to the syntactic relationship between each word and the part-of-speech of each word, the grammatical components played by each word in the search text can be determined. Since the subject, predicate, and object are usually the backbone of a sentence, that is, the key components in a sentence, and the meaning expressed by a sentence is usually contained in the subject, predicate, and object, therefore, in the embodiments of the present application, the words corresponding to the subject, predicate, and object are extracted respectively. It should be noted that since chat conversations are relatively casual, the search text does not necessarily contain all three parts of the subject, predicate, and object at the same time. However, when extracting words, the words corresponding to the subject, predicate, and object will be searched separately, and any word found corresponding to the subject, predicate, or object will be extracted.

[0132] Optionally, in another embodiment of the present application, after performing step S205, the following can be further performed: If there are words among the words corresponding to the subject that meet the merging criteria, then the words that meet the merging criteria are merged into one word.

[0133] Among them, the merging criteria are: there are multiple words corresponding to the subject, and their positions in the search text are continuous.

[0134] Because when multiple words are in the subject position and their positions in the search text are continuous, these words are often a complete name in themselves, but are divided into multiple words during word segmentation. For example, the two words "Tencent, meeting" in the above example should be merged into "Tencent meeting". If not merged, the extracted words will be different from the meaning originally expressed by the search text, and thus provide keywords that do not meet the user's needs to the user.

[0135] In addition, since the search text may not be a complete and smooth sentence that conforms to grammatical rules, there may be some meaningless predicates that have no effect on subsequent searches, so they can be removed. Therefore, in another embodiment of the present application, after performing step S205, it may further include: comparing the vocabulary corresponding to the predicate with a preset target predicate word. If a vocabulary corresponding to the predicate that matches the target predicate word is found, the vocabulary corresponding to the predicate that matches the target predicate word is removed. Among them, the target predicate word refers to a preset meaningless predicate word, such as: "have, know, go, play", etc. For example, for the search text "There are scenic spots in Beijing", the predicate "have" is obviously meaningless, and the extracted "Beijing" and "scenic spots" can obviously meet the user's subsequent search needs, so it is removed.

[0136] S206. Generate a keyword set by using the vocabulary corresponding to the subject, predicate, and object respectively.

[0137] Among them, the keyword set includes: the vocabulary corresponding to the subject, predicate, and object respectively, and the sentences composed of the vocabulary corresponding to the subject, predicate, and object respectively. That is, in the keyword set, it not only contains each extracted vocabulary, but also includes the complete sentences composed of these vocabularies, that is, the complete sentences exactly composed of the subject, predicate, and object together. For example, for the search text "May I ask if the function of adding a graphics tablet needs to be considered in Tencent Meeting?", tokenizing this search text can obtain: May I ask, Tencent, Meeting, in, whether, consider, add, graphics tablet, function. Also referring to Figure 6 the syntactic relationship shown, the vocabulary corresponding to the subject is: Tencent Meeting, the vocabulary corresponding to the predicate is: add, the vocabulary corresponding to the object is: graphics tablet, and the vocabulary corresponding to the subject, predicate, and object respectively can exactly form the sentence "Tencent Meeting adds a graphics tablet", so the generated keyword set is: Tencent Meeting, add, graphics tablet, Tencent Meeting adds a graphics tablet.

[0138] Obviously, the composed sentence is equivalent to the text after the search text is streamlined. When the search results obtained after the user selects the search text for searching are not accurate enough or do not meet the requirements, etc., in subsequent searches, the user will hope to conduct further searches. If the selected feedback vocabulary is used for searching, the obtained search results often deviate greatly from the information that the user originally hoped to search for, resulting in unexpected results. The sentences in the keyword set contain more content than the vocabulary, and are more streamlined than the search text. Therefore, compared with each vocabulary, they can well meet the user's need for further searching.

[0139] It should be noted that since chat conversations are relatively casual, the search text does not necessarily contain all three parts of the subject, predicate, and object at the same time. Therefore, it does not necessarily form a complete sentence, and the keyword set does not necessarily contain the sentences formed by the selected individual words.

[0140] Optionally, in order to provide a keyword set that conforms to the user's search habits, in another embodiment of the present application, after obtaining the keyword set in step S206, it may further include: retrieving the user's historical search records, and then using the historical search records to determine the historical search times of each word in the keyword set, and removing the words with historical search times less than the preset times from the keyword set. After removing the words with fewer user search times, the words and sentences in the keyword set are fed back to the user for the user to select for subsequent information search.

[0141] A method for keywords provided by an embodiment of the present application includes obtaining a search text, and when determining that the search text is in the form of a question, performing part-of-speech analysis on each word in the search text to obtain the part of speech of each word, so that based on the part of speech of each word, the dependency syntax algorithm can be used to perform syntax analysis on each word to obtain the syntactic relationship between every two words with a syntactic relationship. Finally, based on the syntactic relationship between every two words with a syntactic relationship, the words corresponding to the subject, predicate, and object are selected from the words in the search text respectively, and a keyword set including each selected word and the sentences formed by the selected words is generated. Thus, for questions, keywords are no longer extracted using an entity dictionary, making the extraction of keywords more flexible and better applicable to questions, and thus being able to accurately extract keywords that meet the user's needs from questions.

[0142] Another embodiment of the present application provides a keyword extraction device, as Figure 8 shown, including the following units:

[0143] A first acquisition unit 801, configured to acquire a search text.

[0144] A judgment unit 802, configured to judge whether the search text is in the form of a question.

[0145] A part-of-speech analysis unit 803, configured to perform part-of-speech analysis on each word in the search text to obtain the part of speech of each word when the search text is in the form of a question.

[0146] A syntax analysis unit 804, configured to perform syntax analysis on each word based on the part of speech of each word by using a dependency syntax algorithm to obtain the syntactic relationship between every two words with a syntactic relationship.

[0147] An extraction unit 805, configured to select the words corresponding to the subject, predicate, and object respectively from each word in the search text based on the syntactic relationships between every two words with syntactic relationships.

[0148] A generation unit 806, configured to generate a keyword set by using the words corresponding to the subject, predicate, and object respectively.

[0149] Wherein, the keyword set includes: each selected word and / or a sentence composed of the selected words.

[0150] Optionally, in the keyword extraction device provided in another embodiment of the present application, the following units may further be included:

[0151] A first word segmentation unit, configured to perform word segmentation on the search text to obtain each word in the search text.

[0152] A first feature processing unit, configured to perform feature processing on each word respectively to obtain the feature vector of each word.

[0153] A classification unit, configured to call a pre-trained convolutional neural network model to process the feature vector of each word to determine whether the search text is of an interrogative sentence type.

[0154] Optionally, in the keyword extraction device provided in another embodiment of the present application, a model training unit is further included. Wherein, the model training unit, as Figure 9 shown, includes the following units:

[0155] A second acquisition unit 901, configured to acquire a plurality of question titles and a plurality of news titles.

[0156] A sample unit 902, configured to use each question title and each news title as training sample data.

[0157] Wherein, each question title is used as positive training sample data, and each news title is used as negative training sample data.

[0158] A second word segmentation unit 903, configured to perform word segmentation on each training sample data respectively to obtain the sample words corresponding to each training sample data.

[0159] A second feature processing unit 904, configured to perform feature processing on each sample word respectively to obtain the feature vector of each sample word.

[0160] An input unit 905, configured to input the feature vectors of the sample words corresponding to the training sample data into the convolutional neural network model, and perform calculations through the convolutional neural network model to obtain the classification result of the training sample data.

[0161] A parameter adjustment unit 906, configured to adjust the parameters of the convolutional neural network model when the error between the classification result of the training sample data and the label of the training sample data is greater than a preset threshold, and return to execute inputting the feature vectors of the respective sample words corresponding to the training sample data into the convolutional neural network model.

[0162] Among them, the label of the positive training sample data is 1, and the label of the negative training sample data is 0.

[0163] A first determination unit 907, configured to determine that the training of the convolutional neural network model is completed when the error between the classification result of the training sample data and the label of the training sample data is not greater than a preset threshold.

[0164] Optionally, in the keyword extraction device provided in another embodiment of the present application, it may further include:

[0165] A merging unit, configured to merge the words that meet the merging criteria into one word if there are words that meet the merging criteria among the words corresponding to the selected subject.

[0166] Among them, the merging criterion is: there are multiple words corresponding to the subject, and their positions in the search text are continuous.

[0167] Optionally, in the keyword extraction device provided in another embodiment of the present application, it may further include the following units:

[0168] A comparison unit, configured to compare the words corresponding to the predicate with the preset target predicate words.

[0169] A first elimination unit, configured to eliminate the words corresponding to the predicate that match the target predicate words.

[0170] Optionally, in the keyword extraction device provided in another embodiment of the present application, it may further include the following units:

[0171] An access unit, configured to access the user's historical search records.

[0172] A second determination unit, configured to use the historical search records to determine the historical search times of each word in the keyword set.

[0173] A second elimination unit, configured to eliminate the words with historical search times less than a preset number from the keyword set.

[0174] It should be noted that the specific working processes of the respective units provided in the above embodiments of the present application may be correspondingly referred to the specific working processes of the corresponding units in the above method embodiments, and will not be elaborated here.

[0175] A third aspect of the present application provides a computer storage medium for storing a computer program, which, when executed, is used to implement the keyword extraction method described in any one of the above.

[0176] Computer storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data signals and carrier waves.

[0177] Another embodiment of the present application provides an electronic device, as Figure 10 shown, including:

[0178] A memory 1001 and a processor 1002.

[0179] Among them, the memory 1001 is used to store a program, and the processor 1002 is used to execute the program stored in the memory 1001. When the program is executed, it is specifically used to implement the keyword extraction method provided in any one of the above embodiments.

[0180] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0181] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting keywords, characterized in that, include: Get the search text; Determining whether the search text is a question type; If it is determined that the search text is a question type, then performing part-of-speech analysis on each word in the search text to obtain the part-of-speech of each word; Performing syntactic analysis on each of the words using a dependency syntax algorithm to obtain a syntactic relationship between every two words having a syntactic relationship; Based on the syntactic relationship between each two of the words having a syntactic relationship, selecting words corresponding to the subject, the predicate and the object respectively from the words in the search text; If there are words that meet the merging criteria among the words corresponding to the selected subjects, the words that meet the merging criteria are merged into one word; wherein the merging criteria are: there are multiple words corresponding to the subject, and the positions of the words in the search text are continuous; Comparing the vocabulary corresponding to the predicate with the preset target predicate words, and eliminating the vocabulary corresponding to the predicate that matches the target predicate words; Generate a keyword set using the obtained vocabulary corresponding to the subject, predicate and object respectively; wherein the keyword set includes: the vocabulary corresponding to the obtained subject, predicate and object respectively, and a sentence composed of the vocabulary corresponding to the obtained subject, predicate and object respectively; the sentence is a keyword that better meets the user's further search needs than the vocabulary corresponding to the obtained subject, predicate and object respectively; When the search results obtained by searching using the words corresponding to the subject, predicate and object respectively do not meet the expected results, further search is performed using sentences composed of the words corresponding to the subject, predicate and object respectively.

2. The method according to claim 1, wherein The step of determining whether the search text is a question type includes: Segmenting the search text to obtain each word of the search text; Performing feature processing on each of the words respectively to obtain a feature vector of each of the words; A pre-trained convolutional neural network model is called to process the feature vector of each of the words to determine whether the search text is a question type.

3. The method according to claim 2, wherein The training method of the convolutional neural network model comprises: Get multiple question titles and multiple news titles; Each of the question titles and each of the news titles is used as training sample data; wherein each of the question titles is used as positive training sample data, and each of the news titles is used as negative training sample data; Segmenting each of the training sample data respectively to obtain sample vocabulary corresponding to each of the training sample data; Performing feature processing on each of the sample words respectively to obtain a feature vector of each of the sample words; Inputting the feature vectors of each of the sample words corresponding to the training sample data into a convolutional neural network model, and performing calculations through the convolutional neural network model to obtain classification results of the training sample data; If the error between the classification result of the training sample data and the label of the training sample data is greater than a preset threshold, then adjust the parameters of the convolutional neural network model, and return to execute the step of inputting the feature vectors of each of the sample words corresponding to the training sample data into the convolutional neural network model; wherein, the label of the positive training sample data is 1, and the label of the negative training sample data is 0; If the error between the classification result of the training sample data and the label of the training sample data is not greater than the preset threshold, then it is determined that the training of the convolutional neural network model is completed.

4. The method according to claim 1, characterized in that, After generating the keyword set by using the words corresponding to the subject, predicate, and object respectively, it further includes: Retrieve the user's historical search records; Use the historical search records to determine the historical search times of each of the words in the keyword set; Remove the words with historical search times less than the preset times from the keyword set.

5. An apparatus for extracting keywords, characterized in that, It includes: A first acquisition unit for acquiring a search text; A judgment unit for judging whether the search text is of an interrogative sentence type; A part-of-speech analysis unit for performing part-of-speech analysis on each of the words in the search text to obtain the part-of-speech of each of the words when the search text is of an interrogative sentence type; A syntactic analysis unit for performing syntactic analysis on each of the words by using the dependency syntactic algorithm to obtain the syntactic relationship between every two words with a syntactic relationship; An extraction unit for selecting the words corresponding to the subject, predicate, and object respectively from each of the words in the search text based on the syntactic relationship between every two words with a syntactic relationship; A merging unit for merging the words that meet the merging criteria into one word if there are words that meet the merging criteria among the words corresponding to the subject; wherein, the merging criteria are: there are multiple words corresponding to the subject, and their positions in the search text are consecutive; A comparison unit for comparing the word corresponding to the predicate with a preset target predicate word; A first removal unit for removing the word corresponding to the predicate that matches the target predicate word; A generation unit for generating a keyword set by using the words corresponding to the subject, predicate, and object respectively obtained; wherein, the keyword set includes: the words corresponding to the subject, predicate, and object respectively obtained, and the sentences composed of the words corresponding to the subject, predicate, and object respectively obtained; the sentences are keywords that are more in line with the user's further search needs than the words corresponding to the subject, predicate, and object respectively obtained; The device is further configured to, in the case where the search result obtained by searching by using the words corresponding to the subject, predicate, and object respectively obtained does not meet the expected result, perform a further search by using the sentence composed of the words corresponding to the subject, predicate, and object respectively obtained.

6. The device according to claim 5, characterized in that, The judgment unit includes: A first word segmentation unit for performing word segmentation on the search text to obtain each of the words in the search text; A first feature processing unit for performing feature processing on each of the words respectively to obtain the feature vector of each of the words; A classification unit is used to call a pre-trained convolutional neural network model to process the feature vectors of each of the said words and determine whether the search text is of the question type.

7. The device according to claim 6, characterized in that, It further includes: A model training unit; The model training unit includes: A second acquisition unit for acquiring a plurality of question titles and a plurality of news titles; A sample unit for using each of the said question titles and each of the said news titles as training sample data; wherein, each of the said question titles is used as positive training sample data, and each of the said news titles is used as negative training sample data; A second word segmentation unit for segmenting each of the said training sample data to obtain sample words corresponding to each of the said training sample data; A second feature processing unit for performing feature processing on each of the said sample words to obtain feature vectors of each of the said sample words; An input unit for inputting the feature vectors of each of the said sample words corresponding to the training sample data into a convolutional neural network model, and performing calculations through the convolutional neural network model to obtain classification results of the training sample data; A parameter adjustment unit for, if the error between the classification result of the training sample data and the label of the training sample data is greater than a preset threshold, adjusting the parameters of the convolutional neural network model and returning to execute inputting the feature vectors of each of the said sample words corresponding to the training sample data into the convolutional neural network model; wherein, the label of the positive training sample data is 1, and the label of the negative training sample data is 0; A first determination unit for, if the error between the classification result of the training sample data and the label of the training sample data is not greater than a preset threshold, determining that the training of the convolutional neural network model is completed.

8. The device according to claim 5, characterized in that It further includes: A retrieval unit for retrieving the user's historical search records; A second determination unit for using the historical search records to determine the historical search times of each of the words in the keyword set; A second elimination unit for eliminating words with historical search times less than a preset number from the keyword set.

9. A computer storage medium, characterized in that, For storing a computer program, which when executed is used to implement the keyword extraction method according to any one of claims 1 to 4.

10. An electronic device, characterized in that, It includes: A memory and a processor; Wherein, the memory is used to store a program; The processor is used to execute the program, and when the program is executed, it is specifically used to implement the keyword extraction method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method for realizing similarity calculation of questions in vertical field

    CN110532566A

  • Text keyword processing method, device and equipment apparatus

    CN111339751A