A method, apparatus, electronic device, and storage medium for obtaining keywords
By determining candidate keywords in the target text and building a correlation diagram, the problem of inaccurate search results based on selected keywords is solved, and more accurate information acquisition effect is achieved.
Patent Information
- Application Number
- CN202011301926.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-11-19
AI Technical Summary
The search results based on the selected keywords are inaccurate and cannot effectively meet the user's information acquisition needs.
By obtaining multiple keywords in the target text, determining candidate keywords whose vocabulary number is less than or equal to the threshold value with the selected keyword, constructing an association graph to obtain the word importance value of the candidate keyword, and finally obtaining the target keyword from the candidate keyword.
It improves the accuracy of search results, meets users' information acquisition needs, and provides more relevant text content.
Smart Images

Figure CN113392177B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technology, and more specifically, to a method, apparatus, electronic device, and storage medium for obtaining keywords. Background Art
[0002] With the development of the Internet, users can browse texts through electronic devices to obtain corresponding information. During the process of browsing texts, users may have a need to search for other texts related to a certain keyword contained in the text. For example, when a user browses text A through an electronic device and browses the keyword A contained in text A, the user wants to view texts related to keyword A. The user can select keyword A in text A and then conduct a search.
[0003] Currently, the texts obtained by searching based on a selected keyword (such as keyword A) are very likely not the texts that the user needs, that is, the search results are inaccurate. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus, electronic device, and storage medium for obtaining keywords to at least solve the problem that the search results obtained by searching based on a selected keyword are inaccurate.
[0005] This application provides the following technical solutions:
[0006] According to a first aspect of an embodiment of the present disclosure, a method for obtaining keywords is provided, including:
[0007] Obtain multiple keywords included in a target text, where the target text includes a selected keyword in a selected state;
[0008] Based on the first positions of the multiple keywords in the target text respectively, determine at least one candidate keyword from the multiple keywords whose number of intervening words from the second position of the selected keyword in the target text is less than or equal to a first threshold;
[0009] Based on the at least one candidate keyword and the selected keyword, obtain an association graph; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and the weight of the edge between the two nodes is the relevance of the two nodes;
[0010] Based on the association graph, obtain the word importance values corresponding to the at least one candidate keyword respectively;
[0011] Based on the word importance values corresponding to the at least one candidate keyword respectively, obtain a target keyword from the at least one candidate keyword.
[0012] According to a second aspect of the embodiments of the present disclosure, there is provided a keyword acquisition device, including:
[0013] A first acquisition module, configured to acquire a plurality of keywords included in a target text, where the target text includes a selected keyword in a selected state;
[0014] A first determination module, configured to determine, based on first positions of the plurality of keywords in the target text respectively, at least one candidate keyword from the plurality of keywords, where a number of words between a second position of the selected keyword in the target text and the at least one candidate keyword is less than or equal to a first threshold;
[0015] A second acquisition module, configured to obtain an association graph based on the at least one candidate keyword and the selected keyword; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and a weight of the edge between the two nodes is the relevance between the two nodes;
[0016] A third acquisition module, configured to obtain a word importance value corresponding to each of the at least one candidate keyword based on the association graph;
[0017] A screening module, configured to obtain a target keyword from the at least one candidate keyword based on the word importance value corresponding to each of the at least one candidate keyword.
[0018] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0019] A memory, configured to store a program;
[0020] A processor, configured to execute the program, and the program is specifically configured to:
[0021] Acquire a plurality of keywords included in a target text, where the target text includes a selected keyword in a selected state;
[0022] Based on first positions of the plurality of keywords in the target text respectively, determine at least one candidate keyword from the plurality of keywords, where a number of words between a second position of the selected keyword in the target text and the at least one candidate keyword is less than or equal to a first threshold;
[0023] Based on the at least one candidate keyword and the selected keyword, obtain an association graph; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and a weight of the edge between the two nodes is the relevance between the two nodes;
[0024] Based on the associated graph, obtain the word importance values corresponding to at least one candidate keyword respectively;
[0025] Based on the word importance values corresponding to the at least one candidate keyword respectively, obtain a target keyword from the at least one candidate keyword.
[0026] According to the fourth method of the embodiments of the present disclosure, there is provided a storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the keyword acquisition method shown in any one of the first aspect is implemented.
[0027] As can be seen from the above technical solutions, in the keyword acquisition method provided by the embodiments of the present application, if a search operation for a selected keyword included in the target text is detected, it indicates that during the process of browsing the target text, the user needs to view the text related to the selected keyword. It can be understood that since the user conducts the search during the process of browsing the target text, the text related to the selected keyword that the user wants to view has a certain relevance to the target text. Therefore, the embodiments of the present application provide a method for obtaining a target keyword based on an associated graph. The nodes included in the associated graph are: at least one candidate keyword and the selected keyword, and the number of words between the position of the candidate keyword in the target text and the position of the selected keyword in the target text is less than or equal to a first threshold; it can be understood that since the number of words between the position of the candidate keyword in the target text and the position of the selected keyword in the target text is less than or equal to the first threshold, the correlation between the candidate keyword and the selected keyword is relatively strong; if the relevance between any two nodes included in the associated graph is greater than a corresponding threshold, then there is an edge between these two nodes. Therefore, the relevance of the two keywords with an edge is relatively high. So the word importance value of the candidate keyword obtained based on the associated graph can represent the correlation with the selected keyword and the importance degree for the target text; the target keyword obtained from the at least one candidate keyword based on the word importance values corresponding to the at least one candidate keyword respectively has a relatively strong correlation with the selected keyword and a relatively high importance degree for the target text. Therefore, using the target keyword and the selected keyword together as search terms, the obtained search results are more in line with the user's needs, that is, the search results are relatively accurate. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0029] Figures 1a to 1c Schematic diagram of the fingertip search application scenario provided by the embodiment of the present application;
[0030] Figure 2 Architecture diagram of the implementation environment involved in a keyword acquisition method provided by the embodiment of the present application;
[0031] Figure 3 Flowchart of a keyword acquisition method provided by the embodiment of the present application;
[0032] Figure 4 Schematic diagram of an association diagram provided by the embodiment of the present application;
[0033] Figure 5 Structure of a keyword acquisition device provided by the embodiment of the present application;
[0034] Figure 6 Block diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0036] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and to perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0037] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0038] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.
[0039] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0040] The solution provided in the embodiments of this application involves technologies such as natural language processing of artificial intelligence, and will be specifically described through the following embodiments.
[0041] The embodiments of this application provide a keyword acquisition method, device, electronic device, and storage medium. Before introducing the technical solution provided in the embodiments of this application in detail, the application scenarios and implementation environments involved in the embodiments of this application will be briefly introduced here.
[0042] First, the application scenarios involved in the embodiments of this application will be briefly introduced.
[0043] The embodiments of this application can be applied to the fingertip search application scenario. In the fingertip search application scenario, when the user browses the text through the electronic device, the user can select one or more consecutive characters in the text displayed by the electronic device. In the embodiments of this application, "one or more consecutive characters" are called selected keywords, and then a search operation is performed. The above application scenario will be illustrated by examples below.
[0044] As Figures 1a to 1c shown, it is a schematic diagram of the fingertip search application scenario provided by the embodiments of this application.
[0045] The user can browse the text through the electronic device. Exemplarily, the electronic device can be any electronic product that can perform human-computer interaction with the user in one or more ways such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device. For example, a mobile phone, a tablet computer, a handheld computer, a personal computer, a wearable device, a smart TV, etc.
[0046] Figures 1a to 1cThe electronic device is described as a mobile phone as an example.
[0047] exist Figure 1a In the example, a user browses a text on a mobile phone. Assume that the text is news about "B". If the user wants to know other news about "B", he can select "B" in the text. Figure 1a As shown, "B" 11 is the selected keyword 11. If the user performs a corresponding operation, for example, long-pressing the selected keyword 11, a search prompt box 12 will be displayed, for example, Figure 1a A search prompt box 12 is located above the selected keyword 11.
[0048] exist Figure 1b In the example, if the user needs to view news related to the selected keyword 11, the user can click "Search" in the search prompt box 12 to obtain other news related to the selected keyword 11, such as Figure 1c As shown, the selected keyword 11 is displayed as a search term in the search box 13, and multiple other news about the selected keyword 11 are displayed in the display interface of the electronic device, such as news that B officially released a statement, B officially responded, B changed the situation and further postponed the business and will not be sold for the time being, etc.
[0049] Figures 1a to 1c In the example of news as the text type, the fingertip search application scenario involved in the present application is introduced. The present application embodiment is not limited to the text type of the text.
[0050] Exemplarily, the type of text may be any one of news, microblog, blog, encyclopedia, and article.
[0051] When the user determines the selected keyword, he can select one or more consecutive characters. For example, if the user wants to know the merger information about "A" and "B", the user needs to select the two keywords "A" and "B". Figure 1a The positions of "A" and "B" in the file shown are not continuous, so the user needs to select a sentence containing both "A" and "B" as the selected keyword, for example Figure 1a In the example, "A and B" are determined as a whole as the selected keywords. It is not possible to "jump and select" multiple words. For example, the user cannot skip other words, such as "and", and select only "A" and "B" without selecting "and".
[0052] In summary, if the user needs to select multiple words as selected keywords, it is impossible to achieve "jump selection" of multiple words, but can only select a sentence containing the above multiple words as the selected keyword. Since the sentence also includes other words, the search results obtained cannot meet the user's needs, that is, the search results are inaccurate.
[0053] In view of this, the present application provides a keyword acquisition method, which can determine target keywords that a user may need based on selected keywords selected by the user, and use the target keywords and the selected keywords together as search terms, which can improve the accuracy of search results. That is, without the user having to select the target keywords themselves, the keyword acquisition method provided by the embodiments of the present application automatically selects the target keywords for the user.
[0054] Secondly, a brief introduction to the implementation environment involved in the embodiments of the present application is given.
[0055] Figure 2 It is an architecture diagram of the implementation environment involved in a keyword acquisition method provided by an embodiment of the present application. This embodiment environment includes: a server 21 and at least one terminal device 22.
[0056] Exemplarily, the terminal device 22 and the server 21 can establish a connection and communicate through a wireless network.
[0057] Exemplarily, the terminal device 22 can be any kind of electronic product that can perform human-computer interaction with a user through one or more means such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a mobile phone, a tablet computer, a handheld computer, a personal computer, a wearable device, a smart TV, etc.
[0058] Exemplarily, a client is running on the terminal device 22, and the user can browse text based on this client. If the client is an application client, then the terminal device 22 can install this client; if the client is a web version client, then the terminal device 22 can display the web version client through a browser.
[0059] Exemplarily, the server 21 can be a single server, a server cluster composed of multiple servers, or a cloud computing service center. The server 21 can include a processor, a memory, and a network interface, etc.
[0060] Exemplarily, the database stores multiple texts, and texts belonging to the same text type are stored in the same text collection. Exemplarily, texts belonging to different text types are stored in different file collections; that is, the database stores multiple texts in partitions based on the text types of the multiple texts.
[0061] Exemplarily, texts belonging to different text types are stored in the same file collection.
[0062] Exemplarily, the server 21 obtains one or more texts from the database and sends them to the terminal device 22. The terminal device 22 can display one or more texts.
[0063] Exemplarily, the above database may be independent of the server 21 or located in the server 21.
[0064] Figure 2 Merely as an example, Figure 2 One terminal device 22 is shown. In actual applications, the number of terminal devices 22 can be set according to actual needs, and the embodiments of the present disclosure do not limit the number of terminal devices 22.
[0065] In an optional implementation manner, the terminal device 22 is used to display the target text and obtain a search operation implemented on the selected keywords included in the target text. The server 21 is used to, if receiving the search operation implemented on the selected keywords included in the target text sent by the terminal device 22, obtain the target keyword from the multiple keywords included in the target text based on the selected keyword. The server 21 is further used to use the target keyword and the selected keyword as search terms to obtain a search result and send the search result to the terminal device 22. Exemplarily, the terminal device 22 may display the search result, as Figure 1c shown.
[0066] In an optional implementation manner, the terminal device 22 is used to display the target text and detect a search operation implemented on the selected keywords included in the target text; if the terminal device 22 detects the search operation implemented on the selected keywords included in the target text, it obtains the multiple keywords included in the target text from the multiple keywords respectively included in the multiple texts stored in the database by the server 21, determines the target keyword from the multiple keywords based on the selected keyword, and sends the target keyword to the server 21; Exemplarily, the server 21 is further used to use the target keyword and the selected keyword as search terms to obtain a search result and send the search result to the terminal device 22. Exemplarily, the terminal device 22 may display the search result, as Figure 1c shown.
[0067] Exemplarily, obtaining the target keyword from the multiple keywords included in the target text based on the selected keyword involves natural language processing technology of artificial intelligence.
[0068] Those skilled in the art should understand that the above electronic devices and servers are only examples. Other existing or future possible electronic devices or servers that can be applied to the present disclosure should also be included within the protection scope of the present disclosure and are hereby incorporated herein by reference.
[0069] The technical solution provided by the present application will be introduced below with reference to the accompanying drawings.
[0070] Figure 3 is a flowchart of a keyword acquisition method provided by an embodiment of the present application. This method can be applied to Figure 2The server 21 or the terminal device 22 in the shown implementation environment. In the process of implementing this method, the following steps S301 to S305 are included.
[0071] In step S301, if a search operation implemented on the selected keyword included in the target text is detected, multiple keywords included in the target text are obtained.
[0072] If Figure 3 The shown method is applied to the server, then the "search operation implemented on the selected keyword included in the target text" is received from the terminal device 22.
[0073] If Figure 3 The shown method is applied to the terminal device 22, then the "search operation implemented on the selected keyword included in the target text" is detected by the terminal device 22 itself.
[0074] Exemplarily, if Figure 3 The shown method is applied to the terminal device 22, the implementation manners of "detecting the search operation implemented on the selected keyword included in the target text" include but are not limited to the following two implementation manners.
[0075] The first implementation manner: If a preset button pressing operation is detected, it is determined that the search operation implemented on the selected keyword included in the target text is detected.
[0076] Exemplarily, the preset button can be Figure 1b The shown "Search" button.
[0077] The second implementation manner: If a preset voice is detected, it is determined that the search operation implemented on the selected keyword included in the target text is detected.
[0078] Exemplarily, the preset voice includes the selected keyword. For example, the "Search for B" voice.
[0079] If Figure 3 The shown method is applied to the terminal device 22, exemplarily, the display manners of the selected keyword in the target text include but are not limited to the following four manners.
[0080] The first display manner: The selected keyword is displayed in the target text in a "flashing" manner.
[0081] The second display manner: The selected keyword is displayed in the target text in a preset font color.
[0082] For example, if the font in the target text is black, the selected keyword is displayed in the target text in red font.
[0083] Exemplarily, in the second display mode, the present application does not limit the specific color of the preset font color, and any font color that can distinguish the target text and the selected keyword is within the protection scope of the embodiments of the present application.
[0084] The third display mode: The selected keyword is displayed in the target text in a preset font format.
[0085] For example, the preset font format includes "bold, bolded".
[0086] Exemplarily, in the third display mode, the "preset font format" can be one of "bold, bolded", "italic", or "underline". It can be understood that the present application is not limited to the specific font format of the preset font format, and any font format that can distinguish the target text and the selected keyword is within the protection scope of the embodiments of the present application.
[0087] The fourth display mode: The selected keyword is displayed in a manner covered by a selection window.
[0088] Exemplarily, Figure 1a The fourth display mode is taken as an example for illustration.
[0089] Exemplarily, the multiple keywords included in the target text include the selected keyword, or the multiple keywords included in the target text do not include the selected keyword.
[0090] In an optional implementation manner, the multiple keywords included in the target text have been determined before step S301 and are stored in the server 21 or the database or the terminal device 22. When step S301 is executed, they can be obtained from the server 21 or the database or the terminal device 22.
[0091] It can be understood that the server 21 can pre-store the multiple keywords included in multiple texts respectively and store them. When step S301 is executed, the multiple keywords included in the target text can be obtained from the multiple keywords included in each text that have been obtained.
[0092] In an optional implementation manner, the multiple keywords included in the target text are determined in real time after a search operation implemented on the selected keyword included in the target text is detected.
[0093] In step S302, based on the first positions of the multiple keywords in the target text respectively, at least one candidate keyword whose number of intervening words from the second position of the selected keyword in the target text is less than or equal to the first threshold is determined from the multiple keywords.
[0094] It can be understood that the keyword or selected keyword may appear more than once in the target text. Therefore, for each keyword, one or more first positions may be determined in the target text, and for the selected keyword, one or more second positions may be determined in the target text.
[0095] There are various implementation manners for step S302, and the embodiments of the present application provide but are not limited to the following three.
[0096] The first implementation manner of step S302 includes steps A11 to A12.
[0097] In step A11, for each second position corresponding to the selected keyword, one or more words in the target text whose number of intervening words from the second position is less than or equal to the first threshold are determined, so as to obtain the words corresponding to each of at least one second position.
[0098] In step A12, the intersection of the words corresponding to each of at least one second position determined in step A11 and the multiple keywords included in the target text determined in step S301 is determined, so as to obtain one or more candidate keywords.
[0099] The second implementation manner of step S302 includes steps A21 to A22.
[0100] In step A21, for each first position of each keyword, one or more words in the target text whose number of intervening words from the first position is less than or equal to the first threshold are determined, so as to obtain the words corresponding to the keyword.
[0101] In step A22, it is determined whether the words determined in step A21 include the selected keyword. If so, the keyword is determined as a candidate keyword.
[0102] The third implementation manner of step S302 includes steps A31 to A32.
[0103] In step A31, for each keyword, the closest first position and second position are determined from at least one first position corresponding to the keyword and at least one second position corresponding to the selected keyword, so as to obtain a keyword combination corresponding to the keyword, and keyword combinations corresponding to multiple keywords are obtained.
[0104] In step A32, for the keyword combination of each keyword, the number of words between the first position and the second position included in the keyword combination is determined. If the number of words is less than or equal to the first threshold, the keyword is determined as a candidate keyword, so as to obtain one or more candidate keywords.
[0105] Exemplarily, the above "number of words" does not include stop words, for example, "de", "le"; exemplarily, the above "number of words" includes stop words.
[0106] The following uses specific examples to illustrate the three implementation processes of step S302, assuming that the above "number of words" does not include stop words.
[0107] Assume that the target text is: "Programmers (English: Programmer) are professionals engaged in program development and maintenance. Generally, programmers are divided into program designers and program coders, but the boundary between the two is not very clear, especially in China. Software practitioners are divided into four categories: junior programmers, senior programmers, system analysts, and project managers."
[0108] After segmenting the target text and removing the stop words in the target text, the multiple words in the resulting vocabulary set are: programmer, English, program, development, maintenance, professional, personnel, programmer, divided into, program, design, personnel, program, coding, personnel, boundary, especially, China, software, personnel, divided into, programmer, senior, programmer, system, analyst, project, manager.
[0109] In summary, the vocabulary set of the target text includes 28 words, and each word corresponds to a position in the target text. Assume that the positions of the 28 words in the target text are successively position 1, position 2,..., position 28.
[0110] Assume that the selected keyword is "program", and "program" has three second positions in the target text, namely position 3, position 10, and position 13.
[0111] The keyword "programmer" has four first positions in the target text, namely position 1, position 8, position 22, and position 24.
[0112] The following takes the keyword "programmer" and the selected keyword "program" as examples to illustrate the above three implementation methods.
[0113] In the first implementation method of step S302, for each second position of the selected keyword "program", for example, position 3, position 10, and position 13, determine each word in the target text whose number of words separated from the second position is less than or equal to the first threshold.
[0114] Assume that the first threshold is 4, and the words in the target text whose number of words separated from position 3 is less than or equal to the first threshold are respectively: {programmer, English, development, maintenance, professional, personnel}.
[0115] Each of the words in the target text whose number of intervening words from position 10 is less than or equal to the first threshold is: {professional, personnel, programmer, divided into, design, personnel, program, coding}.
[0116] Each of the words in the target text whose number of intervening words from position 13 is less than or equal to the first threshold is: {divided into, program, design, personnel, coding, personnel, boundary, special}.
[0117] Then, for each of the at least one second position where the selected keyword determined in step A11 is "program", the corresponding words are {programmer, English, develop, maintain, professional, divided into, design, personnel, program, coding, boundary, special}.
[0118] Exemplarily, the intersection of {programmer, English, develop, maintain, professional, divided into, design, personnel, program, coding, boundary, special} and the multiple keywords included in the target text is the candidate keyword.
[0119] In the second implementation manner of step S302, for each first position of the keyword "programmer", for example, position 1, position 8, position 22, or position 24, one or more words in the target text whose number of intervening words from position 1 is less than or equal to the first threshold are {English, program, develop, maintain}; one or more words in the target text whose number of intervening words from position 8 is less than or equal to the first threshold are {develop, maintain, professional, personnel, divided into, program, design, personnel}; one or more words in the target text whose number of intervening words from position 22 is less than or equal to the first threshold are {China, software, personnel, divided into, senior, programmer, system, analyst}; one or more words in the target text whose number of intervening words from position 24 is less than or equal to the first threshold are {personnel, divided into, programmer, senior, system, analyst, project, manager}. Then, the corresponding words for the keyword "programmer" are {English, program, develop, maintain, professional, personnel, divided into, design, China, software, senior, programmer, system, analyst, project, manager}. Since {English, program, develop, maintain, professional, personnel, divided into, design, China, software, senior, programmer, system, analyst, project, manager} includes the selected keyword "program", the keyword "programmer" is a candidate keyword.
[0120] In the third implementation manner of step S302, for the keyword "programmer" and the selected keyword "program", from {position 1, position 8, position 22, position 24} and {position 3, position 10, position 13}, determine the closest positions, which are position 1 and position 3, or position 8 and position 10.
[0121] Since the number of words between position 1 and position 3 or between position 8 and position 10 is 1, which is less than the first threshold, the keyword "programmer" is a candidate keyword.
[0122] In step S303, an association graph is obtained based on the at least one candidate keyword and the selected keyword.
[0123] The at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to the corresponding threshold, and the weight of the edge between the two nodes is the relevance between the two nodes.
[0124] It can be understood that for any two nodes included in the association graph, if the relevance between the two nodes is greater than the corresponding threshold, there is an edge between the two nodes; otherwise, there is no edge between the two nodes.
[0125] Exemplarily, two nodes are called a set of node sets, and the thresholds corresponding to different node sets are different; exemplarily, the thresholds corresponding to different node sets are the same.
[0126] Exemplarily, the association graph can be a weighted undirected graph or a weighted directed graph.
[0127] Exemplarily, if the association graph is a directed graph, then the determination method of the direction of the edge between any two nodes with an edge in the association graph includes but is not limited to the following two.
[0128] The first determination method of the direction of the edge between any two nodes in the association graph includes: randomly determining the direction.
[0129] For example, if the two nodes are node A and node B, then the direction of the edge between node A and node B can be from node A to node B, or from node B to node A, or from node A to node B and at the same time from node B to node A.
[0130] The second determination method of the direction of the edge between any two nodes in the association graph includes: determining the direction of the edge between the two nodes based on the positions of the two nodes in the target text.
[0131] In the second implementation manner, the embodiments of the present application provide but are not limited to the following two implementation manners.
[0132] The first: Among any two nodes (the two nodes can both be candidate keywords, or one of the two nodes is a candidate keyword and the other is a selected keyword), the node with a forward position in the target text points to the node with a backward position in the target text; or, the node with a backward position in the target text points to the node with a forward position in the target text.
[0133] Taking "a node located earlier in the target text points to a node located later in the target text" as an example for illustration.
[0134] Exemplarily, since there may be multiple first positions of each candidate keyword in the target text, the direction of the edge between any two candidate keywords may be bidirectional or unidirectional.
[0135] Still taking the above as an example, that is, the multiple words in the vocabulary set of the target text are: programmer, English, program, development, maintenance, professional, personnel, programmer, divided into, program, design, personnel, program, coding, personnel, boundary, special, China, software, personnel, divided into, programmer, senior programmer, system analyst, project manager.
[0136] Assume that "programmer", "coding", and "divided into" are candidate keywords. Assume that the relevance pair of "programmer" and "coding" is greater than or equal to the corresponding threshold. Then, there is an edge between "programmer" and "coding". Since the first positions of "programmer" in the target text are: position 1, position 8, position 22, position 24, and the first position of "coding" in the target text is: position 14. That is, "programmer" appears in the target text before the candidate keyword "coding" appears, and "programmer" also appears after the candidate keyword "coding" appears. Therefore, the direction of the edge between "programmer" and "coding" is bidirectional, that is, it points from "programmer" to "coding" and from "coding" to "programmer".
[0137] Exemplarily, since there may be multiple second positions of the selected keyword in the target text and multiple first positions of each candidate keyword in the target text, the direction of the edge between the selected keyword and the candidate keyword may be bidirectional or unidirectional.
[0138] Assume that "programmer" is a candidate keyword and "program" is a selected keyword. The first positions of "programmer" in the target text are: position 1, position 8, position 22, position 24; the second positions of "program" in the target text are: position 3, position 10, position 13. Therefore, the position of "programmer" in the target text is before "program" and after "program". Therefore, the direction of the edge between "programmer" and "program" is bidirectional, that is, it points from "programmer" to "program" and from "program" to "programmer".
[0139] Second type: There may be multiple second positions of the selected keyword in the target text, and there may be multiple first positions of each candidate keyword in the target text. However, in the process of determining the candidate keyword from at least one keyword, it may be determined based on the closest positions of the selected keyword and the candidate keyword in the target text. For example, in the implementation manner of the third step S302, so if there is an edge between the candidate keyword and the selected keyword, the direction of the edge needs to be determined based on the front-back relationship of the positions in the keyword combination corresponding to the candidate keyword.
[0140] The following takes "the node located in the front position in the target text points to the node located in the back position in the target text" as an example for illustration.
[0141] Suppose "programmer" is the candidate keyword and "program" is the selected keyword. Since "programmer" is determined as the candidate keyword based on the keyword combination {position 1, position 3} or the keyword combination {position 8, position 10}, in both of these keyword combinations, "programmer" is in the front position in the target text and "program" is in the back position in the target text. Therefore, the direction of the edge between "programmer" and "program" is one-way and points from "programmer" to "program".
[0142] Exemplarily, the relevance between two nodes can be represented by any one of the cosine similarity, Euclidean distance, Mahalanobis distance, Manhattan distance, or Hamming distance between the two nodes.
[0143] It can be understood that since the number of words between the position of the candidate keyword in the target text and the position of the selected keyword in the target text is less than or equal to the first threshold, the relevance between the candidate keyword and the selected keyword is relatively strong.
[0144] Relevance is a non-deterministic interdependent relationship existing in objective phenomena. The relevance between the candidate keyword and the selected keyword means that after the user selects the selected keyword, the probability of implicitly selecting the candidate keyword. Here, "implicit selection" means that the user actually does not select the candidate keyword, but after the electronic device or server "thinks" that the user has selected the "selected keyword", the candidate keyword may be what the user intends to select.
[0145] In summary, the relevance between each candidate keyword included in the association graph and the selected keyword is relatively strong.
[0146] In step S304, based on the association graph, obtain the word importance values corresponding to at least one candidate keyword respectively.
[0147] There are multiple implementation manners for step S304. The embodiments of the present application provide but are not limited to the following two implementation manners.
[0148] The first implementation method of step S304 includes, for each node: obtaining the word importance value of the node based on the word vectors respectively corresponding to at least one node connected to the node and the weights of the edges respectively corresponding to at least one node connected to the node.
[0149] Exemplarily, assume that the nodes connected to node 1 are respectively: node 2, node 3, and node 4; the word importance degree value of node 1 = the word vector of node 2 * the weight of the edge connecting node 1 and node 2 + the word vector of node 3 * the weight of the edge connecting node 1 and node 3 + the word vector of node 4 * the weight of the edge connecting node 1 and node 4.
[0150] Exemplarily, the word vector of a node can be obtained through the Word2Vec model.
[0151] The second implementation method of step S304 includes, for each node: based on the word vectors respectively corresponding to at least one first node connected to the node, the word vectors respectively corresponding to at least one second node connected to each of the first nodes, the word vectors respectively corresponding to at least one third node connected to each of the second nodes,..., the word vectors respectively corresponding to at least one Mth node connected to each of the M - 1th nodes, the word vectors respectively corresponding to at least one leaf node connected to each of the Mth nodes, the weights of the edges between the node and each of the first nodes connected to it, the weights of the edges between each of the first nodes and each of the second nodes connected to it,..., the weights of the edges between each of the M - 1th nodes and each of the Mth nodes connected to it, the weights of the edges between each of the Mth nodes and each of the leaf nodes connected to it, obtaining the word importance value of the node.
[0152] Where M is a positive integer greater than or equal to 2.
[0153] Exemplarily, the word importance value of a node can be obtained through the formula , and the word importance value of the node is obtained.
[0154] Where refers to the word importance value of node , refers to the word importance value of node , d is the damping coefficient, generally set to 0.85, is the set of predecessor nodes of node in each node included in the associated graph, is the set of successor nodes of node in each node included in the associated graph, refers to the relevance between node and node , refers to node and the node relevance.
[0155] In an alternative embodiment, if the associated graph is an undirected graph, the predecessor node of a node refers to the node among the nodes included in the associated graph whose position in the target text is in front of the node ; the successor node of a node refers to the node among the nodes included in the associated graph whose position in the target text is behind the node .
[0156] If the position of a node in the target text is in front of the node and behind the node , then this node is both the predecessor node of the node and the successor node of the node .
[0157] In an alternative embodiment, if the associated graph is a directed graph, the predecessor node of a node refers to the node among the nodes included in the associated graph that points to the node , and the successor node of a node refers to the node among the nodes included in the associated graph that the node points to.
[0158] In step S305, based on the word importance values respectively corresponding to the at least one candidate keyword, a target keyword is obtained from the at least one candidate keyword.
[0159] In the keyword acquisition method provided by the embodiments of the present application, if a search operation for a selected keyword included in the target text is detected, it indicates that during the process of browsing the target text, the user needs to view the text related to the selected keyword. It can be understood that since the user conducts the search during the process of browsing the target text, the text that the user wants to view related to the selected keyword has a certain relevance to the target text. Therefore, the embodiments of the present application provide a method for obtaining target keywords based on an association graph. The nodes included in the association graph are: at least one candidate keyword and the selected keyword, and the number of words between the position of the candidate keyword in the target text and the position of the selected keyword in the target text is less than or equal to the first threshold; it can be understood that since the number of words between the position of the candidate keyword in the target text and the position of the selected keyword in the target text is less than or equal to the first threshold, the correlation between the candidate keyword and the selected keyword is relatively strong; if the correlation degree between any two nodes included in the association graph is greater than the corresponding threshold, then there is an edge between these two nodes. Therefore, the correlation degree between the two keywords with an edge is relatively high. So, the word importance value of the candidate keyword obtained based on the association graph can represent the correlation relationship with the selected keyword and the importance degree for the target text; based on the word importance values respectively corresponding to the at least one candidate keyword, the target keyword obtained from the at least one candidate keyword has a relatively strong correlation with the selected keyword and a relatively high importance degree for the target text. Therefore, using the target keyword and the selected keyword as the search terms together, the obtained search results are more in line with the user's needs, that is, the search results are relatively accurate.
[0160] In an alternative embodiment, the specific implementation process of step S303 includes steps B1 to B3.
[0161] In step B1, obtain the first correlation degree between each of the at least one candidate keyword and the selected keyword.
[0162] Exemplarily, the first correlation degree is any one of cosine similarity, Euclidean distance, Mahalanobis distance, Manhattan distance, or Hamming distance.
[0163] Exemplarily, the first correlation degree between the candidate keyword and the selected keyword can be obtained based on the word vector of the candidate keyword and the word vector of the selected keyword.
[0164] Exemplarily, input the candidate keyword into the Word2Vec model to obtain the word vector of the candidate keyword; input the selected keyword into the Word2Vec model to obtain the word vector of the selected keyword.
[0165] The Word2vec model is a word vector calculation model. The Word2vec model is a shallow two-layer neural network, and the word2vec model can be used to map each keyword to a vector.
[0166] Exemplarily, input the candidate keywords into the bert pre-trained model to obtain the word vectors of the candidate keywords; input the selected keywords into the bert pre-trained model to obtain the word vectors of the selected keywords.
[0167] The training process of the Word2Vec model is described below. The training process of the Word2Vec model includes steps C1 to C2.
[0168] In step C1, obtain the target domain to which the target text belongs.
[0169] The domain to which the text belongs can be divided based on the text content. For example, the domain to which the text belongs can be at least one of the entertainment domain and the science domain.
[0170] In step C2, obtain each text belonging to the target domain, and perform word segmentation on each text belonging to the target domain to obtain a word segmentation set corresponding to each text belonging to the target domain.
[0171] Exemplarily, the word segmentation set corresponding to a text may not include stop words.
[0172] In step C3, use the word segmentation sets corresponding to each text belonging to the target domain to train to obtain the Word2Vec model.
[0173] Exemplarily, since the training samples of the Word2Vec model are each text in the target domain to which the target text belongs, for target texts belonging to different domains, the Word2Vec models corresponding to the target texts are different, specifically reflected in the different model parameters of the Word2Vec models.
[0174] Exemplarily, the Word2Vec model can be pre-trained based on the texts in different domains to obtain the Word2Vec models corresponding to each domain.
[0175] Among them, the dimension of the word vector obtained based on the Word2Vec model can be preset. For example, the dimension of the word vector is 200 or 300 or 400. This application does not limit the dimension of the word vector.
[0176] In step B2, for each of the candidate keywords, if the first relevance between the candidate keyword and the selected keyword is greater than or equal to the second threshold, an edge between the candidate keyword and the selected keyword is constructed, and the first relevance is determined as the weight of the edge between the candidate keyword and the selected keyword.
[0177] Exemplarily, the first relevance is any one of cosine similarity, Euclidean distance, Mahalanobis distance, Manhattan distance, or Hamming distance.
[0178] Exemplarily, if the first relevance is cosine similarity, the second threshold is between 0 and 1. For example, the second threshold is 0.6.
[0179] In step B3, if the at least one candidate keyword includes at least two candidate keywords, for any two candidate keywords, if the second relevance between the two candidate keywords is greater than or equal to the third threshold, an edge between the two candidate keywords is constructed, and the second relevance is determined as the weight of the edge between the two candidate keywords.
[0180] Exemplarily, the second relevance is any one of cosine similarity, Euclidean distance, Mahalanobis distance, Manhattan distance, or Hamming distance.
[0181] Exemplarily, if the second relevance is cosine similarity, the third threshold is between 0 and 1. For example, the third threshold is 0.7.
[0182] Exemplarily, the second threshold may be equal to or different from the third threshold in magnitude.
[0183] The construction process of the association graph will be described below by way of example. Figure 4 This is a schematic diagram of an association graph provided by an embodiment of the present application.
[0184] Figure 4 The shown association graph includes five nodes: node A, node B, node C, node D, and node E, where node A, node B, node C, and node D are candidate keywords, and node E is the selected keyword.
[0185] Since the first relevance between node E and node B and between node E and node D is greater than or equal to the second threshold, there is an edge between node E and node B, and there is an edge between node E and node D; since the first relevance between node E and node A and between node E and node C is less than the second threshold, there is no edge between node E and node A, and there is no edge between node E and node C.
[0186] Since the second relevance degrees of node B with respect to node C and node A are greater than or equal to the third threshold, there is an edge between node B and node C, and there is an edge between node B and node A. Since the second relevance degree of node B with respect to node D is less than the third threshold, there is no edge between node B and node D.
[0187] Since the second relevance degree of node D with respect to node A is greater than or equal to the third threshold, there is an edge between node D and node A. Since the second relevance degrees of node D with respect to node C and node B are less than the third threshold, there is no edge between node D and node C, and there is no edge between node D and node B.
[0188] The same applies to node C and will not be elaborated here.
[0189] In the embodiment of the present application, the number of words between each candidate keyword included in the association graph and the selected keyword is less than or equal to the first threshold. Therefore, the correlation between the candidate keyword and the selected keyword is strong, that is, the correlation between each node in the association graph and the selected keyword is strong. That is, if the user selects the selected candidate word, there is a high probability of selecting the candidate keyword. The edge between any two nodes in the association graph is established when the relevance degree between the two nodes is greater than the corresponding threshold. Therefore, the edge between the two nodes represents a strong transitivity between the two nodes. That is, if the user selects one of the nodes, the probability of implicitly selecting the other node connected to the node is relatively high.
[0190] Therefore, the association graph determined in the embodiment of the present application is a node network graph that represents a strong correlation with the selected keyword and has strong transitivity.
[0191] There are multiple implementation processes for step S305. The embodiment of the present application provides but is not limited to the following two implementation manners.
[0192] The first implementation manner of step S305 includes: sorting at least one candidate keyword in descending order of the word importance value, and selecting the first N candidate keywords as the target keywords. N is a positive integer greater than or equal to 1.
[0193] The second implementation manner of step S305 includes steps D1 to D4.
[0194] In step D1, at least one target candidate keyword is obtained from the at least one candidate keyword based on the word importance values respectively corresponding to the at least one candidate keyword.
[0195] Exemplarily, the implementation manner of step D1 includes but is not limited to the following two implementation manners.
[0196] The first implementation method of step D1: Sort at least one candidate keyword in descending order according to the word importance value, and determine the first preset number of candidate keywords as the target candidate keywords.
[0197] Exemplarily, the first preset number is a positive integer greater than or equal to 1 and less than the number of candidate keywords. For example, if the number of candidate keywords is 10, the first preset number is less than or equal to 10.
[0198] The second implementation method of step D1: Obtain at least one target candidate keyword with a word importance value greater than or equal to the fourth threshold from at least one candidate keyword.
[0199] Exemplarily, the fourth threshold can be determined based on the actual situation and is not limited here.
[0200] In step D2, obtain multiple query records included in the query log within a preset time period. The query record includes the selected keyword and at least one of the target candidate keywords.
[0201] Exemplarily, one query record corresponds to one search term. One search term includes one or more keywords.
[0202] The following is an example to illustrate the search term. When the user enters "notebook paper" in the input box of the user interface displayed by the browser, then the search term includes each keyword entered by the user in the input box, that is, the search term includes "notebook" and "paper".
[0203] Exemplarily, the query log includes query records corresponding to each user within a preset time period.
[0204] Exemplarily, the preset time period ends at the first moment of "detecting a search operation on the selected keyword included in the target text".
[0205] Exemplarily, the duration of the preset time period can be L hours or G days, where L is any value greater than 0 and G is any value greater than 0.
[0206] The following is an example to illustrate step D2. Assume that the preset time period is Q days.
[0207] Assume that there are a total of 1 million query records in Q days. Obtain the query records including the selected keyword and at least one target candidate keyword from the 1 million query records. Assume that there are 800,000 query records including the selected keyword and at least one target candidate keyword. Then what step D2 obtains is these 800,000 query records.
[0208] In step D3, for each of the target candidate keywords, determine the first number of query records in the multiple query records that contain the target candidate keyword, so as to obtain the first numbers respectively corresponding to the at least one target candidate keyword.
[0209] Exemplarily, the first number characterizes the number of times that the target candidate keyword and the selected keyword appear simultaneously as "search terms" in the query log, that is, the co-occurrence times.
[0210] In step D4, based on the first numbers respectively corresponding to the at least one target candidate keyword, obtain the target keyword from the at least one target candidate keyword.
[0211] Exemplarily, the implementation manner of step D4 includes but is not limited to the following three implementation manners.
[0212] The first implementation manner of step D4: For the first number corresponding to each target candidate keyword, if the first number is greater than or equal to the fifth threshold, determine that the target candidate keyword is the target keyword.
[0213] Exemplarily, the fifth threshold can be set based on the actual situation and is not limited here.
[0214] The second implementation manner of step D4: Based on the first number corresponding to each target candidate keyword and the sum of the first numbers respectively corresponding to all target candidate keywords, determine the co-occurrence frequency corresponding to each target candidate keyword. If the co-occurrence frequency corresponding to a target candidate keyword is greater than or equal to the sixth threshold, determine that the target candidate keyword is the target keyword.
[0215] Exemplarily, the sixth threshold can be set based on the actual situation and is not limited here.
[0216] Exemplarily, the co-occurrence frequency corresponding to a target candidate keyword = the first number corresponding to the target candidate keyword / the sum of the first numbers respectively corresponding to the at least one target candidate keyword.
[0217] The third implementation manner of step D4: Based on the first numbers respectively corresponding to the at least one target candidate keyword and the sum of the first numbers respectively corresponding to the at least one target candidate keyword, determine the co-occurrence frequencies respectively corresponding to the at least one target candidate keyword; based on the co-occurrence frequencies respectively corresponding to the at least one target candidate keyword, obtain the target keyword from the at least one target candidate keyword.
[0218] Exemplarily, sort the at least one target candidate keyword in descending order according to the co-occurrence frequency, and determine the first second preset number of target candidate keywords as the target keyword.
[0219] The second preset number is any integer greater than or equal to 1.
[0220] Among them, the co-occurrence frequency corresponding to a target candidate keyword characterizes the frequency of the target candidate keyword and the selected keyword appearing simultaneously as "search terms" in multiple query records obtained in step D2.
[0221] Exemplarily, the co-occurrence frequency corresponding to a target candidate keyword = the first number corresponding to the target candidate keyword / the sum of the first numbers respectively corresponding to the at least one target candidate keyword.
[0222] It can be understood that the more times the same search term is input by each user, the more likely it is that the search term corresponds to a meaningful and relatively popular event. The greater the co-occurrence times or co-occurrence frequency corresponding to a target candidate keyword, the more likely it is that the target candidate keyword and the selected keyword correspond to a "popular" event within a preset time period. Therefore, the search text obtained based on the target keyword and the selected keyword can better meet the search intent of the user, that is, the search results obtained based on the target keyword and the selected keyword are more accurate than the search results obtained based only on the selected keyword.
[0223] In an optional implementation manner, there are multiple ways to obtain the "multiple keywords included in the target text" involved in step S301. The embodiments of the present application provide but are not limited to the following two.
[0224] The first implementation manner of obtaining multiple keywords included in the target text includes: performing word segmentation on the target text to obtain multiple keywords included in the target text.
[0225] Exemplarily, perform word segmentation on the target text and remove stop words to obtain multiple keywords, that is, the multiple keywords included in the target text do not include stop words.
[0226] For example, a word segmentation tool can be preset to perform word segmentation on the target text to obtain multiple words, and based on a preset stop word dictionary, remove the stop words in the multiple words to obtain multiple keywords included in the target text.
[0227] Exemplarily, the word segmentation tool can be any one of the postag part-of-speech tagging tool, the jieba word segmenter, the Paoding word segmentation tool, and the IK word segmentation tool.
[0228] Exemplarily, the four word segmentation tools provided by the embodiments of the present application, other existing or future possible word segmentation tools that can be applied to the present application should also be included in the protection scope of the present application and are hereby incorporated herein by reference.
[0229] Exemplarily, the preset stop word dictionary includes multiple stop words, for example, meaningless words such as "de", "le", "ne", etc.
[0230] Exemplarily, the target text is segmented to obtain a plurality of keywords, that is, the plurality of keywords included in the target text include stop words.
[0231] The second implementation manner for obtaining the plurality of keywords included in the target text includes steps E1 to E3.
[0232] In step E1, a plurality of words included in the target text are obtained.
[0233] Exemplarily, the target text can be segmented based on a preset word segmentation tool, and the stop words in the target text are removed to obtain a plurality of words included in the target text. For the specific description of the first implementation manner for obtaining the plurality of keywords included in the target text, reference can be made, and details are not described herein again.
[0234] In step E2, for each of the words, based on the second number of times the word is included in the target text, the total number of words included in the target text, the third number of texts in the preset text set that include the word, the total number of each text included in the preset text set, and the total number of the plurality of words, a text importance value representing the importance of the word to the target text is obtained, so as to obtain the text importance values corresponding to the plurality of words respectively.
[0235] Exemplarily, based on the TF-IDF (term frequency–inverse document frequency) model, the text importance values corresponding to the respective words included in the target text can be determined.
[0236] Exemplarily, the TF value (Term Frequency) of any word represents the frequency of occurrence of the word in the target text.
[0237] Suppose the total number of words included in the target text is 100, and the second number of times the word "program" is included in the target text is 3. Then the word frequency of the word "program" in the target file = the second number / the total number of words = 3 / 100 = 0.03.
[0238] Exemplarily, based on the formula TF i =n i / m, where TF i represents the word frequency of word i, n i refers to the number of times word i appears in the target text (i.e., the second number), and m represents the total number of words included in the target text, such as 100 above.
[0239] Exemplarily, for the IDF value (inverse document frequency) of any vocabulary, it can be obtained by taking the ratio of the number of texts containing the vocabulary (i.e., the third number) in the preset text set where the target file is located to the total number of all texts in the preset text set, and then taking the logarithm of the obtained ratio.
[0240] Exemplarily, it can be based on the formula IDF i =lg(a i / b) to determine the IDF value of the vocabulary. Where IDF i represents the inverse document frequency of vocabulary i, a i represents the number of texts containing vocabulary i in the preset text set, and b represents the total number of all texts in the preset text set.
[0241] Exemplarily, all texts in the preset text set belong to the same text type. For example, all texts in the preset file set belong to news articles, or all texts in the preset file set belong to microblogs posted by users. Exemplarily, all texts in the preset text set can belong to different text types.
[0242] Exemplarily, based on the formula text importance value of vocabulary i = TF i *IDF i , determine the text importance value representing the importance degree of the target text corresponding to vocabulary i.
[0243] For example, if the word "program" appears in 1,000 texts in the preset text set, and the total number of texts in the preset text set is 10,000,000, its inverse document frequency is lg(10,000,000 / 1,000) = 4. The final TF-IDF value of "program" is 0.03 * 4 = 0.12.
[0244] In step E3, based on the text importance values respectively corresponding to the multiple vocabularies, obtain the multiple keywords from the multiple vocabularies.
[0245] Exemplarily, the implementation manner of step E3 of the present application includes but is not limited to the following two implementation manners.
[0246] The first implementation manner of step E3: Determine the first third preset number of keywords in descending order according to the text importance values respectively corresponding to the multiple vocabularies.
[0247] Exemplarily, the third preset number is a positive integer greater than or equal to 1 and less than or equal to the total number of the multiple vocabularies. For example, if the total number of the multiple vocabularies is 100, the second preset number is less than or equal to 100.
[0248] The second implementation manner of step E3: Obtain multiple keywords from multiple words whose text importance values are greater than or equal to the seventh threshold.
[0249] That is, the keyword is a word whose text importance value is greater than or equal to the seventh threshold.
[0250] Exemplarily, the seventh threshold can be determined based on the actual situation, and this application does not limit it.
[0251] In an optional implementation manner, if the keyword acquisition method is applied to the server 21, the method further includes: Obtain a search result based on the target keyword and the selected keyword. Send the search result to the terminal device 22, and the terminal device 22 displays the search result.
[0252] Exemplarily, the text type of the target text is the same as the text type of at least one text included in the search result.
[0253] For example, the text type of the target text is news, and the text type of the text included in the search result is also news; the text type of the target text is a Weibo post, and the text type of the text included in the search result is also a Weibo post.
[0254] Exemplarily, the text type of the target text may be different from the text type of the text included in the search result.
[0255] In an optional implementation manner, if the keyword acquisition method is applied to the terminal device 22, the method further includes: Send the target keyword and the selected keyword to the server 21; The server 21 obtains a search result based on the target keyword and the selected keyword. Send the search result to the terminal device 22, and the terminal device 22 displays the search result.
[0256] In an optional implementation manner, only the selected keyword in the search result displayed by the terminal device 22 is in the selected state.
[0257] Exemplarily, the display manner of the selected keyword in the search result can be any one of the above first display manner, second display manner, third display manner, and fourth display manner.
[0258] In an optional implementation manner, both the selected keyword and the target keyword in the search result displayed by the terminal device 22 are in the selected state.
[0259] Exemplarily, the display manner of the selected keyword in the search result can be any one of the above first display manner, second display manner, third display manner, and fourth display manner.
[0260] Exemplarily, the display manner of the target keyword in the search results may be any one of the above first display manner, second display manner, third display manner, and fourth display manner.
[0261] In the embodiments disclosed in the present application above, the method is described in detail. For the method of the present application, it can be implemented by means of various forms of devices. Therefore, the present application also discloses various devices, and specific embodiments are given below for detailed description.
[0262] In an optional embodiment, the embodiment of the present application provides a keyword acquisition device. As Figure 5 shown, it is a structural diagram of a keyword acquisition device provided by the embodiment of the present application.
[0263] The device includes: a first acquisition module 51, a first determination module 52, a second acquisition module 53, a third acquisition module 54, and a screening module 55.
[0264] Among them, the first acquisition module is used to acquire a plurality of keywords included in the target text, and the target text includes a selected keyword in a selected state.
[0265] The first determination module is used to determine, based on the first positions of the plurality of keywords in the target text respectively, at least one candidate keyword whose number of intervening words from the second position of the selected keyword in the target text is less than or equal to a first threshold.
[0266] The second acquisition module is used to obtain an association graph based on the at least one candidate keyword and the selected keyword; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and the weight of the edge between the two nodes is the relevance of the two nodes.
[0267] The third acquisition module is used to obtain the word importance values respectively corresponding to the at least one candidate keyword based on the association graph.
[0268] The screening module is used to obtain the target keyword from the at least one candidate keyword based on the word importance values respectively corresponding to the at least one candidate keyword.
[0269] In an optional embodiment, the second acquisition module includes:
[0270] The first acquisition unit is used to obtain the first relevance between the at least one candidate keyword and the selected keyword respectively.
[0271] The first construction unit is configured to, for each of the candidate keywords, if the first relevance between the candidate keyword and the selected keyword is greater than or equal to a second threshold, construct an edge between the candidate keyword and the selected keyword, and determine the first relevance as the weight of the edge between the candidate keyword and the selected keyword.
[0272] The second construction unit is configured to, if the at least one candidate keyword includes at least two candidate keywords, for any two candidate keywords, if the second relevance between the two candidate keywords is greater than or equal to a third threshold, construct an edge between the two candidate keywords, and determine the second relevance as the weight of the edge between the two candidate keywords.
[0273] In an alternative embodiment, the screening module includes:
[0274] The second acquisition unit is configured to obtain at least one target candidate keyword from the at least one candidate keyword based on the word importance values respectively corresponding to the at least one candidate keyword.
[0275] The third acquisition unit is configured to obtain a plurality of query records included in the query log within a preset time period, where the query records include the selected keyword and at least one of the target candidate keywords.
[0276] The determination unit is configured to, for each of the target candidate keywords, determine a first number of query records including the target candidate keyword in the plurality of query records, so as to obtain the first numbers respectively corresponding to the at least one target candidate keyword.
[0277] The fourth acquisition unit is configured to obtain a target keyword from the at least one target candidate keyword based on the first numbers respectively corresponding to the at least one target candidate keyword.
[0278] In an alternative embodiment, the first acquisition module includes:
[0279] The fifth acquisition unit is configured to obtain a plurality of words included in the target text.
[0280] The sixth acquisition unit is configured to, for each of the words, based on the second number of times the target text includes the word, the third number of texts in the preset text set that include the word, the total number of words included in the target text, the total number of texts included in the preset text set, and the total number of the plurality of words, obtain a text importance value representing the importance of the word for the target text, so as to obtain the text importance values respectively corresponding to the plurality of words.
[0281] A seventh obtaining unit, configured to obtain the plurality of keywords from the plurality of words based on the text importance values respectively corresponding to the plurality of words.
[0282] In an alternative embodiment, an embodiment of the present application provides an electronic device. Refer to Figure 6 As shown, it is a block diagram of an electronic device provided by an embodiment of the present application.
[0283] Exemplarily, the electronic device may be a terminal device 22 or a server 21.
[0284] The electronic device includes, but is not limited to, components such as an input unit 61, a memory 62, a display unit 63, and a processor 64. Those skilled in the art can understand that Figure 6 the structure shown in
[0285] is only an example of an implementation manner and does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 6 The following specifically introduces each component of the electronic device in conjunction with
[0286] Exemplarily, the input unit 61 can be used to obtain a user's selection operation on a word in the target text. For example, the user touches a certain word in the target text to determine the word as a selected keyword.
[0287] Exemplarily, the input unit 61 may include a touch panel 611 and other input devices 612. The touch panel 611, also known as a touch screen, can collect touch operations of the user thereon (such as operations of the user using a finger, a stylus, or any suitable object or accessory on the touch panel 611), and drive a corresponding connection device according to a preset program (such as driving the keyword obtaining function in the processor 64). Optionally, the touch panel 611 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 64, and can receive commands sent by the processor 64 and execute them. In addition, the touch panel 611 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 611, the input unit 61 may further include other input devices 612. Specifically, the other input devices 612 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0288] Exemplarily, the memory 62 can be used to store software programs and modules. The processor 64 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 62. The memory 62 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the electronic device (for example, a first threshold). In addition, the memory 62 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0289] Exemplarily, the display unit 63 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device. The display unit 63 may include a display panel 631. Optionally, the display panel 631 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode). Further, the touch panel 611 can cover the display panel 631. When the touch panel 611 detects a touch operation on or near it, it transmits the operation to the processor 64 to determine the type of touch event. Subsequently, the processor 64 provides a corresponding visual output on the display panel 631 according to the type of touch event.
[0290] Exemplarily, the touch panel 612 and the display panel 631 can be implemented as two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 612 and the display panel 631 can be integrated to realize the input and output functions of the electronic device.
[0291] The processor 64 is the control center of the electronic device. It connects various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 62, and by calling the data stored in the memory 62, it executes various functions of the electronic device and processes data. Exemplarily, the processor 64 may include one or more processing units; Exemplarily, the processor 64 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 64.
[0292] The electronic device also includes a power source 65 (such as a battery) for supplying power to each component. Exemplarily, the first power source can be logically connected to the processor 64 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0293] Although not shown, the electronic device may further include a camera, a Bluetooth module, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi (wireless fidelity) module, sensors, a network unit, an interface unit, and so on.
[0294] The network unit of the electronic device provides users with wireless broadband Internet access, such as accessing a server.
[0295] The interface unit is an interface for connecting an external device to the electronic device. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headset port, and so on. The interface unit can be used to receive inputs from an external device (such as data information, power, etc.) and transmit the received inputs to one or more components within the electronic device or can be used to transmit data between the electronic device and the external device.
[0296] In the embodiment of the present disclosure, the processor 64 included in the electronic device may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0297] The processor 64 included in the electronic device has the following functions: obtaining a plurality of keywords included in a target text, where the target text includes a selected keyword in a selected state; determining, based on first positions of the plurality of keywords in the target text respectively, at least one candidate keyword from the plurality of keywords whose number of intervening words from the second position of the selected keyword in the target text is less than or equal to a first threshold; obtaining an association graph based on the at least one candidate keyword and the selected keyword; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and the weight of the edge between the two nodes is the relevance between the two nodes; obtaining a word importance value corresponding to each of the at least one candidate keyword based on the association graph; and obtaining a target keyword from the at least one candidate keyword based on the word importance value corresponding to each of the at least one candidate keyword.
[0298] In an alternative embodiment, a storage medium is further provided, which can be directly loaded into the internal memory of a computer, such as the aforementioned memory 62, and contains software code. After being loaded and executed by the computer, the computer program can implement the steps shown in any of the embodiments of the keyword acquisition method applied to the electronic device described above.
[0299] It should be noted that the features described in the various embodiments in this specification can be replaced or combined with each other. For device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0300] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0301] The steps of the method or algorithm described in connection with the embodiments disclosed herein can be implemented directly in hardware, a software module executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0302] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for obtaining keywords, characterized in that, Comprising: If a search operation implemented on a selected keyword included in the target text is detected, obtain a plurality of keywords included in the target text; Based on the first positions of the plurality of keywords in the target text respectively, determine at least one candidate keyword from the plurality of keywords, where the number of intervening words between the second position of the selected keyword in the target text is less than or equal to a first threshold; Based on the at least one candidate keyword and the selected keyword, obtain an association graph; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and the weight of the edge between the two nodes is the relevance between the two nodes; Based on the association graph, obtain the word importance values respectively corresponding to the at least one candidate keyword; Based on the word importance values respectively corresponding to the at least one candidate keyword, obtain a target keyword from the at least one candidate keyword; Wherein, the step of obtaining a target keyword from the at least one candidate keyword based on the word importance values respectively corresponding to the at least one candidate keyword includes: Based on the word importance values respectively corresponding to the at least one candidate keyword, obtain at least one target candidate keyword from the at least one candidate keyword; Obtain a plurality of query records included in the query log within a preset time period, where the query records include the selected keyword and at least one of the target candidate keywords; For each of the target candidate keywords, determine a first number of query records including the target candidate keyword among the plurality of query records, so as to obtain the first numbers respectively corresponding to the at least one target candidate keyword; Based on the first numbers respectively corresponding to the at least one target candidate keyword, obtain a target keyword from the at least one target candidate keyword.
2. The keyword acquisition method according to claim 1, wherein The step of obtaining an association graph based on the at least one candidate keyword and the selected keyword includes: Obtain a first relevance between each of the at least one candidate keyword and the selected keyword; For each of the candidate keywords, if the first relevance between the candidate keyword and the selected keyword is greater than or equal to a second threshold, construct an edge between the candidate keyword and the selected keyword, and determine the first relevance as the weight of the edge between the candidate keyword and the selected keyword; If the at least one candidate keyword includes at least two candidate keywords, for any two candidate keywords, if the second relevance between the two candidate keywords is greater than or equal to a third threshold, construct an edge between the two candidate keywords, and determine the second relevance as the weight of the edge between the two candidate keywords.
3. The keyword acquisition method according to claim 1 or 2, characterized in that The step of obtaining a plurality of keywords included in the target text includes: Obtain a plurality of words included in the target text; For each of the said words, based on the second number of times the word is included in the target text, the total number of words in the target text, the third number of texts in the preset text set that contain the word, the total number of texts in the preset text set, and the total number of the multiple words, obtain a text importance value representing the importance of the word for the target text, so as to obtain the text importance values corresponding to the multiple words respectively; Based on the text importance values corresponding to the multiple words respectively, obtain the multiple keywords from the multiple words.
4. A keyword acquisition device, characterized in that, It includes: A first acquisition module, configured to acquire multiple keywords included in a target text, where the target text includes selected keywords in a selected state; A first determination module, configured to determine, based on the first positions of the multiple keywords in the target text respectively, at least one candidate keyword from the multiple keywords whose number of words spaced from the second position of the selected keyword in the target text is less than or equal to a first threshold; A second acquisition module, configured to obtain an association graph based on the at least one candidate keyword and the selected keyword; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and the weight of the edge between the two nodes is the relevance between the two nodes; A third acquisition module, configured to obtain the word importance values corresponding to the at least one candidate keyword respectively based on the association graph; A screening module, configured to obtain target keywords from the at least one candidate keyword based on the word importance values corresponding to the at least one candidate keyword respectively; Wherein, the screening module includes: A second acquisition unit, configured to obtain at least one target candidate keyword from the at least one candidate keyword based on the word importance values corresponding to the at least one candidate keyword respectively; A third acquisition unit, configured to acquire multiple query records included in a query log within a preset time period, where the query records include the selected keyword and at least one of the target candidate keywords; A determination unit, configured to, for each of the target candidate keywords, determine the first number of query records that include the target candidate keyword in the multiple query records, so as to obtain the first numbers corresponding to the at least one target candidate keyword respectively; A fourth acquisition unit, configured to obtain target keywords from the at least one target candidate keyword based on the first numbers corresponding to the at least one target candidate keyword respectively.
5. The keyword acquisition device according to claim 4, characterized in that The second acquisition module includes: A first acquisition unit, configured to obtain the first relevance between the at least one candidate keyword and the selected keyword respectively; A first construction unit, configured to, for each of the candidate keywords, if the first relevance between the candidate keyword and the selected keyword is greater than or equal to a second threshold, construct an edge between the candidate keyword and the selected keyword, and determine the first relevance as the weight of the edge between the candidate keyword and the selected keyword; A second construction unit, configured to, if the at least one candidate keyword includes at least two candidate keywords, for any two candidate keywords, if the second relevance between the two candidate keywords is greater than or equal to a third threshold, construct an edge between the two candidate keywords and determine the second relevance as the weight of the edge between the two candidate keywords.
6. The keyword acquisition device according to claim 4 or 5, characterized in that The first acquisition module includes: A fifth acquisition unit, configured to acquire a plurality of words included in the target text; A sixth acquisition unit, configured to, for each of the words, based on the second number of times the target text includes the word, the third number of texts in the preset text set that include the word, the total number of words included in the target text, the total number of texts included in the preset text set, and the total number of the plurality of words, obtain a text importance value representing the importance of the word to the target text, so as to obtain the text importance values respectively corresponding to the plurality of words; A seventh acquisition unit, configured to obtain the plurality of keywords from the plurality of words based on the text importance values respectively corresponding to the plurality of words.
7. An electronic device, characterized in that, It includes: A memory, configured to store a program; A processor, configured to execute the program, and the program is specifically configured to: Acquire a plurality of keywords included in a target text, where the target text includes a selected keyword in a selected state; Based on the first positions of the plurality of keywords in the target text respectively, determine at least one candidate keyword from the plurality of keywords, the number of words separated from the second position of the selected keyword in the target text being less than or equal to a first threshold; Based on the at least one candidate keyword and the selected keyword, obtain an association graph; the at least one candidate keyword and the selected keyword are respectively nodes in the association graph, and there is an edge between any two nodes in the association graph whose relevance is greater than or equal to a corresponding threshold, and the weight of the edge between the two nodes is the relevance between the two nodes; Based on the association graph, obtain the word importance values respectively corresponding to the at least one candidate keyword; Based on the word importance values respectively corresponding to the at least one candidate keyword, obtain a target keyword from the at least one candidate keyword; Wherein, the step of obtaining a target keyword from the at least one candidate keyword based on the word importance values respectively corresponding to the at least one candidate keyword includes: Based on the word importance values respectively corresponding to the at least one candidate keyword, obtain at least one target candidate keyword from the at least one candidate keyword; Acquire a plurality of query records included in a query log within a preset time period, where the query records include the selected keyword and at least one of the target candidate keywords; For each of the target candidate keywords, determine the first number of query records including the target candidate keyword in the plurality of query records, so as to obtain the first numbers respectively corresponding to the at least one target candidate keyword; Based on the first numbers respectively corresponding to the at least one target candidate keyword, obtain a target keyword from the at least one target candidate keyword.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the keyword acquisition method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Graph-based keyword expansion
US20090234832A1