Information search method and related equipment
By constructing the core event graph and performing event distribution analysis, the problems of low efficiency and insufficient accuracy of text summary generation in the prior art are solved, and more accurate text summary generation and improved search results quality are achieved.
Patent Information
- Application Number
- CN202111454976.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-01
AI Technical Summary
The prior art is inefficient and inaccurate when generating text summary, and cannot accurately express core events in the text, resulting in low output efficiency and accuracy of search results.
By obtaining the keywords of text information under the same topic, a core event graph is constructed, node search and event distribution analysis are performed based on the weights and association relationships of the keyword group to generate a more accurate text summary.
It improves the efficiency and accuracy of text summary generation, and can express core events in text information more accurately, thereby improving the output efficiency and accuracy of search results.
Smart Images

Figure CN114328820B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an information search method and related equipment. Background Art
[0002] With the development of computer technology, more and more information is disclosed through the Internet. The text information on the Internet has exploded, and people are exposed to a large number of documents every day, such as news and papers. It is very important to extract key information from a large amount of document information. For example, this key information can be used in search, and generating text summaries is a means of extracting key information. Text summarization technology can convert documents into short summaries containing key information, that is, to summarize the core events contained in the text in refined language and generate event summaries of the corresponding text; it can help users quickly understand the content of the document, and can also be applied to the hot topic list on the search page.
[0003] In the current related technologies, the content of the original document can usually be understood based on semantics through a neural network model, and important information can be automatically extracted from the original document to generate a corresponding summary. However, the efficiency of summarizing the summary is low, and the accuracy of the generated text summary is low. It cannot accurately express the core events contained in the text, resulting in low output efficiency and accuracy of the search results. Summary of the invention
[0004] The embodiments of the present application provide an information search method and related equipment. The related equipment may include an information search device, an electronic device, a computer-readable storage medium and a computer program product, which can improve the efficiency and accuracy of generating text summaries, more accurately express the core events contained in the text information, and thereby improve the output efficiency and accuracy of search results.
[0005] The present application provides an information search method, including:
[0006] Acquire keywords corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word;
[0007] Extracting at least one keyword group from the keywords, and determining a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords having an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information;
[0008] According to the keyword groups and the weights corresponding to the keyword groups, a core event graph is constructed, wherein the core event graph includes nodes corresponding to the keywords and lines between the nodes, wherein the lines represent associations between the keywords;
[0009] Performing a node search in the core event graph according to the reference abstract word to obtain a target-related keyword associated with the reference abstract word in the core event graph;
[0010] Performing event distribution analysis on the weights of the target-related keywords and the keyword group containing the reference summary words to obtain event probability information;
[0011] According to the event probability information and the hidden state information of the reference summary word, the candidate summary word and the reference summary word are subjected to summary construction processing to obtain a text summary corresponding to the at least one text information, and a search result is output according to the text summary.
[0012] Accordingly, an embodiment of the present application provides an information search device, including:
[0013] An acquisition unit, configured to acquire a keyword corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word;
[0014] A phrase extraction unit, configured to extract at least one keyword group from the keywords and determine a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords having an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information;
[0015] An event graph construction unit, configured to construct a core event graph according to the keyword groups and the weights corresponding to the keyword groups, wherein the core event graph includes nodes corresponding to the keywords and lines between the nodes, wherein the lines represent associations between the keywords;
[0016] A node search unit, configured to perform a node search in the core event graph according to the reference summary word, and obtain a target associated keyword associated with the reference summary word in the core event graph;
[0017] An analysis unit, configured to perform event distribution analysis on the weights of the target-related keywords and the keyword group in which the reference summary words are located, to obtain event probability information;
[0018] A summary construction unit is used to perform summary construction processing on the candidate summary words and the reference summary words according to the event probability information and the hidden state information of the reference summary words, obtain a text summary corresponding to the at least one text information, and output a search result according to the text summary.
[0019] Optionally, in some embodiments of the present application, the acquisition unit may include a word segmentation subunit, an entity recognition subunit, a frequency analysis subunit, and a keyword selection subunit, as follows:
[0020] The word segmentation subunit is used to perform word segmentation processing on at least one text information under the same topic to obtain at least one word segmentation of the text information;
[0021] An entity recognition subunit, configured to perform entity recognition on each of the segmented words of the text information, so as to determine at least one entity segmented word from each of the segmented words of the text information;
[0022] A frequency analysis subunit, used to perform frequency analysis on each entity segmentation in the text information to obtain a word frequency parameter of each entity segmentation in the text information;
[0023] The keyword selection subunit is used to select keywords of the text information from the entity segmentation based on the word frequency parameter, and obtain a reference summary word and at least one candidate summary word of the text information.
[0024] Optionally, in some embodiments of the present application, the frequency analysis subunit can be specifically used to count the frequency of occurrence of each entity segmentation in the text information to obtain the weight of the entity segmentation in the text information; count the frequency of occurrence of the entity segmentation in sample text information to obtain the reference weight of the entity segmentation; determine the word frequency parameter of the entity segmentation based on the reference weight of the entity segmentation and its weight in the text information.
[0025] Optionally, in some embodiments of the present application, each text message includes at least one text sentence;
[0026] The phrase extraction unit may include a text sentence selection subunit, a target keyword selection subunit, and a phrase construction subunit, as follows:
[0027] The text sentence selection subunit is used to select a key text sentence of the text information from each text sentence of the text information based on the matching degree between the keyword and each text sentence of the text information for each text information;
[0028] A target keyword selection subunit is used to select at least one target keyword from the segmented words corresponding to the key text sentences of each text information based on the keywords and the semantic dependency between the segmented words in the key text sentences of each text information;
[0029] The phrase construction subunit is used to construct at least one keyword group according to the semantic dependency relationship between each target keyword.
[0030] Optionally, in some embodiments of the present application, the text sentence selection subunit can be specifically used to count the number of the keywords contained in each text sentence in each text information; based on the number, determine the matching degree between the text sentence and the keyword; and based on the matching degree, select the key text sentence of the text information from each text sentence in the text information.
[0031] Optionally, in some embodiments of the present application, the target keyword selection subunit can be specifically used to perform syntactic structure analysis on the key text sentence of each text information to determine the semantic dependency relationship between each word segment in the key text sentence; based on the semantic dependency relationship, extract the first target keyword of the key text sentence from each word segment of the key text sentence; based on the keywords contained in the key text sentence, obtain the second target keyword of the key text sentence; and determine the target keyword corresponding to the key text sentence based on the first target keyword and the second target keyword.
[0032] Optionally, in some embodiments of the present application, each text message includes at least one text sentence;
[0033] The summary construction unit may include a selection subunit, a coding subunit and a summary construction subunit, as follows:
[0034] The selection subunit is used to select a target key text sentence from the text sentences corresponding to each text information based on the keyword and the core event graph;
[0035] The encoding and decoding subunit is used to perform encoding and decoding processing on each word segment in the target key text sentence to obtain probability information of the target key text sentence;
[0036] The summary construction subunit is used to select a target summary word from each candidate summary word according to the probability information of the target key text sentence, the event probability information and the hidden state information of the reference summary word, so as to perform summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information.
[0037] Optionally, in some embodiments of the present application, the selection subunit can be specifically used to select, for each text information, a key text sentence of the text information from each text sentence of the text information based on the keyword; for the key text sentence corresponding to each text information, the word frequency parameters corresponding to the keyword contained in the key text sentence are merged to obtain a first important indicator of the key text sentence; according to the phrases formed by each word segment in the key text sentence, a node search is performed in the core event graph to obtain the weight corresponding to the keyword group contained in the key text sentence; the weights corresponding to each keyword group contained in the key text sentence are merged to obtain a second important indicator of the key text sentence; based on the first important indicator and the second important indicator, a target key text sentence is selected from the key text sentences corresponding to each text information.
[0038] Optionally, in some embodiments of the present application, the encoding and decoding subunit can be specifically used to perform feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence; fuse the encoding information of each word in the target key text sentence to obtain context feature information; and decode the context feature information based on current decoding parameters to obtain probability information of the target key text sentence.
[0039] Optionally, in some embodiments of the present application, the step of “performing feature encoding on each segmentation in the target key text sentence to obtain encoding information of each segmentation in the target key text sentence” may include:
[0040] Based on the sequence of each segmentation in the target key text sentence, forward encoding each segmentation in the target key text sentence is performed to obtain forward encoding information of each segmentation in the target key text sentence;
[0041] Based on the sequence of each segmentation in the target key text sentence, each segmentation in the target key text sentence is backward encoded to obtain backward encoding information of each segmentation in the target key text sentence;
[0042] The forward encoding information and the backward encoding information are merged to obtain encoding information of each word segment in the target key text sentence.
[0043] Optionally, in some embodiments of the present application, the analysis unit may include a first fusion subunit, a logistic regression subunit, and a second fusion subunit, as follows:
[0044] The first fusion subunit is used to fuse the weights of each target-related keyword with the keyword group where the reference summary word is located to obtain initial event probability information;
[0045] The logistic regression subunit is used to perform logistic regression processing on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight;
[0046] The second fusion subunit is used to fuse the initial event probability information and the context adaptive weight to obtain event probability information.
[0047] Optionally, in some embodiments of the present application, the summary construction unit may include an attention feature extraction subunit, a calculation subunit, a summary word selection subunit and a construction subunit, as follows:
[0048] The attention feature extraction subunit is used to extract attention features of each candidate summary word according to the hidden state information of the reference summary word, so as to obtain attention feature information corresponding to each candidate summary word;
[0049] A calculation subunit, used to calculate target probability information corresponding to each candidate summary word according to the event probability information and the attention feature information;
[0050] A summary word selection subunit, configured to select a target summary word from each candidate summary word based on the target probability information;
[0051] The construction subunit is used to perform summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information.
[0052] Optionally, in some embodiments of the present application, the attention feature extraction subunit can be specifically used to perform feature extraction on each candidate summary word according to the hidden state information of the reference summary word to obtain the hidden state information corresponding to the candidate summary word; perform logistic regression processing on the hidden state information corresponding to the candidate summary word to obtain the attention feature information corresponding to the candidate summary word.
[0053] An electronic device provided in an embodiment of the present application includes a processor and a memory, wherein the memory stores a plurality of instructions, and the processor loads the instructions to execute the steps in the information search method provided in the embodiment of the present application.
[0054] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps in the information search method provided in the embodiment of the present application.
[0055] In addition, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the information search method provided in the embodiment of the present application.
[0056] The embodiment of the present application provides an information search method and related devices, which can obtain keywords corresponding to at least one text information under the same topic, a reference summary word of the text information, and at least one candidate summary word; extract at least one keyword group from the keywords, and determine the weight corresponding to the keyword group, the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; construct a core event graph according to the keyword group and the weight corresponding to the keyword group, the core event graph includes nodes corresponding to the keywords, and lines between each node, the lines represent the association relationship between the keywords; perform node search in the core event graph according to the reference summary word to obtain a target-related keyword associated with the reference summary word in the core event graph; perform event distribution analysis on the weights of the target-related keyword and the keyword group where the reference summary word is located to obtain event probability information; perform summary construction processing on the candidate summary word and the reference summary word according to the event probability information and the hidden state information of the reference summary word to obtain a text summary corresponding to the at least one text information, and output search results according to the text summary. The embodiment of the present application can construct a core event graph through the keywords contained in the text information under the same topic, which can better refine the core event information, and then construct a text summary based on the core event graph, which is conducive to improving the generation efficiency and accuracy of the text summary, more accurately expressing the core events contained in the text information, and thus improving the output efficiency and accuracy of the search results. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0058] Figure 1a is a scenario diagram of the information search method provided in an embodiment of the present application;
[0059] Figure 1b is a flow chart of an information search method provided by an embodiment of the present application;
[0060] Figure 1c is a page schematic diagram of the information search method provided in an embodiment of the present application;
[0061] Figure 1d is an illustration of an information search method provided by an embodiment of the present application;
[0062] Figure 1e It is an overall architecture diagram of the information search method provided in the embodiment of the present application;
[0063] Figure 2 is another flow chart of the information search method provided by an embodiment of the present application;
[0064] Figure 3 is a schematic diagram of the structure of an information search device provided in an embodiment of the present application;
[0065] Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0067] The embodiment of the present application provides an information search method and related equipment, which may include an information search device, an electronic device, a computer-readable storage medium, and a computer program product. The information search device may be integrated in an electronic device, which may be a terminal or a server.
[0068] It is understandable that the information search method of this embodiment can be executed on a terminal, can be executed on a server, or can be executed by both a terminal and a server. The above examples should not be construed as limiting the present application.
[0069] like Figure 1a As shown, the information search method is performed by a terminal and a server together as an example. The information search system provided in the embodiment of the present application includes a terminal 10 and a server 11, etc. The terminal 10 and the server 11 are connected via a network, such as a wired or wireless network connection, etc., wherein the information search device can be integrated in the server.
[0070] Wherein, the server 11 can be used to: obtain keywords corresponding to at least one text information under the same topic, reference abstract words of the text information, and at least one candidate abstract word; extract at least one keyword group from the keywords, and determine the weight corresponding to the keyword group, the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; construct a core event graph according to the keyword group and the weight corresponding to the keyword group, the core event graph includes nodes corresponding to the keywords, and lines between each node, the lines represent the association relationship between keywords; perform node search in the core event graph according to the reference abstract word to obtain the target associated keyword associated with the reference abstract word in the core event graph; perform event distribution analysis on the weight of the target associated keyword and the keyword group where the reference abstract word is located to obtain event probability information; perform abstract construction processing on the candidate abstract word and the reference abstract word according to the event probability information and the hidden state information of the reference abstract word to obtain the text abstract corresponding to the at least one text information, and output the search result according to the text abstract. Wherein, the server 11 can be a single server, or a server cluster or cloud server composed of multiple servers. The information search method or device disclosed in the present application, wherein multiple servers can be composed into a blockchain, and the servers are nodes on the blockchain.
[0071] The terminal 10 may be used to receive a text summary corresponding to the at least one text message sent by the server 11, and display the text summary. The terminal 10 may include a mobile phone, a tablet computer, a laptop computer, a personal computer (PC), an intelligent voice interaction device, an intelligent home appliance, a vehicle-mounted terminal, etc. A client may also be set on the terminal 10, and the client may be an application client or a browser client, etc.
[0072] The step of generating a text summary by the server 11 may also be executed by the terminal 10 .
[0073] The information search method provided in the embodiment of the present application relates to natural language processing in the field of artificial intelligence. The embodiment of the present application can improve the efficiency and accuracy of generating text summaries and more accurately express the core events contained in the text information. The embodiment can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0074] Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level technology and software-level technology. Among them, artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, smart transportation and other major directions.
[0075] Among them, natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0076] It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments.
[0077] This embodiment will be described from the perspective of an information search device. The information search device may be integrated into an electronic device, which may be a server or a terminal.
[0078] The information search method provided in the present application can be applied in various scenarios where text summaries need to be generated, and specifically, can be applied in the search and display scenarios of news events.
[0079] like Figure 1b As shown, the specific process of the information search method can be as follows:
[0080] 101. Obtain keywords corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word.
[0081] The theme is the central idea to be expressed by the text information, generally refers to the main content, and specifically can also be the core event of the text information. There are many types of themes, and this embodiment does not limit this. For example, each text information in this embodiment can belong to a certain entertainment theme, or can also belong to a certain science and technology theme.
[0082] Specifically, each text information may be text information in an article, a paragraph, or a document. The reference summary word may be a summary word that has been selected as a text summary corresponding to the text information, or may be a reference word for selecting a summary word to form a text summary, which is not limited in this embodiment. The text summary may include at least one summary word.
[0083] This embodiment can select a target abstract word from candidate abstract words based on the reference abstract word, and then determine the text abstract corresponding to the text information under the same topic based on the target abstract word and the reference abstract word. Specifically, the reference abstract word can be updated based on the target abstract word, such as adding the target abstract word to the word set corresponding to the reference abstract word to obtain the updated word set corresponding to the reference abstract word; then obtain new candidate abstract words, which may not contain the selected target abstract word compared to the original candidate abstract words, so that based on the word set corresponding to the updated reference abstract word, a new target abstract word is selected from the new candidate abstract words, so that multiple target abstract words can be obtained to construct a text abstract.
[0084] Among them, the text summary can also be understood as a short description of the event corresponding to one or more text information under the same topic. The short description of the event has important application value for information flow platforms and search systems. For example, the short description of the event can be used in hot topic lists, event query recommendations, related searches and other scenarios.
[0085] This application can identify core events for multiple articles describing the same topic, and summarize the description in refined language. Specifically, the information search method provided by this application can be divided into two stages, namely: the core event identification stage and the summary generation stage (i.e., the short description generation stage). Among them, in the core event identification stage, the core event identification at the graph granularity and sentence granularity can be performed by constructing a core event graph, so that the core event information can be better located, the integrity and semantic coherence of the event extraction can be ensured, and a higher quality text summary can be generated; in the summary generation stage, the core event graph and the information of the target key text sentences can be fused to generate a text summary, which improves the accuracy of the generated text summary. Applying this information search method in search-related scenarios can improve multiple types of online indicators.
[0086] like Figure 1cAs shown in page a, it shows the hot topic list in the search page of a certain application; the solution based on this application can generate short event descriptions with high timeliness, allowing users to perceive hot events on the entire network; automatically generating short event descriptions based on multiple text information can save the manpower cost of operation, and at the same time, compared with manual operation and editing, automatically generating short event descriptions has the characteristics of high timeliness and diversity. Figure 1c Page b shows query tips when searching for events. For example, when searching for "X City", query tips related to "X City" can be displayed on the search page, such as "the cause of the explosion in X City", "the entire X City government resigned", etc. Figure 1c Page c of the search results shows related searches for a certain event. When searching for the event "X City Explosion", related searches such as "X City Explosion Details", "X City Port Explosion Causes", etc. can be displayed on the search page. When searching, providing users with hot event query recommendation services can improve user click-through rate and perception of news hot spots. Since the query prompts and related searches are both large in magnitude, automatically generating short event descriptions based on multiple text information can cover more scenarios, thereby improving user click-through rate and search satisfaction.
[0087] Optionally, in this embodiment, the step of “obtaining keywords corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word” may include:
[0088] Performing word segmentation processing on at least one text information under the same topic to obtain at least one word segmentation of the text information;
[0089] Performing entity recognition on each of the segmented words of the text information to determine at least one entity segmented word from each of the segmented words of the text information;
[0090] Performing frequency analysis on each entity segmentation in the text information to obtain a word frequency parameter of each entity segmentation in the text information;
[0091] Based on the word frequency parameter, keywords of the text information are selected from the entity segmentation words, and a reference summary word and at least one candidate summary word of the text information are obtained.
[0092] The word segmentation of the text information can be a word or a character, which is not limited in this embodiment. The word segmentation process can adopt Jieba word segmentation, which is a length-first word segmentation scheme, which is coarser than other word segmentation schemes and can reduce the difficulty of generation.
[0093] Among them, entity recognition can specifically be the identification of entities with specific meanings in text information, mainly including names of people, places, institutions, proper nouns, etc., as well as text such as time, quantity, currency, and proportional values.
[0094] Optionally, in this embodiment, entity participles whose word frequency parameter is greater than a preset value can be selected as keywords of the text information, and the preset value can be set according to actual conditions; or the entity participles of the text information can be sorted based on the word frequency parameter, such as sorting them from large to small to obtain sorted entity participles, and the first n entity participles after sorting are used as keywords of the text information.
[0095] In some embodiments, before the text information is segmented, the text information may be filtered, for example, some articles with poor news value may be filtered out; some externally linked articles may also be cleaned, so as to ensure the quality of the articles.
[0096] Optionally, in this embodiment, the step of “performing frequency analysis on each entity segmentation in the text information to obtain a word frequency parameter of each entity segmentation in the text information” may include:
[0097] For each entity participle of the text information, counting the frequency of occurrence of the entity participle in the text information, and obtaining the weight of the entity participle in the text information;
[0098] Counting the frequency of occurrence of the entity segmentation in the sample text information to obtain a reference weight of the entity segmentation;
[0099] A word frequency parameter of the entity segmentation is determined according to the reference weight of the entity segmentation and the weight in the text information.
[0100] Among them, the frequency of the entity segmentation in the text information is specifically the word frequency of the entity segmentation in the text information, which can be represented by TF. TF refers to the frequency of a given word in a file. This is a normalization of the number of words to prevent it from being biased towards long files. In this embodiment, the word frequency of the entity segmentation in the text information can be directly used as the weight of the entity segmentation in the text information.
[0101] The sample text information may specifically be text in a document library, and the reference weight of the entity segmentation may be represented by an inverse text frequency. The inverse text frequency of a certain entity segmentation represents the frequency of its occurrence in the corpus.
[0102] Optionally, the step of “determining a frequency parameter of the entity segmentation according to the reference weight of the entity segmentation and the weight in the text information” may include:
[0103] The reference weight of the entity segmentation and the weight in the text information are fused to obtain a word frequency parameter of the entity segmentation.
[0104] There are many ways to merge, for example, the fusion method can be multiplication, etc., which is not limited in this embodiment. Specifically, the frequency parameter of the entity segmentation can be obtained by multiplying the frequency TF and the inverse text frequency IDF of the entity segmentation, and the frequency parameter can be represented by TF-IDF.
[0105] TF-IDF stands for Term Frequency–Inverse Document Frequency, which is a commonly used weighting technique for information retrieval and text mining. TF-IDF is a statistical method used to assess the importance of a word to a document set or a document in a corpus. The importance of a word increases in direct proportion to the number of times it appears in a document, but decreases in inverse proportion to the frequency of its appearance in the corpus.
[0106] Optionally, for the extraction of keywords from text information, a keyword extraction scheme such as TextRank may also be used, which is not limited in this embodiment. The TextRank algorithm is a graph-based ranking algorithm for keyword extraction and document summary generation. It can extract keywords using semantic co-occurrence information between words within a document, and can extract keywords and keyword groups from a given text, and use an extractive automatic summarization method to extract key sentences from the text. The basic idea of the TextRank algorithm is to regard a document as a network of words, and the links in the network represent the semantic relationship between words.
[0107] 102. Extract at least one keyword group from the keywords, and determine a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords having an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information.
[0108] The association relationship between the keywords in the keyword group may specifically include the semantic dependency relationship between the keywords.
[0109] Optionally, in this embodiment, each text message includes at least one text sentence;
[0110] The step of “extracting at least one keyword group from the keywords” may include:
[0111] For each piece of text information, based on the matching degree between the keyword and each text sentence of the text information, selecting a key text sentence of the text information from each text sentence of the text information;
[0112] Based on the keywords and the semantic dependency between the segmented words in the key text sentences of each text information, selecting at least one target keyword from the segmented words corresponding to the key text sentences of each text information;
[0113] At least one keyword group is constructed according to the semantic dependency relationship between the target keywords.
[0114] Specifically, every two target keywords having a semantic dependency relationship may be regarded as a keyword group, that is, each keyword group includes two target keywords having a semantic dependency relationship.
[0115] In some embodiments, the weight corresponding to the keyword group can be counted by co-occurrence, and the weight corresponding to the keyword group can characterize the number of times the keyword group appears in the key text sentences of each text information. For example, the number of sentences of the key text sentences containing a certain keyword group can be used as the weight corresponding to the keyword group. Specifically, when the target keyword included in the keyword group appears in a key text sentence at the same time, and the correlation relationship between the target keywords in the key text sentence is consistent with the correlation relationship between the target keywords in the keyword group, it can be regarded as that the key text sentence contains this keyword group.
[0116] For example, a keyword group includes two target keywords with a subject-predicate relationship. If a key text sentence contains these two target keywords, and these two target keywords are in a subject-predicate relationship in the key text sentence, then the key text sentence can be considered to contain the keyword group. If there are two key text sentences containing these two target keywords, and the target keywords are associated with a subject-predicate relationship in these two key text sentences, then the weight of the keyword group containing these two target keywords can be 2.
[0117] Optionally, in this embodiment, the step of “for each text information, based on the matching degree between the keyword and each text sentence of the text information, selecting a key text sentence of the text information from each text sentence of the text information” may include:
[0118] For each text sentence in each text information, counting the number of the keywords contained in the text sentence;
[0119] Based on the number, determining a degree of match between the text sentence and the keyword;
[0120] According to the matching degree, a key text sentence of the text information is selected from each text sentence of the text information.
[0121] In some embodiments, a text sentence in the text information with a matching degree greater than a preset value can be selected as a key text sentence of the text information, and the preset value can be set according to actual conditions. In other embodiments, the text sentences of the text information can be sorted according to the matching degree, such as sorting from large to small, to obtain sorted text sentences, and then the first n text sentences can be selected as the key text sentences of the text information.
[0122] Specifically, the calculation process of the matching degree between each text sentence and the keyword of the text information can be shown as formula (1):
[0123]
[0124] in, Indicates text information i The corresponding keywords (can also be regarded as a keyword set), Indicates text information i The jth text sentence in Indicates text information i The jth text sentence and the text information d i Corresponding keywords The degree of match between them.
[0125] The numerator in formula (1) represents the text information d i The jth text sentence, and text information d i Corresponding keyword set The number of common keywords.
[0126] Optionally, in some embodiments, for each text information, which has a title and a body, two key text sentences can be selected from each text sentence of the text information, one of which is the title sentence of the text information, and the other is a key text sentence selected from each body text sentence based on the matching degree between the keyword and each body text sentence; if there are N text information, 2N key text sentences can be selected. In this way, the title of each text information (specifically, each article) and an important sentence in the body are selected as key text sentences, which can avoid the situation where some titles do not highlight the theme of the article, or avoid the situation where the title is not related to the theme of the article.
[0127] Optionally, in some embodiments, the step of “for each text message, based on the matching degree between the keyword and each text sentence of the text message, selecting a key text sentence of the text message from each text sentence of the text message” may include:
[0128] For each text message, based on the matching degree between the keyword and each body text sentence of the text message, selecting a body key text sentence of the text message from each body text sentence of the text message;
[0129] The key text sentence of the text information is obtained according to the key text sentence of the main text and the title sentence corresponding to the text information.
[0130] Among them, both the main text key sentence and the title sentence can be used as the key text sentence of the text information.
[0131] Optionally, in this embodiment, the step of “selecting at least one target keyword from the segmented words corresponding to the key text sentences of each text information based on the keyword and the semantic dependency between the segmented words in the key text sentences of each text information” may include:
[0132] For each key text sentence of the text information, a syntactic structure analysis is performed on the key text sentence to determine the semantic dependency relationship between each word segment in the key text sentence;
[0133] Extracting a first target keyword of the key text sentence from each word segment of the key text sentence according to the semantic dependency;
[0134] Based on the keywords included in the key text sentence, obtaining a second target keyword of the key text sentence;
[0135] According to the first target keyword and the second target keyword, a target keyword corresponding to the key text sentence is determined.
[0136] Among them, the syntactic structure analysis can specifically include the grammatical relationship annotation of each participle in the key text sentence. The grammatical relationship corresponding to the participle includes the semantic dependency relationship formed by the context in which the participle is located, etc., which is not limited in this embodiment. Semantic Dependency Parsing (SDP) can analyze the semantic association between each language unit of the sentence and present the semantic association in a dependency structure.
[0137] Among them, the semantic dependency relationship between participles may include subject-predicate relationship, verb-object relationship, indirect object relationship, attributive-predicate relationship, etc., which is not limited in this embodiment.
[0138] Among them, the participles in the key text sentence can be deleted according to the semantic dependency relationship between the participles in the key text sentence. Specifically, in some embodiments, the participles whose semantic dependency relationship meets the preset relationship can be retained. For example, the participles whose semantic dependency relationship is the subject-predicate relationship, the verb-object relationship, and the intermezzo-object relationship can be retained, and these participles are used as the first target keywords of the target key text sentence.
[0139] Specifically, the step of "obtaining the second target keyword of the key text sentence based on the keywords contained in the key text sentence" may include: taking the participles in the key text sentence that have a semantic dependency relationship with the keywords and the keywords contained in the key text sentence as the second target keywords of the key text sentence.
[0140] Among them, the first target keyword and the second target keyword corresponding to the key text sentence may be determined as the target keywords corresponding to the key text sentence.
[0141] 103. Construct a core event graph according to the keyword groups and the weights corresponding to the keyword groups. The core event graph includes nodes corresponding to the keywords and lines between the nodes. The lines represent associations between the keywords.
[0142] Among them, the nodes of the core event graph are keywords rather than sentences. This can better model the association between event dimensions between different sentences and better explore the connection between event information between different articles. Moreover, through keyword extraction, named entity recognition, semantic dependency analysis and other technologies, the constructed core event graph can basically cover important event element information, which can provide a basis for the selection of target key text sentences.
[0143] Specifically, the connection lines between nodes can represent the semantic dependency relationship between the keywords corresponding to the nodes, and the weights corresponding to the keyword groups composed of the keywords.
[0144] Specifically, the nodes of the core event graph may be the target keywords in the above-mentioned embodiments; the target keywords are selected from the segmented words corresponding to the key text sentences of each text information based on the semantic dependencies between the keywords and the segmented words in the key text sentences of each text information.
[0145] In a specific scenario, for the construction of the core event graph of multiple related articles, the input can be a set of multiple related articles (that is, at least one text information under the same topic), denoted as d 1 ,d 2 ,…,d N Respectively represent each article (ie, each text information), and the output may be a core event graph composed of event elements, and each node in the core event graph may represent an event element.
[0146] Specifically, by extracting keywords from each article, we can get the keywords corresponding to each article. (i.e., keyword set), and keyword set kw of article cluster D , where kw D The keywords corresponding to each article Merged.
[0147] In addition to the extraction of keywords, this embodiment can also extract key text sentences from multiple articles as the source of the construction of the core event graph. Specifically, the extraction of key text sentences can refer to the description of the steps in the above embodiment, that is, the specific description of the step "for each text information, based on the matching degree between the keywords and the various text sentences of the text information, select the key text sentences of the text information from the various text sentences of the text information", which will not be repeated here.
[0148] After obtaining the key text sentences of each text information, the semantic relationship between each event element can be extracted through semantic dependency analysis, that is, the semantic dependency relationship between each segmentation in the key text sentence can be analyzed. Specifically, the segmentation in the key text sentence that has a subject-predicate relationship, a verb-object relationship, or a subject-predicate relationship with other segmentations can be selected as the target keyword; in other words, the phrase composed of two segmentations in the key text sentence that have a subject-predicate relationship, a verb-object relationship, or a subject-predicate relationship can be determined as a keyword group. In addition, the keyword set kw in the key text sentence that is related to the article cluster can also be selected. D The keywords in the article are selected as target keywords; in other words, the keyword set kw based on the article cluster D , determine the keywords contained in the key text sentence, and then determine the phrase consisting of the segmented words that have a semantic dependency relationship with the keywords and the corresponding keywords in the key text sentence as the keyword group. Among them, the weight of the keyword group It can be determined based on the number of key text sentences containing the keyword group. Then, based on the keyword group and the weight corresponding to the keyword group, the core event graph G is constructed. e .
[0149] Among them, the edges of the core event graph can be defined as Among them, t ij is the semantic dependency type between nodes, specifically representing the semantic dependency between the i-th target keyword and the j-th target keyword. is the weight corresponding to the keyword group containing the i-th target keyword and the j-th target keyword.
[0150] For example, Figure 1d As shown in the figure, it is a schematic diagram of the construction of the core event graph, which is described in detail as follows:
[0151] Based on the matching degree between the keywords and each text sentence of the text information, the key text sentences selected from the text sentences of each text information are as follows:
[0152] Key text sentence A: On August 4, a huge explosion occurred in the port warehouse area of Y City.
[0153] A huge explosion occurred in the warehouse area of the port of CityY, on August 4.
[0154] Key text sentence B: The explosion in Y city, X country has caused nearly 200 deaths and more than 6,500 injuries.
[0155] City Y bombing in X Country has killed nearly 200 people and injured more than 6500.
[0156] Key text sentence C: A large-scale demonstration broke out in City X of Country X.
[0157] Large-scale demonstrations erupt in City X in Country X.
[0158] Key text sentence D: The explosion has made the economy of Country X even worse. Preliminary estimates show that the economic losses may exceed US$15 billion.
[0159] The huge explosion made Country X's worse economy, and the economic loss is likely to exceed 15 billion dollars according to preliminary estimates.
[0160] …
[0161] Then analyze the semantic dependency relationship between each word segment in the key text sentence; based on the keywords and the semantic dependency relationship between the word segments in each key text sentence, select the target keyword from the word segment corresponding to each key text sentence; and construct the keyword group according to the semantic dependency relationship between each target keyword.
[0162] For example, for the key text sentence A, the semantic dependency relationship between the participles "port" and "occurrence" is a subject-predicate relationship, so "port" and "occurrence" can be selected as target keywords, that is, the phrase formed by "port" and "occurrence" is determined as a keyword group.
[0163] After obtaining the keyword groups and the corresponding weights of the keyword groups, we can construct Figure 1dAs shown in the core event graph, the nodes in the core event graph represent target keywords, and the edges represent the association relationship between target keywords. Specifically, it can represent the weight of the keyword group composed of the target keywords corresponding to the nodes, and the semantic dependency relationship between the target keywords.
[0164] exist Figure 1d In the semantic dependency annotation of , SBV indicates the subject-predicate relationship, VOB indicates the verb-object relationship, ATT indicates the attributive relationship, and TIME indicates the time. For example, in the key text sentence A, "port" and "occurrence" are in the subject-predicate relationship, and the semantic dependency relationship between the two is marked as SBV. And the number of times the keyword group composed of "port" and "occurrence" appears in each key text sentence is 2, then the edge connecting the node corresponding to "port" and the node corresponding to "occurrence" can be marked as (SBV, 2).
[0165] 104. Perform node search in the core event graph according to the reference summary word to obtain a target-related keyword associated with the reference summary word in the core event graph.
[0166] Among them, a node search is performed in the core event graph according to the reference abstract word, that is, in the core event graph, a target node connected to the node corresponding to the reference abstract word is found, and the keyword corresponding to the target node is the target associated keyword associated with the reference abstract word.
[0167] Specifically, the information search method of the present application is to cluster a given article into There are N articles in total. Summarize the core events and generate a concise text summary Y = (y 1 ,y 2 ,…,y M ). Among them, d 1 ,d 2 ,…,d N Represents each article (i.e., a text message), y 1 ,y 2 ,…,y M For each summary word that constitutes the text summary.
[0168] In some embodiments, if the current time (which can be regarded as the tth time) is the time when the tth summary word (which can be recorded as y) constituting the text summary is selected from the candidate summary words t ), the reference summary word can be the summary word selected as the text summary corresponding to the text information last time (which can be regarded as the t-1th time), recorded as y t-1 If there are k target-related keywords associated with the reference summary word in the core event graph, they are recorded as Then it can be considered that there are k event words associated with the reference summary word y t-1 , reference abstract word yt-1 The associated event set is
[0169] 105. Perform event distribution analysis on the weights of the target-related keywords and the keyword group containing the reference summary words to obtain event probability information.
[0170] Among them, in the core event graph, the target-related keywords are connected to the nodes corresponding to the reference summary words, and the weight of the keyword group in which the two are located is the weight corresponding to the edge connecting the nodes.
[0171] Optionally, in this embodiment, the step of “performing event distribution analysis on the weights of the target-related keywords and the keyword group containing the reference summary words to obtain event probability information” may include:
[0172] The weights of the keyword groups containing the reference summary words are combined to obtain initial event probability information;
[0173] Performing logistic regression processing on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight;
[0174] The initial event probability information and the context adaptive weight are fused to obtain event probability information.
[0175] There are many ways to combine the weights of each target-related keyword and the keyword group containing the reference abstract word, which is not limited in this embodiment. For example, the combination can be multiplication or weighted operation.
[0176] Specifically, the step of “merging the weights of each target-related keyword with the keyword group in which the reference summary word is located to obtain initial event probability information” may include:
[0177] Determining initial event sub-probability information corresponding to each target-related keyword based on the weight of each target-related keyword and the keyword group in which the reference summary word is located;
[0178] The initial event sub-probability information corresponding to each target-related keyword is fused to obtain the initial event probability information.
[0179] In some embodiments, if there is only one target-related keyword associated with the reference abstract word, the target-related keyword may be recorded as e. Indicates the target key word e and the reference summary word y t-1 The weight of the keyword group, the initial event probability information can be expressed as Wherein, N is the amount of text information, specifically the number of articles in the article cluster.
[0180] In some other embodiments, if there are multiple target-related keywords associated with the reference abstract word, these target-related keywords can be recorded as The initial event sub-probability information corresponding to each target-related keyword can be recorded as in, Then the initial event probability information can be expressed as
[0181] Optionally, in this embodiment, the step of “performing a logistic regression process on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight” may include:
[0182] The semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information are integrated to obtain a context-adaptive initial weight;
[0183] The context adaptive initial weight is normalized to obtain the context adaptive weight.
[0184] There are many ways to fuse, such as using a softmax classifier to fuse the semantic feature information corresponding to the target-related keywords, the latent state information of the candidate summary words, and the context feature information. The softmax function can convert the output values of multiple classifications into a probability distribution in the range of [0, 1].
[0185] Among them, logistic regression processing, also known as logistic regression analysis, is often used in data mining. Softmax regression is generally used in the output layer of a neural network, in which case the output layer is called a softmax layer.
[0186] The semantic feature information corresponding to the target-related keyword may specifically be the word vector of the target-related keyword. The context feature information may specifically be the context feature information corresponding to the target key text sentence.
[0187] The hidden state information of the candidate summary word is the implicit vector representation of the candidate summary word, which can represent the semantic information of the candidate summary word, and its influencing factor can include the hidden state information corresponding to the reference summary word.
[0188] In a specific embodiment, the semantic feature information corresponding to the target-related keyword can be recorded as (Specifically, it can be regarded as an event element vector), the hidden state information of the candidate summary word can be recorded as h i , the context feature information is recorded as ct , then the calculation process of context adaptive weight can be shown as formula (2):
[0189]
[0190] Where j∈[1,k], W h is a trainable parameter represents the contextual adaptation weight, It can be understood as the target-related keyword associated with the reference summary word at the tth time, which is represented as the input feature vector.
[0191] Specifically, the hidden state information h of the candidate summary word i It can be based on the hidden state information h of the reference summary word i-1 The word vector w corresponding to the candidate summary word itself i The calculation is as shown in formula (3):
[0192] h i =f LSTM (h i-1 ,w i ) (3)
[0193] Among them, f LSTM Indicates that the data is processed by the LSTM neural network model.
[0194] Optionally, in this embodiment, the step of “merging the initial event probability information and the context adaptive weight to obtain event probability information” may include:
[0195] Based on the preset weights, the initial event probability information and the context adaptive weights are weighted and fused to obtain event probability information.
[0196] The preset weights may be set according to actual conditions, and this embodiment does not impose any limitation on this.
[0197] By integrating context adaptive weights, the probability distribution corresponding to the event probability information can change with the context, that is, the relevant information of the context can be introduced in the process of generating text summaries.
[0198] Specifically, the initial event probability information can be regarded as the probability information based on the core event graph. By integrating the initial event probability information and the context adaptive weight, the total probability distribution of event elements at time t can be obtained, that is, the event probability information. The calculation process is shown in formula (4):
[0199]
[0200] in, represents the initial event probability information, represents the contextual adaptation weight, Represents event probability information. γ is an adjustable parameter, such as 0.2.
[0201] 106. Perform summary construction processing on the candidate summary words and the reference summary words according to the event probability information and the hidden state information of the reference summary words to obtain a text summary corresponding to the at least one text information, and output a search result according to the text summary.
[0202] Specifically, the candidate summary words may come from the segmented words in the text information, from the key text sentences of each text information, from the keywords in the core event graph, or from the preset word list, which is not limited in this embodiment. The preset word list can be set according to actual conditions.
[0203] Optionally, in this embodiment, each text message includes at least one text sentence;
[0204] The step of “performing summary construction processing on the candidate summary words and the reference summary words according to the event probability information and the hidden state information of the reference summary words to obtain a text summary corresponding to the at least one text information” may include:
[0205] Based on the keywords and the core event graph, selecting a target key text sentence from the text sentences corresponding to each text information;
[0206] Performing encoding and decoding processing on each word segment in the target key text sentence to obtain probability information of the target key text sentence;
[0207] A target summary word is selected from each candidate summary word according to the probability information of the target key text sentence, the event probability information and the hidden state information of the reference summary word, so as to perform summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information.
[0208] Among them, the target probability information corresponding to each candidate summary word can be calculated according to the probability information of the target key text sentence, the event probability information and the hidden state information of the reference summary word; and then based on the target probability information, the target summary word is selected from each candidate summary word.
[0209] Among them, the step of “calculating the target probability information corresponding to each candidate summary word according to the probability information of the target key text sentence, the event probability information and the hidden state information of the reference summary word” may include:
[0210] Extracting attention features of each candidate summary word according to the hidden state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word;
[0211] The target probability information corresponding to each candidate summary word is calculated according to the probability information of the target key text sentence, the event probability information and the attention feature information.
[0212] Specifically, the probability information of the target key text sentence, the event probability information and the attention feature information may be weighted and fused to obtain the target probability information corresponding to the candidate summary word.
[0213] Optionally, in this embodiment, the step of “selecting target key text sentences from text sentences corresponding to each text information based on the keywords and the core event graph” may include:
[0214] For each piece of text information, based on the keyword, selecting a key text sentence of the text information from each text sentence of the text information;
[0215] For each key text sentence corresponding to the text information, the word frequency parameters corresponding to the keywords contained in the key text sentence are merged to obtain the first important index of the key text sentence;
[0216] According to the phrases formed by each word segment in the key text sentence, a node search is performed in the core event graph to obtain the weight corresponding to the keyword phrase contained in the key text sentence;
[0217] The weights corresponding to the key word groups contained in the key text sentence are merged to obtain the second important index of the key text sentence;
[0218] Based on the first important indicator and the second important indicator, a target key text sentence is selected from the key text sentences corresponding to each text information.
[0219] The specific selection process of the step "selecting key text sentences from various text sentences of text information based on keywords" can refer to the description in the above embodiment, which will not be repeated here. The word frequency parameter corresponding to the keyword can be obtained by multiplying the word frequency TF and the inverse text frequency IDF of the keyword. The word frequency parameter can be represented by TF-IDF. The specific calculation process can refer to the description in step 101.
[0220] There are many ways to merge the word frequency parameters corresponding to the keywords contained in the key text sentence, which is not limited in this embodiment. For example, the fusion method may be addition.
[0221] Among them, every two words in the key text sentence can form a phrase, and the weight corresponding to the phrase in the key text sentence is determined according to the core event graph. Specifically, if the phrase does not belong to the keyword group, the weight of the phrase can be set to 0; if the phrase belongs to the keyword group, the node corresponding to the phrase can be searched in the core event graph, and the weight corresponding to the connection between the nodes is determined as the weight of the phrase.
[0222] There are many ways to merge the weights corresponding to the key word groups contained in the key text sentence, which is not limited in this embodiment. For example, the fusion method may be addition.
[0223] Specifically, the step of “selecting a target key text sentence from key text sentences corresponding to each text information based on the first important indicator and the second important indicator” may include:
[0224] The first important index and the second important index are combined to obtain a target important index of the key text sentence;
[0225] According to the target important indicators, the target key text sentences are selected from the key text sentences corresponding to each text information.
[0226] In some embodiments, the target importance index can be obtained by adding the first importance index and the second importance index. In this embodiment, the key text sentences whose target importance index is greater than a preset value can be selected as the target key text sentences, and the preset value can be set according to the actual situation; the key text sentences of each text information can also be sorted according to the target importance index, such as sorting from large to small, to obtain sorted key text sentences, and then the first n key text sentences in the sorted key text sentences are selected as the target key text sentences.
[0227] In a specific embodiment, after obtaining the core event graph and the key text sentences of each text information, the first similarity between each key text sentence and the keyword, and the second similarity between each key text sentence and the core event graph can be calculated. Based on the first similarity and the second similarity, the target importance index of each key text sentence is determined. Then, according to the target importance index, the target key text sentence that best expresses the core event of the article cluster is selected from each key text sentence. For each key text sentence, the target importance index is calculated as shown in formula (5):
[0228]
[0229] in, Indicates a key text sentence, kw D The keyword representing the article cluster, G e Represents the core event graph. It refers to the first similarity between the key text sentence and the keyword, which is specifically the first important indicator of the key text sentence described in the above embodiment; It is the second similarity between the key text sentence and the core event graph, and is also the second most important indicator of the key text sentence. Represents the target importance index of the key text sentence, that is, the importance score.
[0230] in, Indicates keyword w j The corresponding word frequency parameter, Represents the weight of the phrase corresponding to the j-th participle and the t-th participle in the key text sentence.
[0231] Among them, L represents the number of segmentations contained in the key text sentence, that is, the length of the sentence; λ is an adjustable parameter used to adjust the relationship between the first similarity and the second similarity; N represents the number of articles contained in the article cluster.
[0232] Among them, N t It refers to the key text sentence and the keyword set kw D The number of words in This is to prevent the algorithm from selecting long sentences.
[0233] After obtaining the importance scores of each key text sentence, the key text sentence with the highest score can be selected as the target key text sentence, and the target key text sentence can be regarded as the sentence most similar to the core event graph.
[0234] The information search method provided by the present application can first extract core events from multiple articles (text information) from which text summaries need to be extracted, specifically, identifying the core event element information in multiple articles, and then generating text summaries based on the identified core event element information. For the core event extraction stage, the present application can extract core events from two dimensions, which are the graph dimension and the sentence dimension. In the graph dimension, a core event graph can be constructed based on the keyword group extracted from the text information. The core event graph can be associated with the content in each text information, and the event element information contained is more comprehensive, but the semantic fluency of the core event graph is poor. In the sentence dimension, the target key text sentence can be selected from the text sentences corresponding to each text information based on the core event graph and keywords. The semantic integrity of the target key text sentence is strong, but it cannot cover the event element information between multiple text information. By combining the core event element information extracted from two dimensions, the present application can improve the semantic integrity of the extracted core events, and can also cover the content in each text information more comprehensively.
[0235] Optionally, in this embodiment, the step of “encoding and decoding each word in the target key text sentence to obtain probability information of the target key text sentence” may include:
[0236] Performing feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence;
[0237] The encoding information of each word segment in the target key text sentence is integrated to obtain context feature information;
[0238] Based on the current decoding parameters, the context feature information is decoded to obtain the probability information of the target key text sentence.
[0239] In some embodiments, a bidirectional LSTM can be used as an encoder and a unidirectional LSTM can be used as a decoder. The bidirectional LSTM is used to perform feature encoding on each word in the target key text sentence to obtain encoding information of each word; and the unidirectional LSTM is used to decode the context feature information.
[0240] Specifically, the neural network model can be used to perform encoding and decoding processing on each word segment in the target key text sentence to obtain the probability information of the target key text sentence.
[0241] Among them, the neural network model can be LSTM (Long Short-Term Memory), RNN (Recurrent Neural Network), GRU (Gate Recurrent Unit), transformer, etc., but it should be understood that the neural network model of this embodiment is not limited to the types listed above.
[0242] Optionally, in this embodiment, the step of “performing feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence” may include:
[0243] Based on the sequence of each segmentation in the target key text sentence, forward encoding each segmentation in the target key text sentence is performed to obtain forward encoding information of each segmentation in the target key text sentence;
[0244] Based on the sequence of each segmentation in the target key text sentence, each segmentation in the target key text sentence is backward encoded to obtain backward encoding information of each segmentation in the target key text sentence;
[0245] The forward encoding information and the backward encoding information are merged to obtain encoding information of each word segment in the target key text sentence.
[0246] Among them, the forward encoding of a certain word in the target key text sentence can be specifically based on other word segments that are located before the word segment in the target key text sentence, and the word segment is encoded to obtain the forward encoding information of the word segment. Specifically, the forward encoding information corresponding to the previous word segment of the word segment in the target key text sentence can be obtained, and the word segment is encoded according to the word vector of the word segment and the forward encoding information corresponding to the previous word segment to obtain the forward encoding information of the word segment.
[0247] Among them, the backward encoding of a certain word in the target key text sentence can be specifically based on other word segments in the target key text sentence that are located after the word segment, and the word segment is encoded to obtain the backward encoding information of the word segment. Specifically, the backward encoding information corresponding to the next word segment of the word segment in the target key text sentence can be obtained, and the word segment is encoded according to the word vector of the word segment and the backward encoding information corresponding to the next word segment to obtain the backward encoding information of the word segment.
[0248] There are many ways to fuse the forward coded information and the backward coded information, which are not limited in this embodiment. For example, the fusion method may be concatenation or weighted fusion.
[0249] In a specific embodiment, the target key text sentence contains N segmented words. Based on the above method, the forward encoding information (specifically, the forward latent vector) and the backward encoding information (specifically, the backward latent vector) of each segmented word in the target key text sentence can be generated. The forward encoding information of each segmented word can be recorded as The backward coded information can be recorded as By concatenating the forward encoding information and the backward encoding information, we can obtain the encoding information corresponding to each word segment (specifically, the hidden state information). For example, for the i-th word segment in the target key text sentence, its encoding information can be expressed as ).
[0250] Specifically, the calculation process of the forward encoding information or the backward encoding information of the word segmentation in the target key text sentence can be expressed as h i =f LSTM (h i-1 ,w i ). Among them, w i Indicates the current segment to be encoded in the target key text sentence. When the encoding of the segment is forward encoding, h i-1 Indicates the forward encoding information corresponding to the previous segmentation of the current segmentation to be encoded, h iIndicates the forward encoding information of the current word to be encoded. When the encoding of the word is backward encoding, h i-1 Indicates the backward encoding information corresponding to the next segmentation of the current segmentation to be encoded, h i Indicates the backward encoding information of the current word to be encoded.
[0251] Among them, the step of "merging the encoding information of each word in the target key text sentence to obtain context feature information" may include:
[0252] Obtain the weights corresponding to each word in the target key text sentence;
[0253] According to the weights, a weighted operation is performed on the encoding information of each word in the target key text sentence to obtain context feature information.
[0254] In a specific embodiment, the encoding information of each word in the target key text sentence can be recorded as h i , the corresponding weight can be recorded as a t,i , then the calculation process of context feature information can be shown as formula (6):
[0255]
[0256] Among them, c t represents context feature information, and N represents the number of words contained in the target key text sentence. Among them, the corresponding weight a t,i Specifically, it can be the attention feature information corresponding to the word segmentation in the target key text sentence, and its calculation process is shown in formula (7):
[0257] a t,i =softmax(v T tanh(W h h i +W d d t +b)) (7)
[0258] Among them, the softmax function is for normalization. Through the softmax function processing, a weight distribution mapped to the input position can be obtained. The tanh function, that is, the hyperbolic tangent, can be used as an activation function in neural networks in the field of deep learning.
[0259] Among them, d t Indicates the current decoding parameters, v, W h , W d , b is a trainable parameter, a t,i Represents the attention feature information corresponding to the i-th word in the target key text sentence at the t-th moment.
[0260] Among them, the step of "decoding the context feature information based on the current decoding parameters to obtain the probability information of the target key text sentence" may include:
[0261] Fusing the current decoding parameter and the context feature information to obtain fused context feature information;
[0262] The fused context feature information is normalized to obtain the probability information of the target key text sentence.
[0263] The fusion method of the current decoding parameter and the context feature information may be a splicing process, etc., which is not limited in this embodiment.
[0264] Among them, the current decoding parameter d t represents the decoding parameters at the current time t (specifically, the parameters in the decoder). The influencing factors of the current decoding parameters may include the decoding parameters d at the previous time (i.e., time t-1). t-1 , and reference abstract word y t-1 .
[0265] In some embodiments, the process of calculating the probability information of the target key text sentence at the tth time based on the context feature information by the decoder can be shown as formula (8):
[0266] P vocab (w) = P(y t |y <t ,S event ;θ)
[0267] =softmax(W 2 (W 1 [d t ,c t ]+b 1 )+b 2 ) (8)
[0268] Among them, d t is the implicit representation of the decoder at time t (specifically, the current decoding parameters), c t is the context feature information at time t, W 1 , W 2 , b 1 , b 2 is a trainable parameter.
[0269] Among them, θ can represent W 1 , W 2 , b 1 , b 2 Equal learning parameters, y <t Indicates the reference abstract word, S eventrepresents the target key text sentence; P vocab (w) represents the probability information of the target key text sentence.
[0270] It should be noted that in the training phase, the word input to the decoder is the word corresponding to the text summary. In actual application, the input to the decoder can be the summary word generated at the previous moment, that is, the reference summary word y obtained at time t-1 t-1 , according to the reference abstract word y t-1 , the decoding parameter d of the previous moment can be t-1 Update to get the current decoding parameter d t The maximum number of decoding steps can be set to 12, which is not limited in this embodiment.
[0271] Optionally, decoding can be performed using beam search, and the size of beam search can be set to 8, so that multiple candidate text summaries can be obtained to meet the requirements of different scenarios for the length, entity words, and fluency of the text summaries, thereby screening the text summaries required by the business scenarios according to actual business needs. In addition, text summaries with better quality can be selected from the candidate text summaries to achieve quality control.
[0272] Optionally, in this embodiment, the step of “performing summary construction processing on the candidate summary words and the reference summary words according to the event probability information and the hidden state information of the reference summary words to obtain a text summary corresponding to the at least one text information” may include:
[0273] Extracting attention features of each candidate summary word according to the hidden state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word;
[0274] Calculating target probability information corresponding to each candidate summary word according to the event probability information and the attention feature information;
[0275] Based on the target probability information, selecting a target summary word from each candidate summary word;
[0276] A summary construction process is performed on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information.
[0277] In some embodiments, the candidate summary words whose target probability information is greater than a preset value can be selected as the target summary words, and the preset value can be set according to the actual situation. In other embodiments, the candidate summary words can also be sorted according to the target probability information, such as sorting from large to small, to obtain sorted candidate summary words, and then the first n candidate summary words in the sorted candidate summary words are selected as the target summary words, and n can be set according to the actual situation, such as the fluency of the text summary to be generated, the length of the text summary, and other factors.
[0278] Specifically, the reference abstract words may include y 1 ,y 2 ,…,y t-1 , the selected target summary word can be recorded as y t , the summary construction process of the target summary word and the reference summary word can be specifically to concatenate the target summary word and the reference summary word to obtain a text summary. In a specific scenario, the text summary to be generated has a certain length requirement. For example, M summary words need to be generated at present. If the total number of reference summary words and target summary words is less than M, the reference summary word needs to be updated to y. 1 ,y 2 ,…,y t-1 ,y t , select a new target abstract word y from the candidate abstract words based on the reference abstract word t+1 The process can refer to the description in the above embodiment, until the total number of target abstract words and reference abstract words selected is not less than M, that is, the abstract word y is obtained. 1 ,y 2 ,…,y M , so as to concatenate these summary words and obtain the text summary Y=(y 1 ,y 2 ,…,y M ).
[0279] In one embodiment, the target probability information corresponding to each candidate summary word can be calculated based on the probability information of the target key text sentence, the event probability information and the attention feature information. Specifically, for each candidate summary word, the attention feature information corresponding to the candidate summary word, the probability information of the target key text sentence, and the event probability information can be fused, such as weighted fusion, to obtain the target probability information corresponding to the candidate summary word.
[0280] Specifically, the calculation process of the target probability information corresponding to the candidate summary word is shown in formula (9):
[0281]
[0282] Among them, P final(w) represents the target probability information corresponding to the candidate summary word w, P vocab (w) represents the probability information of the target key text sentence, a ti Represents attention feature information, Represents event probability information.
[0283] Among them, p gen The generation of the control text summary mainly depends on the target key text sentence or the core event graph. The calculation formula is shown in formula (10):
[0284]
[0285] Among them, σ is a sigmoid function, b gen is a trainable parameter. t is the implicit representation of the decoder at time t (specifically, the current decoding parameters), c t is the context feature information at time t.
[0286] Among them, the sigmoid function, that is, the S-shaped growth curve, can be used as an activation function in a neural network or in logistic regression processing to map variables to a numerical range from zero to one.
[0287] Optionally, in this embodiment, the step of “extracting attention features for each candidate summary word according to the latent state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word” may include:
[0288] For each candidate summary word, according to the hidden state information of the reference summary word, feature extraction is performed on the candidate summary word to obtain the hidden state information corresponding to the candidate summary word;
[0289] Logistic regression processing is performed on the latent state information corresponding to the candidate summary word to obtain the attention feature information corresponding to the candidate summary word.
[0290] Specifically, the hidden state information h of the candidate summary word i Based on the hidden state information h of the reference summary word i-1 The word vector w corresponding to the candidate summary word itself i The calculation is as shown in formula (11):
[0291] h i =f LSTM (h i-1 ,w i ) (11)
[0292] Among them, f LSTMIndicates that the data is processed by the LSTM neural network model.
[0293] Specifically, the hidden state information of each candidate summary word can be recorded as h i , the attention feature information of the candidate summary word can be recorded as a t,i , the calculation process can be shown as formula (12):
[0294] a t,i =softmax(v T tanh(W h h i +W d d t +b)) (12)
[0295] Among them, through the softmax function processing, a weight distribution mapped to the input position can be obtained. The tanh function, that is, the hyperbolic tangent, can be used as an activation function in neural networks in the field of deep learning.
[0296] Among them, d t Indicates the current decoding parameters, v, W h , W d , b is a trainable parameter, a t,i Represents the attention feature information corresponding to the candidate summary word at the tth moment.
[0297] In a specific embodiment, if Figure 1e As shown in the figure, it is an overall architecture diagram of the information search method. The information search method provided by the present application can search for at least one text information under the same topic. Specifically, multiple articles are summarized and summarized into core events to generate corresponding text summaries Y = (y 1 ,y 2 ,…,y M ). The text summary is obtained by selecting the target summary word from the candidate summary words for multiple times, and then fusing (such as splicing) the target summary words selected each time. For the generation process of each summary word in the text summary, the summary words that have been selected as components of the text summary can be used as reference summary words, and the target summary word is selected from the candidate summary words based on the reference summary words; for example, the target summary word y needs to be selected from the candidate summary words at present. t , then the reference abstract word can be y t-1 .
[0298] Among them, the generation process of each summary word in the text summary mainly includes two stages, namely the core event identification stage and the summary generation stage. In the core event identification stage, the core event identification in the graph dimension and sentence dimension can be carried out by constructing a core event graph, which can better locate the core event information, ensure the integrity and semantic coherence of event extraction, and generate a higher quality text summary; in the summary generation stage, the text summary can be generated through the event-oriented pointer network, integrating the core event graph and the target key text sentence information, which improves the accuracy of the generated text summary.
[0299] Specifically, in the core event identification stage, keywords of multiple text information can be extracted first, and then for each text information, based on the matching degree between keywords and each text sentence in the text information, key text sentences are selected from each text sentence in the text information; then, based on the semantic dependency between keywords and word segments in key text sentences, keyword groups are obtained, so as to construct a core event graph based on keyword groups. Then, based on keywords and the core event graph, target key text sentences S are selected from the key text sentences corresponding to each text information. event .
[0300] In the summary generation stage, the main process is to calculate the target probability information corresponding to each candidate summary word, and then select the target summary word from each candidate summary word according to the size of the target probability information, so as to perform summary construction based on the target summary word and the reference summary word.
[0301] Among them, the calculation of the target probability information of the candidate summary words can include three dimensions, namely: probability information obtained based on the target key text sentence, event probability information obtained based on the core event graph, and attention feature information obtained based on the candidate summary words; the probability information of the target key text sentence, event probability information, and attention feature information can be fused to obtain the target probability information of the candidate summary words.
[0302] The process of obtaining the probability information of the target key text sentence is as follows: event , through the bidirectional long short-term memory network to the target key text sentence S event Each word in the feature encoding is performed to obtain the target key text sentence S event The encoding information of each word in is recorded as h = (h 1 ,…,h i ), and then the target key text sentence S event The encoding information of each word in is fused to obtain the context feature information c t ; Thus, through the long short-term memory network based on the current decoding parameter d t, that is, the decoder state information, decodes the context feature information to obtain the probability information of the target key text sentence.
[0303] Among them, for the attention feature information a t,i , can be based on the reference abstract word y t-1 The hidden state information of the candidate summary words is extracted to obtain the hidden state information corresponding to the candidate summary words, and then the hidden state information is processed by logistic regression to obtain the attention feature information corresponding to the candidate summary words.
[0304] The process of obtaining event probability information can be as follows: t-1 Node search is performed in the core event graph to obtain target-related keywords (which can be regarded as candidate event elements) associated with the reference summary words in the core event graph, and the weights of the keyword groups where the target-related keywords and the reference summary words are located are fused to obtain the initial event probability information; then the semantic feature information corresponding to each target-related keyword, the hidden state information of the candidate summary words, and the context feature information c t Logistic regression processing is performed to obtain context-adaptive weights, thereby fusing the initial event probability information with the context-adaptive weights to obtain event probability information.
[0305] During the generation process of this application, by integrating attention feature information, on the one hand, the semantic coherence of the core sentence can be retained, and at the same time, by mapping the probability distribution of the core event graph, the global event information outside the sentence can be covered. Therefore, in the process of generating text summaries, coherence can be guaranteed while covering global event information.
[0306] In actual business scenarios, some scenarios, such as the display of hot event rankings, have high requirements for the accuracy of text summaries, so the quality of the generated text summaries can be tested. Specifically, the quality of text summaries can be tested through a trained quality control model, which can identify the quality of text summaries and score the generated text summaries, which can further improve the accuracy of the generated text summaries.
[0307] Among them, the quality control model can be a neural network model, and the neural network model can specifically include BERT (Bidirectional Encoder Representations from Transformers), etc., which is not limited in this embodiment.
[0308] It should be noted that the quality control model can be trained by other equipment and then provided to the information search device, or can also be trained by the information search device itself.
[0309] If the information search device performs training by itself, the information search method may further include:
[0310] Acquire training data, wherein the training data includes a sample text summary and an expected quality score corresponding to the sample text summary;
[0311] Extracting features of the sample text summary through a quality control model to obtain feature information of the sample text summary;
[0312] Predicting an actual quality score of the sample text summary based on the feature information of the sample text summary;
[0313] According to the actual quality score and the expected quality score, the parameters of the quality control model are adjusted to obtain a trained quality control model.
[0314] The training process is to first calculate the actual quality score of the sample text summary, and then use the back propagation algorithm to adjust the parameters of the quality control model. Based on the actual quality score and the expected quality score of the sample text summary, the parameters of the quality control model are optimized so that the actual quality score of the sample text summary approaches the expected quality score, thereby obtaining a trained quality control model. Specifically, the loss value between the actual quality score and the expected quality score of the sample text summary can be made smaller than a preset value, which can be set according to actual conditions.
[0315] This application can generate multiple text summaries of text information. Through the quality scoring of the quality control model, the text summary with the highest quality can be selected from these multiple text summaries, thereby improving the accuracy of the generated text summary in extracting key entity information of the article.
[0316] As can be seen from the above, this embodiment can obtain keywords corresponding to at least one text information under the same topic, reference abstract words of the text information, and at least one candidate abstract word; extract at least one keyword group from the keywords, and determine the weight corresponding to the keyword group, the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; construct a core event graph according to the keyword group and the weight corresponding to the keyword group, the core event graph includes nodes corresponding to the keywords, and lines between each node, the lines represent the association relationship between the keywords; perform node search in the core event graph according to the reference abstract word to obtain the target associated keyword associated with the reference abstract word in the core event graph; perform event distribution analysis on the weights of the target associated keyword and the keyword group where the reference abstract word is located to obtain event probability information; perform abstract construction processing on the candidate abstract words and the reference abstract words according to the event probability information and the hidden state information of the reference abstract word to obtain the text abstract corresponding to the at least one text information, and output the search result according to the text abstract. The embodiment of the present application can construct a core event graph through the keywords contained in the text information under the same topic, which can better refine the core event information, and then construct a text summary based on the core event graph, which is conducive to improving the generation efficiency and accuracy of the text summary, more accurately expressing the core events contained in the text information, and thus improving the output efficiency and accuracy of the search results.
[0317] According to the method described in the previous embodiment, the information search device will be further described in detail below by taking the example of being specifically integrated in a server.
[0318] The present application embodiment provides an information search method, such as Figure 2 As shown, the specific process of the information search method can be as follows:
[0319] 201. A server obtains keywords corresponding to at least one text message under the same topic, reference abstract words of the text message, and at least one candidate abstract word, wherein each text message includes at least one text sentence.
[0320] Specifically, each text information may be text information in an article, a paragraph, or a document. The reference summary word may be a summary word that has been selected as a text summary corresponding to the text information, or may be a reference word for selecting a summary word to form a text summary, which is not limited in this embodiment. The text summary may include at least one summary word.
[0321] This embodiment can select a target abstract word from candidate abstract words based on the reference abstract word, and then determine the text abstract corresponding to the text information under the same topic based on the target abstract word and the reference abstract word. Specifically, the reference abstract word can be updated based on the target abstract word, such as adding the target abstract word to the word set corresponding to the reference abstract word to obtain the updated word set corresponding to the reference abstract word; then obtain new candidate abstract words, which may not contain the selected target abstract word compared to the original candidate abstract words, so that based on the word set corresponding to the updated reference abstract word, a new target abstract word is selected from the new candidate abstract words, so that multiple target abstract words can be obtained to construct a text abstract.
[0322] This application can identify core events for multiple articles describing the same topic, and summarize the description in refined language. Specifically, the information search method provided by this application can be divided into two stages, namely: the core event identification stage and the summary generation stage (i.e., the short description generation stage). Among them, in the core event identification stage, the core event identification at the graph granularity and sentence granularity can be performed by constructing a core event graph, so that the core event information can be better located, the integrity and semantic coherence of the event extraction can be ensured, and a higher quality text summary can be generated; in the summary generation stage, the core event graph and the information of the target key text sentences can be fused to generate a text summary, which improves the accuracy of the generated text summary. Applying this information search method in search-related scenarios can improve multiple types of online indicators.
[0323] 202. For each piece of text information, the server selects a key text sentence of the text information from the text sentences of the text information based on the matching degree between the keyword and the text sentences of the text information.
[0324] Optionally, in this embodiment, the step of “for each text information, based on the matching degree between the keyword and each text sentence of the text information, selecting a key text sentence of the text information from each text sentence of the text information” may include:
[0325] For each text sentence in each text information, counting the number of the keywords contained in the text sentence;
[0326] Based on the number, determining a degree of match between the text sentence and the keyword;
[0327] According to the matching degree, a key text sentence of the text information is selected from each text sentence of the text information.
[0328] Optionally, in some embodiments, the step of “for each text message, based on the matching degree between the keyword and each text sentence of the text message, selecting a key text sentence of the text message from each text sentence of the text message” may include:
[0329] For each text message, based on the matching degree between the keyword and each body text sentence of the text message, selecting a body key text sentence of the text message from each body text sentence of the text message;
[0330] The key text sentence of the text information is obtained according to the key text sentence of the main text and the title sentence corresponding to the text information.
[0331] Among them, both the main text key sentence and the title sentence can be used as the key text sentence of the text information.
[0332] 203. The server selects at least one target keyword from the segmented words corresponding to the key text sentences of each text information based on the keywords and the semantic dependency between the segmented words in the key text sentences of each text information; and constructs at least one keyword group according to the semantic dependency between the target keywords, wherein the keyword group includes at least two target keywords having an associated relationship.
[0333] The association relationship between the keywords in the keyword group may specifically include the semantic dependency relationship between the keywords.
[0334] Specifically, every two target keywords having a semantic dependency relationship may be regarded as a keyword group, that is, each keyword group includes two target keywords having a semantic dependency relationship.
[0335] Optionally, in this embodiment, the step of “selecting at least one target keyword from the segmented words corresponding to the key text sentences of each text information based on the keyword and the semantic dependency between the segmented words in the key text sentences of each text information” may include:
[0336] For each key text sentence of the text information, a syntactic structure analysis is performed on the key text sentence to determine the semantic dependency relationship between each word segment in the key text sentence;
[0337] Extracting a first target keyword of the key text sentence from each word segment of the key text sentence according to the semantic dependency;
[0338] Based on the keywords included in the key text sentence, obtaining a second target keyword of the key text sentence;
[0339] According to the first target keyword and the second target keyword, a target keyword corresponding to the key text sentence is determined.
[0340] Among them, the semantic dependency relationship between participles may include subject-predicate relationship, verb-object relationship, indirect object relationship, attributive-predicate relationship, etc., which is not limited in this embodiment.
[0341] Among them, the participles in the key text sentence can be deleted according to the semantic dependency relationship between the participles in the key text sentence. Specifically, in some embodiments, the participles whose semantic dependency relationship meets the preset relationship can be retained. For example, the participles whose semantic dependency relationship is the subject-predicate relationship, the verb-object relationship, and the intermezzo-object relationship can be retained, and these participles are used as the first target keywords of the target key text sentence.
[0342] Specifically, the step of "obtaining the second target keyword of the key text sentence based on the keywords contained in the key text sentence" may include: taking the participles in the key text sentence that have a semantic dependency relationship with the keywords and the keywords contained in the key text sentence as the second target keywords of the key text sentence.
[0343] 204. The server determines a weight corresponding to the keyword group, where the weight represents the number of times the keyword group appears in the at least one text message.
[0344] In some embodiments, the weight corresponding to the keyword group can be counted by co-occurrence, and the weight corresponding to the keyword group can characterize the number of times the keyword group appears in the key text sentences of each text information. For example, the number of sentences of the key text sentences containing a certain keyword group can be used as the weight corresponding to the keyword group. Specifically, when the target keyword included in the keyword group appears in a key text sentence at the same time, and the correlation relationship between the target keywords in the key text sentence is consistent with the correlation relationship between the target keywords in the keyword group, it can be regarded as that the key text sentence contains this keyword group.
[0345] 205. The server constructs a core event graph according to the keyword groups and the weights corresponding to the keyword groups. The core event graph includes nodes corresponding to target keywords and lines between the nodes. The lines represent associations between the target keywords.
[0346] Among them, the nodes of the core event graph are keywords rather than sentences. This can better model the association between event dimensions between different sentences and better explore the connection between event information between different articles. Moreover, through keyword extraction, named entity recognition, semantic dependency analysis and other technologies, the constructed core event graph can basically cover important event element information, which can provide a basis for the selection of target key text sentences.
[0347] Specifically, the connection lines between nodes can represent the semantic dependency relationship between the keywords corresponding to the nodes, and the weights corresponding to the keyword groups composed of the keywords.
[0348] 206. The server performs a node search in the core event graph according to the reference abstract word to obtain a target-related keyword associated with the reference abstract word in the core event graph.
[0349] Among them, a node search is performed in the core event graph according to the reference abstract word, that is, in the core event graph, a target node connected to the node corresponding to the reference abstract word is found, and the keyword corresponding to the target node is the target associated keyword associated with the reference abstract word.
[0350] 207. The server performs event distribution analysis on the weights of the target-related keyword and the keyword group in which the reference summary word is located to obtain event probability information.
[0351] Optionally, in this embodiment, the step of “performing event distribution analysis on the weights of the target-related keywords and the keyword group containing the reference summary words to obtain event probability information” may include:
[0352] The weights of the keyword groups containing the reference summary words are combined to obtain initial event probability information;
[0353] Performing logistic regression processing on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight;
[0354] The initial event probability information and the context adaptive weight are fused to obtain event probability information.
[0355] Specifically, the step of “merging the weights of each target-related keyword with the keyword group in which the reference summary word is located to obtain initial event probability information” may include:
[0356] Determining initial event sub-probability information corresponding to each target-related keyword based on the weight of each target-related keyword and the keyword group in which the reference summary word is located;
[0357] The initial event sub-probability information corresponding to each target-related keyword is fused to obtain the initial event probability information.
[0358] Optionally, in this embodiment, the step of “performing a logistic regression process on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight” may include:
[0359] The semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information are integrated to obtain a context-adaptive initial weight;
[0360] The context adaptive initial weight is normalized to obtain the context adaptive weight.
[0361] There are many ways to fuse, such as using a softmax classifier to fuse the semantic feature information corresponding to the target-related keywords, the latent state information of the candidate summary words, and the context feature information. The softmax function can convert the output values of multiple classifications into a probability distribution in the range of [0, 1].
[0362] By integrating context adaptive weights, the probability distribution corresponding to the event probability information can change with the context, that is, the relevant information of the context can be introduced in the process of generating text summaries.
[0363] 208. The server selects a target key text sentence from key text sentences corresponding to each text information based on the keyword and the core event graph.
[0364] Optionally, in this embodiment, the step of “selecting a target key text sentence from key text sentences corresponding to each text information based on the keyword and the core event graph” may include:
[0365] For each key text sentence corresponding to the text information, the word frequency parameters corresponding to the keywords contained in the key text sentence are merged to obtain the first important index of the key text sentence;
[0366] According to the phrases formed by each word segment in the key text sentence, a node search is performed in the core event graph to obtain the weight corresponding to the keyword phrase contained in the key text sentence;
[0367] The weights corresponding to the key word groups contained in the key text sentence are merged to obtain the second important index of the key text sentence;
[0368] Based on the first important indicator and the second important indicator, a target key text sentence is selected from the key text sentences corresponding to each text information.
[0369] Among them, every two words in the key text sentence can form a phrase, and the weight corresponding to the phrase in the key text sentence is determined according to the core event graph. Specifically, if the phrase does not belong to the keyword group, the weight of the phrase can be set to 0; if the phrase belongs to the keyword group, the node corresponding to the phrase can be searched in the core event graph, and the weight corresponding to the connection between the nodes is determined as the weight of the phrase.
[0370] Specifically, the step of “selecting a target key text sentence from key text sentences corresponding to each text information based on the first important indicator and the second important indicator” may include:
[0371] The first important index and the second important index are combined to obtain a target important index of the key text sentence;
[0372] According to the target important indicators, the target key text sentences are selected from the key text sentences corresponding to each text information.
[0373] The information search method provided by the present application can first extract core events from multiple articles (text information) from which text summaries need to be extracted, specifically, identifying the core event element information in multiple articles, and then generating text summaries based on the identified core event element information. For the core event extraction stage, the present application can extract core events from two dimensions, which are the graph dimension and the sentence dimension. In the graph dimension, a core event graph can be constructed based on the keyword group extracted from the text information. The core event graph can be associated with the content in each text information, and the event element information contained is more comprehensive, but the semantic fluency of the core event graph is poor. In the sentence dimension, the target key text sentence can be selected from the text sentences corresponding to each text information based on the core event graph and keywords. The semantic integrity of the target key text sentence is strong, but it cannot cover the event element information between multiple text information. By combining the core event element information extracted from two dimensions, the present application can improve the semantic integrity of the extracted core events, and can also cover the content in each text information more comprehensively.
[0374] 209. The server performs encoding and decoding processing on each word segment in the target key text sentence to obtain probability information of the target key text sentence.
[0375] Optionally, in this embodiment, the step of “encoding and decoding each word in the target key text sentence to obtain probability information of the target key text sentence” may include:
[0376] Performing feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence;
[0377] The encoding information of each word segment in the target key text sentence is integrated to obtain context feature information;
[0378] Based on the current decoding parameters, the context feature information is decoded to obtain the probability information of the target key text sentence.
[0379] In some embodiments, a bidirectional LSTM can be used as an encoder and a unidirectional LSTM can be used as a decoder. The bidirectional LSTM is used to perform feature encoding on each word in the target key text sentence to obtain encoding information of each word; and the unidirectional LSTM is used to decode the context feature information.
[0380] Optionally, in this embodiment, the step of “performing feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence” may include:
[0381] Based on the sequence of each segmentation in the target key text sentence, forward encoding each segmentation in the target key text sentence is performed to obtain forward encoding information of each segmentation in the target key text sentence;
[0382] Based on the sequence of each segmentation in the target key text sentence, each segmentation in the target key text sentence is backward encoded to obtain backward encoding information of each segmentation in the target key text sentence;
[0383] The forward encoding information and the backward encoding information are merged to obtain encoding information of each word segment in the target key text sentence.
[0384] There are many ways to fuse the forward coded information and the backward coded information, which are not limited in this embodiment. For example, the fusion method may be concatenation or weighted fusion.
[0385] Among them, the step of "decoding the context feature information based on the current decoding parameters to obtain the probability information of the target key text sentence" may include:
[0386] Fusing the current decoding parameter and the context feature information to obtain fused context feature information;
[0387] The fused context feature information is normalized to obtain the probability information of the target key text sentence.
[0388] 2010. The server extracts attention features of each candidate summary word according to the hidden state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word.
[0389] Specifically, the candidate summary words may come from the segmented words in the text information, from the key text sentences of each text information, from the keywords in the core event graph, or from the preset word list, which is not limited in this embodiment. The preset word list can be set according to actual conditions.
[0390] Optionally, in this embodiment, the step of “extracting attention features for each candidate summary word according to the latent state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word” may include:
[0391] For each candidate summary word, according to the hidden state information of the reference summary word, feature extraction is performed on the candidate summary word to obtain the hidden state information corresponding to the candidate summary word;
[0392] Logistic regression processing is performed on the latent state information corresponding to the candidate summary word to obtain the attention feature information corresponding to the candidate summary word.
[0393] 2011. The server calculates target probability information corresponding to each candidate summary word according to the probability information of the target key text sentence, the event probability information and the attention feature information.
[0394] In one embodiment, the target probability information corresponding to each candidate summary word can be calculated based on the probability information of the target key text sentence, the event probability information and the attention feature information. Specifically, for each candidate summary word, the attention feature information corresponding to the candidate summary word, the probability information of the target key text sentence, and the event probability information can be fused, such as weighted fusion, to obtain the target probability information corresponding to the candidate summary word.
[0395] 2012. The server selects a target summary word from each candidate summary word based on the target probability information; performs summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information, and outputs a search result according to the text summary.
[0396] In some embodiments, the candidate summary words whose target probability information is greater than a preset value can be selected as the target summary words, and the preset value can be set according to the actual situation. In other embodiments, the candidate summary words can also be sorted according to the target probability information, such as sorting from large to small, to obtain sorted candidate summary words, and then the first n candidate summary words in the sorted candidate summary words are selected as the target summary words, and n can be set according to the actual situation, such as the fluency of the text summary to be generated, the length of the text summary, and other factors.
[0397] As can be seen from the above, this embodiment can obtain keywords corresponding to at least one text information under the same topic, reference abstract words of the text information, and at least one candidate abstract word through a server, wherein each text information includes at least one text sentence; for each text information, based on the matching degree between the keywords and each text sentence of the text information, select the key text sentence of the text information from each text sentence of the text information; based on the keywords and the semantic dependency between the segmented words in the key text sentences of each text information, select at least one target keyword from the segmented words corresponding to the key text sentences of each text information; according to the semantic dependency between each target keyword, construct at least one keyword group, wherein the keyword group includes at least two target keywords with an associated relationship; determine the weight corresponding to the keyword group, wherein the weight represents the number of times the keyword group appears in the at least one text information. The server constructs a core event graph according to the keyword group and the weight corresponding to the keyword group, wherein the core event graph includes nodes corresponding to the target keywords and lines between each node, wherein the lines represent the associated relationship between the target keywords. The server performs node search in the core event graph according to the reference summary word to obtain the target-related keyword associated with the reference summary word in the core event graph; performs event distribution analysis on the weights of the target-related keyword and the keyword group in which the reference summary word is located to obtain event probability information. The server selects a target key text sentence from the key text sentences corresponding to each text information based on the keyword and the core event graph; performs encoding and decoding processing on each word segment in the target key text sentence to obtain probability information of the target key text sentence. According to the hidden state information of the reference summary word, the attention feature extraction is performed on each candidate summary word to obtain attention feature information corresponding to each candidate summary word. The server calculates the target probability information corresponding to each candidate summary word according to the probability information of the target key text sentence, the event probability information and the attention feature information; selects the target summary word from each candidate summary word based on the target probability information; performs summary construction processing on the target summary word and the reference summary word to obtain the text summary corresponding to the at least one text information, and outputs the search result according to the text summary.
[0398] The embodiment of the present application can construct a core event graph through the keywords contained in the text information under the same topic, which can better refine the core event information, and then construct a text summary based on the core event graph, which is conducive to improving the generation efficiency and accuracy of the text summary, more accurately expressing the core events contained in the text information, and thus improving the output efficiency and accuracy of the search results.
[0399] In order to better implement the above method, the present application embodiment also provides an information search device, such as Figure 3 As shown, the information search device may include an acquisition unit 301, a phrase extraction unit 302, an event graph construction unit 303, a node search unit 304, an analysis unit 305, and a summary construction unit 306, as follows:
[0400] (1) Acquisition unit 301;
[0401] The acquisition unit 301 is configured to acquire keywords corresponding to at least one text message under the same topic, a reference abstract word of the text message, and at least one candidate abstract word.
[0402] Optionally, in some embodiments of the present application, the acquisition unit may include a word segmentation subunit, an entity recognition subunit, a frequency analysis subunit, and a keyword selection subunit, as follows:
[0403] The word segmentation subunit is used to perform word segmentation processing on at least one text information under the same topic to obtain at least one word segmentation of the text information;
[0404] An entity recognition subunit, configured to perform entity recognition on each of the segmented words of the text information, so as to determine at least one entity segmented word from each of the segmented words of the text information;
[0405] A frequency analysis subunit, used to perform frequency analysis on each entity segmentation in the text information to obtain a word frequency parameter of each entity segmentation in the text information;
[0406] The keyword selection subunit is used to select keywords of the text information from the entity segmentation based on the word frequency parameter, and obtain a reference summary word and at least one candidate summary word of the text information.
[0407] Optionally, in some embodiments of the present application, the frequency analysis subunit can be specifically used to count the frequency of occurrence of each entity segmentation in the text information to obtain the weight of the entity segmentation in the text information; count the frequency of occurrence of the entity segmentation in sample text information to obtain the reference weight of the entity segmentation; determine the word frequency parameter of the entity segmentation based on the reference weight of the entity segmentation and its weight in the text information.
[0408] (2) phrase extraction unit 302;
[0409] The phrase extraction unit 302 is used to extract at least one keyword group from the keywords and determine the weight corresponding to the keyword group, wherein the keyword group includes at least two associated keywords, and the weight represents the number of times the keyword group appears in the at least one text information.
[0410] Optionally, in some embodiments of the present application, each text message includes at least one text sentence;
[0411] The phrase extraction unit may include a text sentence selection subunit, a target keyword selection subunit, and a phrase construction subunit, as follows:
[0412] The text sentence selection subunit is used to select a key text sentence of the text information from each text sentence of the text information based on the matching degree between the keyword and each text sentence of the text information for each text information;
[0413] A target keyword selection subunit is used to select at least one target keyword from the segmented words corresponding to the key text sentences of each text information based on the keywords and the semantic dependency between the segmented words in the key text sentences of each text information;
[0414] The phrase construction subunit is used to construct at least one keyword group according to the semantic dependency relationship between each target keyword.
[0415] Optionally, in some embodiments of the present application, the text sentence selection subunit can be specifically used to count the number of the keywords contained in each text sentence in each text information; based on the number, determine the matching degree between the text sentence and the keyword; and based on the matching degree, select the key text sentence of the text information from each text sentence in the text information.
[0416] Optionally, in some embodiments of the present application, the target keyword selection subunit can be specifically used to perform syntactic structure analysis on the key text sentence of each text information to determine the semantic dependency relationship between each word segment in the key text sentence; based on the semantic dependency relationship, extract the first target keyword of the key text sentence from each word segment of the key text sentence; based on the keywords contained in the key text sentence, obtain the second target keyword of the key text sentence; and determine the target keyword corresponding to the key text sentence based on the first target keyword and the second target keyword.
[0417] (3) event graph construction unit 303;
[0418] The event graph construction unit 303 is used to construct a core event graph according to the keyword groups and the weights corresponding to the keyword groups. The core event graph includes nodes corresponding to the keywords and lines between the nodes, and the lines represent the association relationship between the keywords.
[0419] (4) Node search unit 304;
[0420] The node search unit 304 is configured to perform a node search in the core event graph according to the reference summary word to obtain a target associated keyword associated with the reference summary word in the core event graph.
[0421] (5) analysis unit 305;
[0422] The analyzing unit 305 is configured to perform event distribution analysis on the weights of the target-related keywords and the keyword group in which the reference summary words are located to obtain event probability information.
[0423] Optionally, in some embodiments of the present application, the analysis unit may include a first fusion subunit, a logistic regression subunit, and a second fusion subunit, as follows:
[0424] The first fusion subunit is used to fuse the weights of each target-related keyword with the keyword group where the reference summary word is located to obtain initial event probability information;
[0425] The logistic regression subunit is used to perform logistic regression processing on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight;
[0426] The second fusion subunit is used to fuse the initial event probability information and the context adaptive weight to obtain event probability information.
[0427] (6) summary construction unit 306;
[0428] The summary construction unit 306 is used to perform summary construction processing on the candidate summary words and the reference summary words according to the event probability information and the hidden state information of the reference summary words, obtain a text summary corresponding to the at least one text information, and output a search result according to the text summary.
[0429] Optionally, in some embodiments of the present application, each text message includes at least one text sentence;
[0430] The summary construction unit may include a selection subunit, a coding subunit and a summary construction subunit, as follows:
[0431] The selection subunit is used to select a target key text sentence from the text sentences corresponding to each text information based on the keyword and the core event graph;
[0432] The encoding and decoding subunit is used to perform encoding and decoding processing on each word segment in the target key text sentence to obtain probability information of the target key text sentence;
[0433] The summary construction subunit is used to select a target summary word from each candidate summary word according to the probability information of the target key text sentence, the event probability information and the hidden state information of the reference summary word, so as to perform summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information.
[0434] Optionally, in some embodiments of the present application, the selection subunit can be specifically used to select, for each text information, a key text sentence of the text information from each text sentence of the text information based on the keyword; for the key text sentence corresponding to each text information, the word frequency parameters corresponding to the keyword contained in the key text sentence are merged to obtain a first important indicator of the key text sentence; according to the phrases formed by each word segment in the key text sentence, a node search is performed in the core event graph to obtain the weight corresponding to the keyword group contained in the key text sentence; the weights corresponding to each keyword group contained in the key text sentence are merged to obtain a second important indicator of the key text sentence; based on the first important indicator and the second important indicator, a target key text sentence is selected from the key text sentences corresponding to each text information.
[0435] Optionally, in some embodiments of the present application, the encoding and decoding subunit can be specifically used to perform feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence; fuse the encoding information of each word in the target key text sentence to obtain context feature information; and decode the context feature information based on current decoding parameters to obtain probability information of the target key text sentence.
[0436] Optionally, in some embodiments of the present application, the step of “performing feature encoding on each segmentation in the target key text sentence to obtain encoding information of each segmentation in the target key text sentence” may include:
[0437] Based on the sequence of each segmentation in the target key text sentence, forward encoding each segmentation in the target key text sentence is performed to obtain forward encoding information of each segmentation in the target key text sentence;
[0438] Based on the sequence of each segmentation in the target key text sentence, each segmentation in the target key text sentence is backward encoded to obtain backward encoding information of each segmentation in the target key text sentence;
[0439] The forward encoding information and the backward encoding information are merged to obtain encoding information of each word segment in the target key text sentence.
[0440] Optionally, in some embodiments of the present application, the summary construction unit may include an attention feature extraction subunit, a calculation subunit, a summary word selection subunit and a construction subunit, as follows:
[0441] The attention feature extraction subunit is used to extract attention features of each candidate summary word according to the hidden state information of the reference summary word, so as to obtain attention feature information corresponding to each candidate summary word;
[0442] A calculation subunit, used to calculate target probability information corresponding to each candidate summary word according to the event probability information and the attention feature information;
[0443] A summary word selection subunit, configured to select a target summary word from each candidate summary word based on the target probability information;
[0444] The construction subunit is used to perform summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information.
[0445] Optionally, in some embodiments of the present application, the attention feature extraction subunit can be specifically used to perform feature extraction on each candidate summary word according to the hidden state information of the reference summary word to obtain the hidden state information corresponding to the candidate summary word; perform logistic regression processing on the hidden state information corresponding to the candidate summary word to obtain the attention feature information corresponding to the candidate summary word.
[0446] As can be seen from the above, in this embodiment, the acquisition unit 301 can acquire keywords corresponding to at least one text information under the same topic, the reference summary word of the text information and at least one candidate summary word; the phrase extraction unit 302 extracts at least one keyword group from the keywords, and determines the weight corresponding to the keyword group, wherein the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; the event graph construction unit 303 constructs a core event graph according to the keyword group and the weight corresponding to the keyword group, wherein the core event graph includes nodes corresponding to the keywords and the connections between the nodes. The connection line represents the association relationship between keywords; the node search unit 304 performs a node search in the core event graph according to the reference summary word to obtain the target associated keyword associated with the reference summary word in the core event graph; the analysis unit 305 performs event distribution analysis on the weights of the target associated keyword and the keyword group where the reference summary word is located to obtain event probability information; the summary construction unit 306 performs summary construction processing on the candidate summary word and the reference summary word according to the event probability information and the hidden state information of the reference summary word to obtain the text summary corresponding to the at least one text information, and outputs the search result according to the text summary. The embodiment of the present application can construct a core event graph through the keywords contained in the text information under the same topic, so that the core event information can be better refined, and then the text summary is constructed based on the core event graph, which is conducive to improving the generation efficiency and accuracy of the text summary, more accurately expressing the core events contained in the text information, and then improving the output efficiency and accuracy of the search results.
[0447] The present application also provides an electronic device, such as Figure 4 As shown, it shows a schematic diagram of the structure of an electronic device involved in an embodiment of the present application, and the electronic device may be a terminal or a server, etc. Specifically:
[0448] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will appreciate that Figure 4 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0449] The processor 401 is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, the processor 401 performs various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 401.
[0450] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0451] The electronic device also includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 403 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.
[0452] The electronic device may further include an input unit 404, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0453] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, thereby realizing various functions, as follows:
[0454] The method comprises the steps of: obtaining keywords corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word; extracting at least one keyword group from the keywords, and determining a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; constructing a core event graph according to the keyword group and the weight corresponding to the keyword group, wherein the core event graph includes nodes corresponding to the keywords and lines between the nodes, wherein the lines represent the associated relationship between the keywords; performing a node search in the core event graph according to the reference abstract word to obtain a target associated keyword associated with the reference abstract word in the core event graph; performing event distribution analysis on the weights of the target associated keyword and the keyword group where the reference abstract word is located to obtain event probability information; performing abstract construction processing on the candidate abstract word and the reference abstract word according to the event probability information and the hidden state information of the reference abstract word to obtain a text abstract corresponding to the at least one text information, and outputting a search result according to the text abstract.
[0455] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0456] As can be seen from the above, this embodiment can obtain keywords corresponding to at least one text information under the same topic, reference abstract words of the text information, and at least one candidate abstract word; extract at least one keyword group from the keywords, and determine the weight corresponding to the keyword group, the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; construct a core event graph according to the keyword group and the weight corresponding to the keyword group, the core event graph includes nodes corresponding to the keywords, and lines between each node, the lines represent the association relationship between the keywords; perform node search in the core event graph according to the reference abstract word to obtain the target associated keyword associated with the reference abstract word in the core event graph; perform event distribution analysis on the weights of the target associated keyword and the keyword group where the reference abstract word is located to obtain event probability information; perform abstract construction processing on the candidate abstract words and the reference abstract words according to the event probability information and the hidden state information of the reference abstract word to obtain the text abstract corresponding to the at least one text information, and output the search result according to the text abstract. The embodiment of the present application can construct a core event graph through the keywords contained in the text information under the same topic, which can better refine the core event information, and then construct a text summary based on the core event graph, which is conducive to improving the generation efficiency and accuracy of the text summary, more accurately expressing the core events contained in the text information, and thus improving the output efficiency and accuracy of the search results.
[0457] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0458] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the information search methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0459] The method comprises the steps of: obtaining keywords corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word; extracting at least one keyword group from the keywords, and determining a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords with an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; constructing a core event graph according to the keyword group and the weight corresponding to the keyword group, wherein the core event graph includes nodes corresponding to the keywords and lines between the nodes, wherein the lines represent the associated relationship between the keywords; performing a node search in the core event graph according to the reference abstract word to obtain a target associated keyword associated with the reference abstract word in the core event graph; performing event distribution analysis on the weights of the target associated keyword and the keyword group where the reference abstract word is located to obtain event probability information; performing abstract construction processing on the candidate abstract word and the reference abstract word according to the event probability information and the hidden state information of the reference abstract word to obtain a text abstract corresponding to the at least one text information, and outputting a search result according to the text abstract.
[0460] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0461] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0462] Since the instructions stored in the computer-readable storage medium can execute the steps in any information search method provided in the embodiments of the present application, the beneficial effects that can be achieved by any information search method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0463] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional implementations of the above-mentioned information search aspect.
[0464] The above is a detailed introduction to an information search method and related equipment provided in an embodiment of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An information search method, It is characterized in that include: Acquire a keyword corresponding to at least one text information under the same topic, a reference summary word of the text information, and at least one candidate summary word, wherein the reference summary word is a summary word that has been selected to constitute a text summary corresponding to the text information, or the reference summary word is a reference word for selecting a summary word to constitute a text summary; Extracting at least one keyword group from the keywords, and determining a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords having an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; According to the keyword groups and the weights corresponding to the keyword groups, a core event graph is constructed, wherein the core event graph includes nodes corresponding to the keywords and lines between the nodes, wherein the lines represent associations between the keywords; Performing a node search in the core event graph according to the reference abstract word to obtain a target-related keyword associated with the reference abstract word in the core event graph; Performing event distribution analysis on the weights of the target-related keywords and the keyword group containing the reference summary words to obtain event probability information; Selecting a target summary word from each candidate summary word according to the event probability information and the hidden state information of the reference summary word; The target summary word and the reference summary word are subjected to summary construction processing to obtain a text summary corresponding to the at least one text information, and a search result is output according to the text summary.
2. The method according to claim 1, It is characterized in that The step of obtaining a keyword corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word includes: Performing word segmentation processing on at least one text information under the same topic to obtain at least one word segmentation of the text information; Performing entity recognition on each of the segmented words of the text information to determine at least one entity segmented word from each of the segmented words of the text information; Performing frequency analysis on each entity segmentation in the text information to obtain a word frequency parameter of each entity segmentation in the text information; Based on the word frequency parameter, keywords of the text information are selected from the entity segmentation words, and a reference summary word and at least one candidate summary word of the text information are obtained.
3. The method according to claim 2, It is characterized in that The performing frequency analysis on each entity segmentation of the text information to obtain a word frequency parameter of each entity segmentation in the text information includes: For each entity participle of the text information, counting the frequency of occurrence of the entity participle in the text information, and obtaining the weight of the entity participle in the text information; Counting the frequency of occurrence of the entity segmentation in the sample text information to obtain a reference weight of the entity segmentation; A word frequency parameter of the entity segmentation is determined according to the reference weight of the entity segmentation and the weight in the text information.
4. The method according to claim 1, It is characterized in that Each of the text messages includes at least one text sentence; The step of extracting at least one keyword group from the keywords comprises: For each piece of text information, based on the matching degree between the keyword and each text sentence of the text information, selecting a key text sentence of the text information from each text sentence of the text information; Based on the keywords and the semantic dependency between the segmented words in the key text sentences of each text information, selecting at least one target keyword from the segmented words corresponding to the key text sentences of each text information; At least one keyword group is constructed according to the semantic dependency relationship between the target keywords.
5. The method according to claim 4, It is characterized in that The step of selecting, for each text information, a key text sentence of the text information from each text sentence of the text information based on the matching degree between the keyword and each text sentence of the text information, comprises: For each text sentence in each text information, counting the number of the keywords contained in the text sentence; Based on the number, determining a degree of match between the text sentence and the keyword; According to the matching degree, a key text sentence of the text information is selected from each text sentence of the text information.
6. The method according to claim 4, It is characterized in that The selecting at least one target keyword from the segmented words corresponding to the key text sentences of each text information based on the keyword and the semantic dependency between the segmented words in the key text sentences of each text information comprises: For each key text sentence of the text information, a syntactic structure analysis is performed on the key text sentence to determine the semantic dependency relationship between each word segment in the key text sentence; Extracting a first target keyword of the key text sentence from each word segment of the key text sentence according to the semantic dependency; Based on the keywords included in the key text sentence, obtaining a second target keyword of the key text sentence; According to the first target keyword and the second target keyword, a target keyword corresponding to the key text sentence is determined.
7. The method according to claim 1, It is characterized in that Each of the text messages includes at least one text sentence; The step of selecting a target summary word from each candidate summary word according to the event probability information and the hidden state information of the reference summary word comprises: Based on the keywords and the core event graph, selecting a target key text sentence from the text sentences corresponding to each text information; Performing encoding and decoding processing on each word segment in the target key text sentence to obtain probability information of the target key text sentence; A target summary word is selected from each candidate summary word according to the probability information of the target key text sentence, the event probability information and the hidden state information of the reference summary word.
8. The method according to claim 7, It is characterized in that The step of selecting a target key text sentence from text sentences corresponding to each text information based on the keyword and the core event graph includes: For each piece of text information, based on the keyword, selecting a key text sentence of the text information from each text sentence of the text information; For each key text sentence corresponding to the text information, the word frequency parameters corresponding to the keywords contained in the key text sentence are merged to obtain the first important index of the key text sentence; According to the phrases formed by each word segment in the key text sentence, a node search is performed in the core event graph to obtain the weight corresponding to the keyword phrase contained in the key text sentence; The weights corresponding to the key word groups contained in the key text sentence are merged to obtain the second important index of the key text sentence; Based on the first important indicator and the second important indicator, a target key text sentence is selected from the key text sentences corresponding to each text information.
9. The method according to claim 7, It is characterized in that The encoding and decoding of each word in the target key text sentence to obtain probability information of the target key text sentence includes: Performing feature encoding on each word in the target key text sentence to obtain encoding information of each word in the target key text sentence; The encoding information of each word segment in the target key text sentence is integrated to obtain context feature information; Based on the current decoding parameters, the context feature information is decoded to obtain the probability information of the target key text sentence.
10. The method according to claim 9, It is characterized in that The feature encoding of each word in the target key text sentence to obtain encoding information of each word in the target key text sentence includes: Based on the sequence of each segmentation in the target key text sentence, forward encoding each segmentation in the target key text sentence is performed to obtain forward encoding information of each segmentation in the target key text sentence; Based on the sequence of each segmentation in the target key text sentence, each segmentation in the target key text sentence is backward encoded to obtain backward encoding information of each segmentation in the target key text sentence; The forward encoding information and the backward encoding information are merged to obtain encoding information of each word segment in the target key text sentence.
11. The method according to claim 9, It is characterized in that The event distribution analysis is performed on the weights of the target-related keywords and the keyword group in which the reference summary words are located to obtain event probability information, including: The weights of the keyword groups containing the reference summary words are combined to obtain initial event probability information; Performing logistic regression processing on the semantic feature information corresponding to each target-related keyword, the latent state information of the candidate summary word, and the context feature information to obtain a context adaptive weight; The initial event probability information and the context adaptive weight are fused to obtain event probability information.
12. The method according to claim 1, It is characterized in that The step of selecting a target summary word from each candidate summary word according to the event probability information and the hidden state information of the reference summary word comprises: Extracting attention features of each candidate summary word according to the hidden state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word; Calculating target probability information corresponding to each candidate summary word according to the event probability information and the attention feature information; Based on the target probability information, a target summary word is selected from each candidate summary word.
13. The method according to claim 12, It is characterized in that The step of extracting attention features from each candidate summary word according to the hidden state information of the reference summary word to obtain attention feature information corresponding to each candidate summary word includes: For each candidate summary word, according to the hidden state information of the reference summary word, feature extraction is performed on the candidate summary word to obtain the hidden state information corresponding to the candidate summary word; Logistic regression processing is performed on the latent state information corresponding to the candidate summary word to obtain the attention feature information corresponding to the candidate summary word.
14. An information search device, It is characterized in that include: an acquisition unit, configured to acquire a keyword corresponding to at least one text information under the same topic, a reference abstract word of the text information, and at least one candidate abstract word, wherein the reference abstract word is an abstract word that has been selected to constitute a text abstract corresponding to the text information, or the reference abstract word is a reference word for selecting an abstract word to constitute a text abstract; A phrase extraction unit, configured to extract at least one keyword group from the keywords and determine a weight corresponding to the keyword group, wherein the keyword group includes at least two keywords having an associated relationship, and the weight represents the number of times the keyword group appears in the at least one text information; An event graph construction unit, configured to construct a core event graph according to the keyword groups and the weights corresponding to the keyword groups, wherein the core event graph includes nodes corresponding to the keywords and lines between the nodes, wherein the lines represent associations between the keywords; A node search unit, configured to perform a node search in the core event graph according to the reference summary word, and obtain a target associated keyword associated with the reference summary word in the core event graph; An analysis unit, configured to perform event distribution analysis on the weights of the target-related keywords and the keyword group in which the reference summary words are located, to obtain event probability information; The summary construction unit is used to select a target summary word from each candidate summary word according to the event probability information and the hidden state information of the reference summary word; perform summary construction processing on the target summary word and the reference summary word to obtain a text summary corresponding to the at least one text information, and output a search result according to the text summary.
15. An electronic device, It is characterized in that It comprises a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to execute the operations in the information search method according to any one of claims 1 to 13.
16. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the information search method according to any one of claims 1 to 13.
17. A computer program product comprising a computer program or instructions, It is characterized in that When the computer program or instruction is executed by a processor, the steps of the information search method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Text processing method and device, storage medium and equipment
CN113392641A
Event abstract generation method and device and electronic equipment
CN113535940A