Policy document key information extraction method and device based on artificial intelligence, terminal equipment and storage medium
By conducting multi-level analysis of initial news information, we guide the extraction of key information of policy documents, and solve the problem of insufficient correlation between key information of policy documents and social issues in the existing technology, and achieve accurate information docking and efficient implementation of policies.
Patent Information
- Application Number
- CN202510647461.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the existing technology to effectively link the key information in the policy documents with the social problems to be solved, making it difficult for the extracted key information to effectively solve actual social problems and affect the implementation effect of the policy.
By collecting initial news information, topic identification, sentiment analysis, influence analysis and event clustering, the causes and results of target events are obtained, thereby guiding the extraction of target key information from relevant policy documents.
The precise connection between key information in the policy documents and social issues is achieved, ensuring that the extracted information can effectively solve actual social problems and improve the implementation effect of policies.
Smart Images

Figure CN120179822A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, terminal device and storage medium for extracting key information of policy documents based on artificial intelligence. Background Art
[0002] In the aspect of policy document interpretation, the existing technologies mainly focus on the policy document text itself. These technologies often use methods such as natural language processing and machine learning to identify and extract the terms, regulations, goals, etc. in the policy document, concentrating on the information mining at the literal level, so that the key information extracted cannot be effectively associated with the social problems that the policy document aims to solve. Furthermore, due to the lack of this association, the key information extracted is difficult to effectively solve the actual social problems in practical applications. Thus, the limitations of the existing technologies in extracting key information from policy documents have greatly affected the implementation effect of policies and the solution efficiency of social problems. Summary of the Invention
[0003] The main purpose of the embodiments of the present invention is to provide a method, device, terminal device and storage medium for extracting key information of policy documents based on artificial intelligence, aiming to solve the problem in the related technologies that the key information extracted from the policy document cannot be effectively associated with the social problems that the policy document aims to solve, and thus the key information extracted is difficult to effectively solve the actual social problems in practical applications, thereby affecting the implementation effect of the policy.
[0004] In the first aspect, the embodiments of the present invention provide a method for extracting key information of policy documents based on artificial intelligence, including:
[0005] Collect initial news information, and perform topic recognition on the initial news information to obtain the initial topic corresponding to the initial news information and the associated news information corresponding to the initial topic;
[0006] Perform sentiment analysis on the initial topic according to the associated news information to obtain the target sentiment type corresponding to the initial topic;
[0007] Perform influence analysis on the initial topic according to the target sentiment type and the associated news information to obtain the target topic and the target news information corresponding to the target topic;
[0008] Perform event analysis on the target news information to obtain the initial event cause and the initial event result corresponding to the target topic;
[0009] Perform event clustering according to the initial event cause and the initial event result to obtain the target event cause corresponding to the target topic and the target event result corresponding to the target event cause;
[0010] Obtain relevant policy documents corresponding to the target topic;
[0011] Extract key information from the relevant policy documents according to the target event reason and the target event result, and obtain the target key information corresponding to the target event reason and the target event result in the relevant policy documents.
[0012] In a second aspect, an embodiment of the present invention provides an apparatus for extracting key information from policy documents based on artificial intelligence, including:
[0013] A topic recognition module, configured to collect initial news information, and perform topic recognition on the initial news information to obtain an initial topic corresponding to the initial news information and associated news information corresponding to the initial topic;
[0014] An emotion analysis module, configured to perform emotion analysis on the initial topic according to the associated news information to obtain a target emotion type corresponding to the initial topic;
[0015] An influence analysis module, configured to perform influence analysis on the initial topic according to the target emotion type and the associated news information to obtain a target topic and target news information corresponding to the target topic;
[0016] An event analysis module, configured to perform event analysis on the target news information to obtain an initial event reason and an initial event result corresponding to the target topic;
[0017] An event clustering module, configured to perform event clustering according to the initial event reason and the initial event result to obtain a target event reason corresponding to the target topic and a target event result corresponding to the target event reason;
[0018] A document acquisition module, configured to obtain relevant policy documents corresponding to the target topic;
[0019] An information extraction module, configured to extract key information from the relevant policy documents according to the target event reason and the target event result, and obtain the target key information corresponding to the target event reason and the target event result in the relevant policy documents.
[0020] In a third aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the steps of any method for extracting key information from policy documents based on artificial intelligence provided in the specification of the present invention are implemented.
[0021] Fourthly, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the artificial intelligence-based key information extraction methods provided in the specification of the present invention.
[0022] The embodiment of the present invention provides an artificial intelligence-based key information extraction method, device, terminal device, and storage medium. The method includes: obtaining an initial topic and associated news information corresponding to the initial topic by collecting initial news information and performing topic recognition on the initial news information, and then performing sentiment analysis on the initial topic to obtain a target sentiment type, so as to timely grasp the attitude and mood of the masses towards the initial topic. Then, perform influence analysis on the initial topic based on the target sentiment type and associated news information to obtain a target topic and target news information corresponding to the target topic. Next, perform event analysis on the target news information to obtain an initial event cause and an initial event result corresponding to the target news information, and perform event clustering based on the initial event cause and the initial event result to obtain a target event cause corresponding to the target topic and a target event result corresponding to the target event cause. Then, obtain relevant policy documents corresponding to the target topic, and extract target key information from the relevant policy documents according to the target event cause and the target event result, so as to timely and accurately obtain the target key information corresponding to the target event cause and the target event result from the relevant policy documents, and ensure that the collected target key information matches the target event cause and the target event result, thereby being able to timely mine relevant policy information corresponding to the target event cause from the relevant policy documents. The method extracts target key information from relevant policy documents guided by the target event cause and the target event result, breaks the gap between information extraction and actual problem-solving in traditional technologies, realizes the precise docking of policy information and social problems, and provides a solid guarantee for the efficient implementation of policies and the effective solution of social problems. It also solves the problem in related technologies that there is no effective association between the key information extracted from policy documents and the social problems that the policy documents aim to solve, so that the extracted key information is difficult to effectively solve actual social problems in practical applications, thus affecting the implementation effect of policies. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1A flowchart showing a method for extracting key information from policy documents based on artificial intelligence provided by an embodiment of the present invention;
[0025] Figure 2 A schematic block diagram of the module structure of a device for extracting key information from policy documents based on artificial intelligence provided by an embodiment of the present invention;
[0026] Figure 3 A schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] The flowchart shown in the accompanying drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0029] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0030] An embodiment of the present invention provides a method, device, terminal device, and storage medium for extracting key information from policy documents based on artificial intelligence. Among them, the method for extracting key information from policy documents based on artificial intelligence can be applied to a terminal device, and the terminal device can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The terminal device can be a server or a server cluster.
[0031] Next, some embodiments of the present invention will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0032] Please refer to Figure 1 , Figure 1 A flowchart showing a method for extracting key information from policy documents based on artificial intelligence provided by an embodiment of the present invention.
[0033] As Figure 1As shown, the artificial intelligence-based policy document key information extraction method includes steps S101 to S107.
[0034] Step S101: Collect initial news information, and perform topic identification on the initial news information to obtain an initial topic corresponding to the initial news information and associated news information corresponding to the initial topic.
[0035] Exemplarily, the initial news information is obtained by automatically crawling corresponding news information from online news platforms such as Sina News, Tencent News, Toutiao, etc. using Octopus collectors, Locomotive collectors, etc.
[0036] For example, since different news sources may publish the same news, the collected news information needs to be deduplicated to avoid data redundancy in subsequent analysis, and irrelevant content such as advertisements, copyright statements, navigation links, etc. in the news needs to be removed, and only the core news text is retained to obtain the processed initial news information.
[0037] Exemplarily, the processed initial news information is converted into corresponding text vectors using a bag-of-words model or a BERT model, and the similarity between any two processed initial news information is calculated based on cosine similarity or Euclidean distance, and then the processed initial news information is clustered according to the similarity to obtain a news clustering result, and then a keyword extraction algorithm such as TextRank is used to extract keywords from each sub-news data in each cluster in the news clustering result, and then the relevant keywords corresponding to each sub-news data are obtained, and then the commonalities and thematic tendencies of the relevant keywords corresponding to each cluster are analyzed, and then the keyword combination that can represent the core content of the entire cluster is found through statistical analysis, and then these keyword combinations are summarized into one or several initial topics that can summarize the news theme of the cluster, and then the relevant cluster corresponding to the initial topic is determined as the associated news information corresponding to the initial topic.
[0038] In some embodiments, the topic identification of the initial news information to obtain the initial topic corresponding to the initial news information and the related news information corresponding to the initial topic includes: preprocessing the initial news information to obtain initial keywords corresponding to the initial news information, and randomly combining the initial keywords to obtain multiple keyword groups, the keyword groups including a first keyword and a second keyword; obtaining a first text vector corresponding to the first keyword according to a text representation model and a second text vector corresponding to the second keyword according to the text representation model; obtaining first frequency information and first spatial information corresponding to the first keyword in the initial news information; obtaining second frequency information and second spatial information corresponding to the second keyword in the initial news information; obtaining third frequency information corresponding to when the first keyword and the second keyword appear in a sub-sentence of the initial news information at the same time; determining the first frequency information and the second frequency information corresponding to the first keyword and the second keyword according to the first frequency information, the second frequency information, and the third frequency information. a correlation degree; determining the text similarity between the first keyword and the second keyword according to the first text vector and the second text vector; determining the second correlation between the first keyword and the second keyword according to the first spatial information, the second spatial information and the text similarity; determining the correlation weight corresponding to the keyword group by fusing the first correlation degree and the second correlation degree, and scoring the importance of the keywords in the keyword group according to the correlation weight to obtain a target score; screening the initial keywords according to the target score to obtain the target keywords corresponding to the initial news information; performing news clustering on the initial news information according to the target keywords to obtain news clustering results; performing attention analysis on each sub-clustering result in the news clustering result to obtain the attention change value corresponding to the sub-clustering result; performing topic screening on the sub-clustering result according to the attention change value to obtain the initial topic corresponding to the initial news information, and determining the associated news information corresponding to the initial topic according to the sub-clustering result.
[0039] Exemplarily, noise in the initial news information is removed, such as HTML tags, special symbols, stop words (such as "的", "是", "在" and other words without actual meaning), and then the cleaned news text is segmented into individual words using tools such as Jieba word segmentation, and then keyword extraction algorithms such as the word frequency inverse document frequency algorithm are used to calculate the importance of each word in the initial news information, and then the words with higher scores are screened out as the initial keywords corresponding to the initial news information.
[0040] Exemplarily, the initial keywords corresponding to each initial news information are combined in pairs to form a plurality of keyword groups, each keyword group including a first keyword and a second keyword.
[0041] Exemplarily, a text representation model such as Word2Vec, GloVe, etc. is used to perform vector representation on the first keyword and the second keyword, so as to obtain a first text vector corresponding to the first keyword and a second text vector corresponding to the second keyword.
[0042] Exemplarily, the initial news information is traversed, the number of times the first keyword appears is counted to obtain first frequency information, and the positions where the first keyword appears in the initial news information are calculated in total to obtain first spatial information corresponding to the first keyword.
[0043] Exemplarily, the initial news information is traversed, the number of times the second keyword appears is counted to obtain second frequency information, and the positions where the second keyword appears in the initial news information are calculated in total to obtain second spatial information corresponding to the second keyword.
[0044] Exemplarily, the number of times the first keyword and the second keyword appear simultaneously in a sub-statement of the initial news information is counted to obtain third frequency information. The sub-statement can be a complete sentence in the initial news information.
[0045] Exemplarily, the first correlation degree is determined according to the first frequency information, the second frequency information, and the third frequency information. Generally speaking, if the third frequency information that the first keyword and the second keyword appear simultaneously is relatively high compared to their respective first frequency information and second frequency information, then the correlation degree between the first keyword and the second keyword is high. For example, twice the value of the third frequency information is obtained as the first value, and then the sum of the first frequency information and the second frequency information is obtained as the second value, so as to divide the first value by the second value to obtain the first correlation degree.
[0046] Exemplarily, methods such as cosine similarity are used to calculate the similarity between the first text vector and the second text vector to obtain text similarity, and then the second correlation degree is determined by combining the first spatial information, the second spatial information, and the text similarity. If the spatial positions of two keywords in the news are close and the text similarity is high, then their second correlation degree is high. For example, the reciprocal of each first position in the first spatial information is calculated and then the reciprocal sum is obtained to obtain the first spatial representation corresponding to the first spatial information, the reciprocal of each second position in the second spatial information is calculated and then the reciprocal sum is obtained to obtain the second spatial representation corresponding to the second spatial information, so as to multiply the first spatial representation and the second spatial representation and then divide by the text similarity to obtain the second correlation degree.
[0047] Exemplarily, the first correlation degree and the second correlation degree are multiplied and fused to obtain the correlation weight corresponding to the keyword group. Thus, after calculating the scores of the initial keywords using TF-IDF, the scores of each keyword in the initial keywords are adjusted using the correlation weight to obtain the target scores. Furthermore, a threshold of the target scores is set, and the keywords with target scores greater than the threshold are screened out as the target keywords corresponding to the initial news information.
[0048] Exemplarily, clustering algorithms such as K-Means and DBSCAN are selected, and the initial news information is clustered according to the target keywords until the news clustering results are obtained.
[0049] Exemplarily, the popularity data corresponding to each sub-clustering result in the news clustering results is obtained through a data interface or web crawler technology. The popularity data includes, but is not limited to, data such as the number of views, likes, comments, and forwards of each news data record in the sub-clustering result on the network platform. Furthermore, the change situation of the popularity data of each sub-clustering result in different time periods is calculated to obtain the attention change value, and multiple attention change values are obtained through the data analysis of multiple different time periods. Thus, the change trend corresponding to the sub-clustering result is determined through the multiple attention change values. There are various types of change trends, such as gradually increasing, remaining unchanged, and gradually decreasing, etc. If a series of attention change values show a continuous upward trend, it indicates that the attention of the news in this sub-clustering is gradually increasing; if the change values fluctuate within a small range around zero, it can be considered that the attention remains relatively stable; if the change values are continuously negative and the absolute value gradually increases, it means that the attention is gradually decreasing.
[0050] Exemplarily, when it is determined that the change trend of a certain sub-clustering result is gradually increasing, this indicates that the news in this sub-clustering is receiving more and more attention from the public, has high news value and potential influence. Furthermore, topic recognition needs to be carried out on this sub-clustering result to determine the initial topic corresponding to this sub-clustering result, and the sub-clustering result corresponding to this initial topic is determined as the associated news information corresponding to this initial topic.
[0051] In some embodiments, obtaining the attention change value corresponding to each sub-clustering result in the news clustering result through attention analysis includes: obtaining the first reporting time corresponding to each sub-news data in the sub-clustering result, and determining the first time span corresponding to the sub-clustering result according to the first reporting time; determining a first target window, and partitioning the data of the sub-clustering result according to the first target window and the first time span to obtain a first segmentation result corresponding to the sub-clustering result and first associated data corresponding to the first segmentation result; obtaining the second reporting time corresponding to the first associated data, and determining the second time span corresponding to the first associated data according to the second reporting time; determining a second target window, and partitioning the data of the first associated data according to the second target window and the second time span to obtain a second segmentation result corresponding to the first associated data, where the second target window is smaller than the first target window; determining the number of windows according to the second target window and the second time span; determining second associated data corresponding to the second target window from the first associated data according to the second target window and the second reporting time; obtaining the data view volume corresponding to the second associated data, and determining the attention characterization value corresponding to the first associated data under the second time span according to the second target window, the number of windows, the data view volume, and the second associated data; determining the attention change value corresponding to the sub-clustering result according to the attention characterization value of the first associated data; where the attention characterization value is obtained according to the following formula:
[0052]
[0053] where represents the attention characterization value corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the second target window corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the number of windows corresponding to the second target window corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the window size corresponding to the second target window corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the number of news of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the t-th second target window, Denote the data view volume of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window in the t-th second target window. Denote the data view volume of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window in the k-th second target window. Denote the number of news of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window in the k-th second target window.
[0054] Exemplarily, obtain the reporting time corresponding to each sub-news data in the sub-clustering result, determine this time as the first reporting time corresponding to this sub-news data, and subtract the earliest first reporting time and the latest first reporting time in the sub-clustering result to obtain the first time span corresponding to this sub-clustering result.
[0055] Exemplarily, set a time length as the first target window. According to the first target window and the first time span, divide the sub-clustering result in chronological order. After division, several subsets of time periods are obtained, and these subsets are the first segmentation result, and the news data contained in each subset is the corresponding first associated data.
[0056] Exemplarily, obtain the reporting time corresponding to each news data in the first associated data and determine it as the second reporting time. Find the earliest and latest second reporting times in the first associated data, and subtract the two to obtain the second time span corresponding to the first associated data. This further narrows the time range and focuses on the time characteristics of the first associated data.
[0057] Exemplarily, set a time length smaller than the first target window as the second target window, and then divide the first associated data again according to the second target window and the second time span to obtain several subsets of smaller time periods, that is, the second segmentation result. This more detailed division helps to discover the change law of the data in a shorter time.
[0058] Exemplarily, according to the second target window and the second time span, divide the second time span by the second target window to obtain the number of windows included in the second time span, and then screen out the news data in each second target window time period from the first associated data according to the second target window and in combination with the second reporting time. These data are the second associated data. This can accurately focus on the news information within a specific time period.
[0059] Exemplarily, obtain the view count corresponding to the second associated data from relevant platforms of news data. The view count can intuitively reflect the degree of attention received by the news. Thus, according to the following formula, combining the second target window, the number of windows, the data view count, and the second associated data, determine the attention characterization value corresponding to the first associated data within the second time span. The attention characterization value can quantify the attention received by news data during this time period:
[0060]
[0061] Wherein, represents the attention characterization value corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the second target window corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the number of windows corresponding to the second target window corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the window size of the second target window corresponding to the j-th first associated data of the i-th sub-clustering result under the first target window, represents the number of news of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the t-th second target window, represents the data view count of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the t-th second target window, represents the data view count of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the k-th second target window, represents the number of news of the second associated data corresponding to the j-th first associated data of the i-th sub-clustering result under the k-th second target window.
[0062] Exemplarily, the above formula comprehensively considers various factors such as the second target window, the number of windows, the data view count, and the number of news, and can more comprehensively and accurately quantify the attention received by news data within a certain time period. Compared with the simple view count, the attention characterization value can avoid evaluation errors caused by unreasonable time period division or differences in the number of news. Thus, by calculating the attention characterization values corresponding to different sub-clustering results and different first associated data, the attention received by news in different time periods can be compared, providing good support for obtaining the attention change values corresponding to the sub-clustering results subsequently.
[0063] Exemplarily, after obtaining the attention characterization value corresponding to the first associated data under the first target window, arrange the attention characterization values according to the segmentation order of the first target window in the sub-clustering result, so as to obtain the sorting sequence of the attention characterization values corresponding to the sub-clustering result. Furthermore, calculate the difference according to the adjacent positions in the sorting sequence to obtain the attention change value corresponding to the sub-clustering result, and thus determine the change trend corresponding to the sub-clustering result through multiple attention change values. The attention change value can reflect the change of the attention degree of the sub-clustering result at different time periods.
[0064] Exemplarily, by calculating the attention characterization value and the attention change value, the attention degree and its change trend of the news at different time periods can be accurately grasped.
[0065] Step S102: Perform sentiment analysis on the initial topic according to the associated news information to obtain the target sentiment type corresponding to the initial topic.
[0066] Exemplarily, determine a sentiment classification model. For example, the sentiment classification model is a machine learning model such as a Naive Bayes classifier, a support vector machine, etc. Then, through learning and training on a large amount of data with sentiment labels, a trained sentiment classification model is obtained.
[0067] Exemplarily, before inputting the associated news information into the sentiment classification model, the data needs to be preprocessed. This includes removing noise information in the text, such as punctuation marks, stop words, etc.; performing word segmentation on the associated news information to split it into individual words, and then performing sentiment classification on the preprocessed text according to the sentiment classification model to output the corresponding initial sentiment type. The initial sentiment type is one of positive, negative, and neutral. Furthermore, after obtaining the initial sentiment type corresponding to each associated news information, perform statistical analysis on the initial sentiment types of all the associated news information and adopt the majority voting principle, that is, determine the target sentiment type of the initial topic as the one with the largest number of news of that sentiment type, so as to obtain the target sentiment type corresponding to the initial topic.
[0068] Step S103: Perform influence analysis on the initial topic according to the target sentiment type and the associated news information to obtain the target topic and the target news information corresponding to the target topic.
[0069] Exemplarily, when the target emotion type corresponding to the initial topic is a negative type, the number and type of media that release relevant news for the relevant news information corresponding to the initial topic are counted, and interaction data such as the number of comments, likes, and shares of the relevant news information on social media are counted, and the discussion heat of the initial topic on social media is monitored, such as the keyword search volume, the usage frequency of topic tags, etc., so as to determine the influence characterization value corresponding to the initial topic according to the number and type of media, interaction data, discussion heat, etc. combined with the neural network prediction model, and then compare the influence characterization value with a preset value. When the influence characterization value is greater than or equal to the preset value, the initial topic is determined as the target topic, and the relevant news information corresponding to the initial topic is determined as the target news information corresponding to the target topic.
[0070] In some embodiments, the obtaining of the target topic and the target news information corresponding to the target topic by performing influence analysis on the initial topic according to the target emotion type and the relevant news information includes: determining a preset emotion type. When the target emotion type is the preset emotion type, keyword distribution analysis is performed on the relevant news information to obtain the relevant keywords corresponding to the initial topic and the initial distribution information corresponding to the relevant keywords; screening the relevant keywords according to the initial distribution information to obtain the topic keywords corresponding to the initial topic and the target distribution information corresponding to the topic keywords; obtaining the number of news in which the topic keywords exist from the relevant news information, and determining the influence characterization value corresponding to the initial topic according to the target distribution information, the number of news, and the relevant news information; screening the initial topic according to the influence characterization value to obtain the target topic, and obtaining the target news information corresponding to the target topic from the relevant news information according to the target topic.
[0071] Exemplarily, the preset emotion type is a negative type. Then, when the target emotion type is a negative type, a keyword extraction algorithm is used to extract keywords from the relevant news information to obtain the relevant keywords corresponding to the initial topic and the initial distribution information corresponding to the relevant keywords.
[0072] Exemplarily, the initial distribution information is compared with the preset distribution information. When the initial distribution information is greater than or equal to the preset distribution information, the initial keyword is determined as the topic keyword corresponding to the initial topic, and the initial distribution information corresponding to the initial keyword is determined as the target distribution information corresponding to the topic keyword.
[0073] Exemplarily, the number of news in which the topic keywords appear in all the relevant news information is counted, so as to obtain the influence characterization value corresponding to the initial topic according to the following formula:
[0074]
[0075] Among them, represents the influence characterization value corresponding to the g-th initial topic, and num represents the number of topic keywords corresponding to the g-th initial topic. represents the target distribution information corresponding to the y-th topic keyword of the g-th initial topic, and lg represents the logarithmic function with base 10. represents the number of associated news information corresponding to the g-th initial topic. represents the number of news corresponding to the y-th topic keyword of the g-th initial topic.
[0076] Exemplarily, a preset characterization value is determined. Then, when the influence characterization value is greater than or equal to the preset characterization value, the initial topic corresponding to the influence characterization value is determined as the target topic, and the associated news information of the initial topic corresponding to the influence characterization value is determined as the target news information corresponding to the target topic.
[0077] Step S104: Perform event analysis on the target news information to obtain the initial event cause and initial event result corresponding to the target topic.
[0078] Exemplarily, event elements such as event subject, event time, event location, event behavior, event object, etc. are determined, and corresponding rule templates are created. Then, the target news information is matched with the formulated event extraction rules. Scan the news text sentence by sentence to determine whether it conforms to a certain rule template. If it conforms, extract the corresponding event elements, and integrate the extracted event elements to form complete event information. For example, combine the extracted elements such as subject, time, location, behavior, etc. into a clear event description.
[0079] Exemplarily, the type of association relationship to be analyzed is determined, such as time sequence relationship, conditional relationship, parallel relationship, causal relationship, etc. Then, semantic analysis is performed on any two sub-events in the event information to determine the association relationship between them. Then, from all the analyzed association relationships, filter out the sub-event pairs with causal relationships. For each group of sub-events with causal relationships, determine the sub-event corresponding to the cause as the initial event cause corresponding to the target topic, and determine the sub-event corresponding to the result as the initial event result.
[0080] In some embodiments, obtaining the initial event cause and the initial event result corresponding to the target topic through event analysis of the target news information includes: extracting event features from the target news information to obtain a target event and related event information corresponding to the target event; obtaining a first event and a second event from the target event, and obtaining first information corresponding to the first event and second information corresponding to the second event from the related event information; determining a first trigger word corresponding to the first event and a second trigger word corresponding to the second event; determining a logical correlation value between the first event and the second event according to the first trigger word and the second trigger word in combination with the first information and the second information; determining the relationship type between the first event and the second event according to the logical correlation value, and determining the initial event cause and the initial event result corresponding to the target topic from the first event and the second event according to the relationship type.
[0081] Exemplarily, determine the event features to be concerned in the target news information, such as the event subject (person, organization, institution, etc.), the time, place, specific action, and objects involved in the event. These features are the key basis for identifying events. Based on the determined event features, conduct a detailed analysis of the target news information, find out the events that meet the features, and determine them as target events. At the same time, collect various information closely related to the target event, such as the time, place, and objects of the event, and these information constitute the related event information.
[0082] Exemplarily, select two events from the target event and determine them as the first event and the second event respectively, and obtain the first information corresponding to the first event and the second information corresponding to the second event from the related event information.
[0083] Exemplarily, a trigger word is a word that can reflect the core behavior or key state change of an event, and thus obtain the first trigger word corresponding to the first event and the second trigger word corresponding to the second event from the target news data.
[0084] Exemplarily, a pre-trained word vector model (such as Word2Vec, GloVe, etc.) is used to perform text encoding on the first trigger word, the second trigger word, the first information, and the second information, and then the encoded trigger word vectors and information vectors are combined. The first trigger word vector and the first information vector can be concatenated together to form a comprehensive vector of the first event, and similarly, a comprehensive vector of the second event can be obtained, so as to collect a large amount of sample data containing information related to the first event and the second event, and manually annotate it according to the actual logical association between the events. For example, it can be annotated as "strong association", "medium association", "weak association" or represented by a specific value (such as a score between 0 and 1) to represent the degree of association, and then the annotated data is divided into a training set, a validation set, and a test set. The training set is used to train the neural network model, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the final performance of the model. Furthermore, a neural network model such as a recurrent neural network and a loss function are used to perform model training using the training set, the validation set, and the test set to obtain the target neural network model.
[0085] Exemplarily, the trained target neural network model is used to perform data fusion on the first trigger word, the second trigger word, the first information, and the second information and then perform data prediction, and then output the logical association value between the first event and the second event. The logical association value is used to characterize the strength of the association between the first event and the second event.
[0086] Exemplarily, according to the magnitude of the logical association value, the relationship type between the first event and the second event is divided into different categories. Common relationship types include causal relationship, time sequence relationship, parallel relationship, conditional relationship, etc. For example, when the logical association value is relatively high and there is an obvious causal derivation relationship, it is determined as a causal relationship. If the relationship type between the first event and the second event is a causal relationship, then the event corresponding to the "cause" in the causal relationship is determined as the initial event cause corresponding to the target topic, and the event corresponding to the "effect" is determined as the initial event result. For other relationship types, it is necessary to further analyze whether the initial event cause and result can be determined according to the specific situation. If it cannot be determined, it is necessary to re-examine the event selection or association analysis process.
[0087] In some embodiments, determining the logical correlation value between the first event and the second event by combining the first information and the second information according to the first trigger word and the second trigger word includes: determining the corresponding sequence order between the first event and the second event under the event logic according to the first trigger word and the second trigger word; determining the target co-occurrence times of the first trigger word and the second trigger word corresponding in the target news information according to the sequence order, and determining the third frequency information of the first trigger word in the target news information and the fourth frequency information of the second trigger word in the target news information; determining the corresponding third correlation degree between the first event and the second event under the first trigger word and the second trigger word according to the target co-occurrence times, the third frequency information and the fourth frequency information; determining the first mutual correlation value between the first event and the second event according to the first trigger word and the second information, and determining the second mutual correlation value between the first event and the second event according to the second trigger word and the first information; fusing the first mutual correlation value and the second mutual correlation value to determine the fourth correlation degree between the first event and the second event; performing correlation degree analysis according to the first information and the second information to obtain the fifth correlation degree between the first event and the second event; fusing the third correlation degree, the fourth correlation degree and the fifth correlation degree to determine the logical correlation value corresponding between the first event and the second event under the sequence order.
[0088] Exemplarily, determine the behaviors or states represented by the first trigger word and the second trigger word, and judge the sequence order of the first event and the second event according to the inherent logical rules of the event. For example, for "sign a contract" and "execute a contract", logically, "sign a contract" must precede "execute a contract". Or for "receive betrothal gifts" and "resist betrothal gifts", logically, "receive betrothal gifts" must precede "resist betrothal gifts".
[0089] Exemplarily, determine the corresponding appearance order of the first trigger word and the second trigger word according to the sequence order, so as to check sentence by sentence in the target news information the number of times the first trigger word and the second trigger word appear simultaneously in the appearance order of the first trigger word and the second trigger word, and determine this number as the target co-occurrence times. The more the co-occurrence times, the closer the correlation between the two events may be, and separately count the number of times the first trigger word and the second trigger word appear alone in the target news information to obtain the third frequency information of the first trigger word and the fourth frequency information of the second trigger word. The frequency information reflects the importance and the frequency of appearance of each trigger word in the news.
[0090] Exemplarily, multiply the third frequency information and the fourth frequency information to obtain a first processed value, and add the third frequency information and the fourth frequency information to obtain a second processed value. Thus, divide the target co-occurrence times by the first processed value to obtain a third processed value, and divide the target constant 2 by the second processed value to obtain a fourth processed value. Furthermore, take the function value of the logarithm with base 10 of the third processed value to obtain a first target value, and at the same time take the function value of the logarithm with base 10 of the fourth processed value to obtain a second target value. Then, divide the first target value and the second target value to obtain the third correlation degree corresponding to the first event and the second event under the first trigger word and the second trigger word.
[0091] Exemplarily, determine information such as the second subject, the second location, and the second behavior included in the second event in the second information, and then calculate the correlation degree between the first trigger word and each piece of information in the second information. For example, calculate the co-occurrence frequency information corresponding to the co-occurrence of the first trigger word and the second subject, and obtain the fifth frequency information corresponding to the second subject. Then, multiply the third frequency information and the fifth frequency information to obtain a fifth processed value, and add the third frequency information and the fifth frequency information to obtain a sixth processed value. Thus, divide the co-occurrence frequency information by the fifth processed value to obtain a seventh processed value, and divide the target constant 2 by the sixth processed value to obtain an eighth processed value. Furthermore, take the function value of the logarithm with base 10 of the seventh processed value to obtain a third target value, and at the same time take the function value of the logarithm with base 10 of the eighth processed value to obtain a fourth target value. Then, divide the third target value and the fourth target value to obtain the data correlation value corresponding to the first event and the second event under the first trigger word and the second subject, and then obtain the data correlation value between the first trigger word and each piece of data in the first information. Thus, sum and average all the data correlation values to obtain the first mutual correlation value between the first trigger word and the first information.
[0092] Exemplarily, based on the same principle as obtaining the first mutual correlation value, then determine the second mutual correlation value between the second trigger word and the second information according to the second trigger word and the first information. Thus, sum the first mutual correlation value and the mutual correlation value to obtain the fourth correlation degree between the first event and the second event.
[0093] Exemplarily, pair up the information such as the first subject, the first location, and the first behavior in the first information with the information such as the second subject, the second location, and the second behavior in the second information to obtain the third mutual correlation value between any piece of information in the first information and any piece of information in the second information. The calculation method of the third mutual correlation value is the same as that of the first mutual correlation value, and will not be elaborated in this application.
[0094] Exemplarily, obtain all the third mutual correlation values and then sum and average them to obtain the fifth correlation degree between the first event and the second event.
[0095] Exemplarily, the third correlation degree, the fourth correlation degree, and the fifth correlation degree are multiplied in sequence to obtain the logical correlation value corresponding to the first event and the second event in the sequence order. The logical correlation value can more accurately reflect the logical correlation degree between the first event and the second event.
[0096] Step S105: Perform event clustering according to the initial event cause and the initial event result to obtain the target event cause corresponding to the target topic and the target event result corresponding to the target event cause.
[0097] Exemplarily, use clustering algorithms such as hierarchical clustering and DBSCAN to cluster the initial event cause and the initial event result to obtain the event clustering result. Then, find the cluster with the largest number of sub-clusters in the event clustering result, and determine the event cause corresponding to this cluster as the target event cause corresponding to the target topic. This means that this event cause has high representativeness and universality among all events. After determining the target event cause, determine the event result corresponding to this target event cause as the target event result corresponding to the target event cause. These results are closely related to the target event cause and can reflect the most common causal relationship under the target topic.
[0098] Step S106: Obtain the relevant policy documents corresponding to the target topic.
[0099] Exemplarily, perform keyword mapping and expansion on the target topic to obtain the relevant keywords corresponding to the target topic. For example, if the target topic is "development of new energy vehicle industry", the relevant keywords may include "new energy vehicle", "industrial policy", "subsidy policy", "new energy technology standard", etc. Then, infer the official websites that may issue relevant policy documents based on the field of the target topic, and then retrieve and download on the official websites according to the relevant keywords to obtain the relevant policy documents of the target topic.
[0100] Step S107: Extract key information from the relevant policy documents according to the target event cause and the target event result to obtain the target key information corresponding to the target event cause and the target event result in the relevant policy documents.
[0101] Exemplarily, based on common sentence-ending symbols such as full stops and semicolons, the policy document is segmented into individual independent initial statement information. Then, each initial statement information is associated with the target event cause and input into the relationship classification model. The relationship classification model is trained based on a large amount of sample data and can identify and judge the types of association relationships between different pieces of information. The model will conduct an in-depth analysis of the logical connection between the initial statement information and the target event cause. For example, if the target event cause is "insufficient enterprise innovation ability" and the initial statement information is "the government provides innovation subsidies", the model may judge that there is a causal relationship between the two because the government providing innovation subsidies is likely to solve the problem of insufficient enterprise innovation ability. Through the analysis and judgment of the model, the first relationship type corresponding to each initial statement information and the target event cause can be obtained. This relationship type may include causal relationship, correlation relationship, no relationship, etc.
[0102] Similarly, each initial statement information is associated with the target event result and input into the relationship classification model. The model will evaluate the logical relationship between the initial statement information and the target event result. Suppose the target event result is "decline in industrial competitiveness" and the initial statement information is "strengthening industry-university-research cooperation". The model may judge that there is a causal relationship between the two because strengthening industry-university-research cooperation may improve the problem of the decline in industrial competitiveness. Through the processing of the model, the second relationship type corresponding to each initial statement information and the target event result can be obtained, and its type is similar to the first relationship type. Thus, after obtaining the first relationship type and the second relationship type, when both the first relationship type and the second relationship type corresponding to an initial statement information are causal relationships, it indicates that this initial statement information has a causal connection with both the target event cause and the target event result in the policy document. This means that this statement is likely to be the measure taken in the policy document for the target event cause, and these measures can directly solve the target event result, thereby determining this initial statement information as the target key information closely related to the target event cause and the target event result.
[0103] In some embodiments, the key information extraction of the relevant policy document according to the target event cause and the target event result, to obtain the target key information corresponding to the target event cause and the target event result in the relevant policy document, includes: disassembling the relevant policy document into multiple initial statement information, and generating a first generated text by text generation of the initial statement information and the target event cause according to a first splicing rule; generating a second generated text by text generation of the initial statement information and the target event result according to a second splicing rule; using a causal relationship recognition model to classify the relationship of the first generated text to obtain a first relationship type corresponding to the first generated text; using the causal relationship recognition model to classify the relationship of the second generated text to obtain a second relationship type corresponding to the second generated text; screening the initial statement information according to the first relationship type and the second relationship type to obtain intermediate statement information; performing a statement key degree scoring on the intermediate statement information to obtain a key degree score corresponding to the intermediate statement information; and obtaining the target key information corresponding to the target event cause and the target event result in the relevant policy document from the intermediate statement information according to the key degree score.
[0104] Exemplarily, based on common statement ending symbols such as full stops and semicolons, the policy document is segmented into independent initial statement information one by one. Then, the initial statement information and the target event cause are spliced according to the first splicing rule to obtain the first generated text. For example, the first splicing rule is to place the target event cause before the initial statement information and connect them with a specific connecting word (such as "For (target event cause), adopt the means of (initial statement information) to solve"), so as to splice each initial statement information and the target event cause to generate the first generated text.
[0105] For example, the first splicing rule can place the target event cause before the initial statement information and use a specific connecting word to connect the two. Specifically, the connecting word can adopt the expression "For (target event cause), adopt the means of (initial statement information) to solve". Suppose the target event cause is "The problem of serious excessive sewage discharge of enterprises" and the initial statement information is "Establish a real-time sewage discharge monitoring system", then the first generated text generated according to the first splicing rule is "For the problem of serious excessive sewage discharge of enterprises, adopt the means of establishing a real-time sewage discharge monitoring system to solve". Through such splicing, a clear association can be established between the specific measures in the policy document and the target event cause, facilitating further analysis of the relationship between the two subsequently.
[0106] Exemplarily, in addition to generating the first generated text related to the cause of the target event, it is also necessary to construct the second generated text related to the result of the target event. This also requires formulating appropriate splicing rules to accurately reflect the logical connection between the initial statement information and the result of the target event. Place the result of the target event after the initial statement information and connect them with appropriate conjunctions. For example, the conjunction can be "After adopting (initial statement information), (result of the target event) can be solved". Suppose the result of the target event is "The air quality in the city continues to deteriorate" and the initial statement information is "Increase the intensity of dust pollution control", then the second generated text generated according to the second splicing rule is "After adopting the increase in the intensity of dust pollution control, the continuous deterioration of the air quality in the city can be solved". Through this splicing method, the potential connection between the measures in the policy document and the result of the target event can be clearly seen, providing the necessary text materials for subsequent causal relationship analysis.
[0107] Exemplarily, in order to accurately judge the causal relationship implied in the first generated text and the second generated text, a causal relationship recognition model needs to be used for classification. Before using this model, it needs to be trained. Generated statements corresponding to statements with causal relationships and false statements corresponding to statements without causal relationships can be constructed. Statements with causal relationships can be selected from sources such as actual policy cases and relevant research reports, such as "Because the fuel tax has been increased, the fuel consumption of cars has decreased". False statements without causal relationships can be constructed reasonably, such as "The sun rises and Xiaoming has breakfast", and there is no direct causal connection between the two. These generated statements and false statements are used as training data and input into the neural network model for training. During the training process, the model will continuously learn and adjust parameters to identify whether there is a causal relationship in the input statements. After sufficient training, a causal relationship recognition model that can relatively accurately judge the causal relationship can be obtained.
[0108] Exemplarily, input the first generated text into this causal relationship recognition model, and the model will analyze and judge it to obtain the first relationship type corresponding to the first generated text. This relationship type may be one of causal relationship and non-causal relationship. Similarly, input the second generated text into the causal relationship recognition model, and the model will also perform corresponding processing to obtain the second relationship type corresponding to the second generated text.
[0109] Exemplarily, when both the first relationship type and the second relationship type are causal relationships, the initial statement information involved in the first relationship type and the second relationship type is determined as the intermediate statement information.
[0110] Exemplarily, obtain the first probability corresponding to the first relationship type being a causal relationship and the second probability corresponding to the second relationship type being a causal relationship, thereby summing and averaging the first probability and the second probability to obtain the target probability, and using the TextRank algorithm to perform a sentence key degree scoring on the intermediate sentence information to obtain the scoring value corresponding to the intermediate sentence information. Furthermore, after converting the target probability and the scoring value to the same dimension, sum and average them to obtain the key degree value corresponding to the intermediate sentence information.
[0111] Exemplarily, determine a threshold of the key degree value, and then screen out the sentences in the intermediate sentence information whose key degree values are higher than the threshold. These sentences are the target key information corresponding to the target event cause and the target event result in the relevant policy documents. It is also possible to sort the key degree values from high to low and select a certain number of sentences with the top rankings as the target key information.
[0112] Please refer to Figure 2 , Figure 2 There is provided a device 200 for extracting key information of a policy document based on artificial intelligence according to an embodiment of the present application. The device 200 for extracting key information of a policy document based on artificial intelligence includes a topic recognition module 201, a sentiment analysis module 202, an impact analysis module 203, an event analysis module 204, an event clustering module 205, a file acquisition module 206, and an information extraction module 207. Among them, the topic recognition module 201 is configured to collect initial news information and perform topic recognition on the initial news information to obtain the initial topic corresponding to the initial news information and the associated news information corresponding to the initial topic; the sentiment analysis module 202 is configured to perform sentiment analysis on the initial topic according to the associated news information to obtain the target sentiment type corresponding to the initial topic; the impact analysis module 203 is configured to perform an impact analysis on the initial topic according to the target sentiment type and the associated news information to obtain the target topic and the target news information corresponding to the target topic; the event analysis module 204 is configured to perform event analysis on the target news information to obtain the initial event cause and the initial event result corresponding to the target topic; the event clustering module 205 is configured to perform event clustering according to the initial event cause and the initial event result to obtain the target event cause corresponding to the target topic and the target event result corresponding to the target event cause; the file acquisition module 206 is configured to obtain the relevant policy document corresponding to the target topic; the information extraction module 207 is configured to perform key information extraction on the relevant policy document according to the target event cause and the target event result to obtain the target key information corresponding to the target event cause and the target event result in the relevant policy document.
[0113] In some embodiments, the AI-based key information extraction device 200 for policy documents can be applied to a terminal device.
[0114] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described AI-based key information extraction device 200 for policy documents can refer to the corresponding process in the foregoing embodiments of the AI-based key information extraction method for policy documents, and will not be elaborated herein.
[0115] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention.
[0116] As Figure 3 shown, the terminal device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a bus 303, and this bus is, for example, an I2C (Inter-integrated Circuit) bus.
[0117] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and this processor 301 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.
[0118] Specifically, the memory 302 can be a Flash chip, a read-only memory (ROM), a magnetic disk, an optical disc, a USB flash drive, or a mobile hard disk, etc.
[0119] Those skilled in the art can understand that Figure 3 the structure shown in
[0120] Among them, the processor is used to run a computer program stored in the memory, and when executing the computer program, implement any one of the methods for extracting key information of policy documents based on artificial intelligence provided by the embodiments of the present invention.
[0121] In one embodiment, the processor is used to run a computer program stored in the memory, and when executing the computer program, implement the following steps:
[0122] Collect initial news information, and perform topic recognition on the initial news information to obtain the initial topic corresponding to the initial news information and the associated news information corresponding to the initial topic;
[0123] Perform sentiment analysis on the initial topic according to the associated news information to obtain the target sentiment type corresponding to the initial topic;
[0124] Perform influence analysis on the initial topic according to the target sentiment type and the associated news information to obtain the target topic and the target news information corresponding to the target topic;
[0125] Perform event analysis on the target news information to obtain the initial event cause and the initial event result corresponding to the target topic;
[0126] Perform event clustering according to the initial event cause and the initial event result to obtain the target event cause corresponding to the target topic and the target event result corresponding to the target event cause;
[0127] Obtain the relevant policy documents corresponding to the target topic;
[0128] Extract key information from the relevant policy documents according to the target event cause and the target event result, and obtain the target key information corresponding to the target event cause and the target event result in the relevant policy documents.
[0129] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described terminal device can refer to the corresponding process in the embodiment of the method for extracting key information of policy documents based on artificial intelligence described above, and will not be repeated here.
[0130] The embodiments of the present invention also provide a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the methods for extracting key information of policy documents based on artificial intelligence provided in the specification of the embodiments of the present invention.
[0131] Among them, the storage medium may be an internal storage unit of the terminal device described in the foregoing embodiments, such as the hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0132] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiment, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.
[0133] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or system comprising the element.
[0134] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for extracting key information from policy documents based on artificial intelligence, characterized in that: The method comprises: Collecting initial news information, and performing topic identification on the initial news information to obtain an initial topic corresponding to the initial news information and associated news information corresponding to the initial topic; Performing sentiment analysis on the initial topic according to the associated news information to obtain a target sentiment type corresponding to the initial topic; Performing influence analysis on the initial topic according to the target sentiment type and the associated news information to obtain a target topic and target news information corresponding to the target topic; Performing event analysis on the target news information to obtain the initial event cause and initial event result corresponding to the target topic; Perform event clustering according to the initial event cause and the initial event result to obtain a target event cause corresponding to the target topic and a target event result corresponding to the target event cause; Obtain relevant policy documents corresponding to the target topic; Key information is extracted from the relevant policy documents according to the target event cause and the target event result to obtain target key information corresponding to the target event cause and the target event result in the relevant policy documents.
2. The method according to claim 1, characterized in that The step of performing topic identification on the initial news information to obtain an initial topic corresponding to the initial news information and associated news information corresponding to the initial topic includes: Preprocessing the initial news information to obtain initial keywords corresponding to the initial news information, and randomly combining the initial keywords to obtain a plurality of keyword groups, wherein the keyword groups include a first keyword and a second keyword; Obtaining a first text vector corresponding to the first keyword according to a text representation model and obtaining a second text vector corresponding to the second keyword according to the text representation model; Obtaining first frequency information and first spatial information corresponding to the first keyword in the initial news information; Obtaining second frequency information and second spatial information corresponding to the second keyword in the initial news information; Obtaining third frequency information corresponding to when the first keyword and the second keyword appear simultaneously in a sub-sentence of the initial news information; Determine a first degree of association corresponding to the first keyword and the second keyword according to the first frequency information, the second frequency information, and the third frequency information; Determine the text similarity between the first keyword and the second keyword according to the first text vector and the second text vector; Determine a second correlation degree corresponding to the first keyword and the second keyword according to the first spatial information, the second spatial information and the text similarity; The first relevance degree and the second relevance degree are combined to determine the relevance weight corresponding to the keyword group, and the keywords in the keyword group are scored for importance according to the relevance weight to obtain a target score; Filtering the initial keywords according to the target scores to obtain target keywords corresponding to the initial news information; Performing news clustering on the initial news information according to the target keyword to obtain a news clustering result; Performing attention analysis on each sub-clustering result in the news clustering result to obtain an attention change value corresponding to the sub-clustering result; The sub-clustering results are subjected to topic screening according to the attention change value to obtain the initial topic corresponding to the initial news information, and the associated news information corresponding to the initial topic is determined according to the sub-clustering results.
3. The method according to claim 2, characterized in that The performing attention analysis on each sub-clustering result in the news clustering result to obtain an attention change value corresponding to the sub-clustering result includes: Obtaining a first reporting time corresponding to each sub-news data in the sub-clustering result, and determining a first time span corresponding to the sub-clustering result according to the first reporting time; Determine a first target window, and perform data segmentation on the sub-clustering result according to the first target window and the first time span to obtain a first segmentation result corresponding to the sub-clustering result and first associated data corresponding to the first segmentation result; Obtaining a second reporting time corresponding to the first associated data, and determining a second time span corresponding to the first associated data according to the second reporting time; Determine a second target window, and perform data segmentation on the first associated data according to the second target window and the second time span to obtain a second segmentation result corresponding to the first associated data, wherein the second target window is smaller than the first target window; determining the number of windows according to the second target window and the second time span; Determine second associated data corresponding to the second target window from the first associated data according to the second target window combined with the second reporting time; Obtaining a data browsing amount corresponding to the second associated data, and determining an attention representation value corresponding to the first associated data in the second time span according to the second target window and the number of windows combined with the data browsing amount and the second associated data; Determining a focus change value corresponding to the sub-clustering result according to the focus representation value of the first association data; The attention characterization value is obtained according to the following formula: ; in, represents the attention characterization value corresponding to the jth first associated data of the i-th sub-clustering result in the first target window, represents the second target window corresponding to the jth first associated data of the i-th sub-clustering result under the first target window, represents the window number corresponding to the second target window corresponding to the jth first associated data of the i-th sub-clustering result under the first target window, represents the window size corresponding to the second target window corresponding to the jth first associated data of the i-th sub-clustering result under the first target window, represents the number of news of the second associated data corresponding to the jth first associated data of the i-th sub-clustering result in the first target window in the t-th second target window, represents the data views of the second associated data corresponding to the jth first associated data of the i-th sub-clustering result in the first target window in the t-th second target window, represents the data views of the second associated data corresponding to the jth first associated data of the i-th sub-clustering result in the first target window in the k-th second target window, It represents the number of news of the second associated data corresponding to the jth first associated data of the i-th sub-clustering result in the first target window in the k-th second target window.
4. The method according to claim 1, characterized in that: The step of performing influence analysis on the initial topic according to the target sentiment type and the associated news information to obtain a target topic and target news information corresponding to the target topic includes: Determine a preset emotion type, and when the target emotion type is the preset emotion type, perform keyword distribution analysis on the associated news information to obtain associated keywords of the initial topic and initial distribution information corresponding to the associated keywords; Filtering the associated keywords according to the initial distribution information to obtain topic keywords corresponding to the initial topic and target distribution information corresponding to the topic keywords; Obtaining the number of news items containing the topic keyword from the associated news information, and determining the influence representation value corresponding to the initial topic according to the target distribution information and the number of news items combined with the associated news information; The initial topic is screened according to the influence representation value to obtain the target topic, and the target news information corresponding to the target topic is obtained from the related news information according to the target topic.
5. The method according to claim 1, characterized in that The performing event analysis on the target news information to obtain the initial event cause and initial event result corresponding to the target topic includes: Performing event feature extraction on the target news information to obtain a target event and related event information corresponding to the target event; Obtain a first event and a second event from the target event, and obtain first information corresponding to the first event and second information corresponding to the second event from the related event information; Determine a first trigger word corresponding to the first event and a second trigger word corresponding to the second event; Determine a logical association value between the first event and the second event according to the first trigger word and the second trigger word in combination with the first information and the second information; The relationship type between the first event and the second event is determined according to the logical association value, and the initial event cause and the initial event result corresponding to the target topic are determined from the first event and the second event according to the relationship type.
6. The method according to claim 5, characterized in that The determining the logical association value between the first event and the second event according to the first trigger word and the second trigger word in combination with the first information and the second information includes: Determine the corresponding sequence between the first event and the second event under event logic according to the first trigger word and the second trigger word; Determine the target co-occurrence times of the first trigger word and the second trigger word in the target news information according to the sequence, and determine the third frequency information of the first trigger word in the target news information and the fourth frequency information of the second trigger word in the target news information; Determine, according to the target co-occurrence count, the third frequency information, and the fourth frequency information, a third degree of association corresponding to the first event and the second event under the first trigger word and the second trigger word; determining a first mutual correlation value between the first event and the second event according to the first trigger word and the second information, and determining a second mutual correlation value between the first event and the second event according to the second trigger word and the first information; fusing the first correlation value and the second correlation value to determine a fourth correlation degree between the first event and the second event; Performing a correlation analysis based on the first information and the second information to obtain a fifth correlation between the first event and the second event; The third degree of association, the fourth degree of association, and the fifth degree of association are integrated to determine the logical association value corresponding to the first event and the second event in the sequence.
7. The method according to claim 1, characterized in that The extracting key information from the relevant policy document according to the target event cause and the target event result to obtain the target key information corresponding to the target event cause and the target event result in the relevant policy document includes: Decomposing the relevant policy document to obtain a plurality of initial sentence information, and generating text based on the initial sentence information and the target event cause according to a first splicing rule to obtain a first generated text; Performing text generation on the initial sentence information and the target event result according to a second splicing rule to obtain a second generated text; Using a causal relationship recognition model to perform relationship classification on the first generated text to obtain a first relationship type corresponding to the first generated text; Using the causal relationship recognition model to perform relationship classification on the second generated text to obtain a second relationship type corresponding to the second generated text; Perform sentence screening on the initial sentence information according to the first relationship type and the second relationship type to obtain intermediate sentence information; Performing sentence criticality scoring on the intermediate sentence information to obtain a criticality score corresponding to the intermediate sentence information; The target key information corresponding to the target event cause and the target event result in the relevant policy document is obtained from the intermediate statement information according to the criticality score.
8. An artificial intelligence-based policy document key information extraction device, characterized in that: include: A topic identification module is used to collect initial news information and perform topic identification on the initial news information to obtain an initial topic corresponding to the initial news information and related news information corresponding to the initial topic; A sentiment analysis module, configured to perform sentiment analysis on the initial topic according to the associated news information to obtain a target sentiment type corresponding to the initial topic; An influence analysis module, configured to perform influence analysis on the initial topic according to the target sentiment type and the associated news information to obtain a target topic and target news information corresponding to the target topic; An event analysis module, used to perform event analysis on the target news information to obtain the initial event cause and initial event result corresponding to the target topic; An event clustering module, used for performing event clustering according to the initial event cause and the initial event result to obtain a target event cause corresponding to the target topic and a target event result corresponding to the target event cause; A file acquisition module, used to obtain relevant policy documents corresponding to the target topic; The information extraction module is used to extract key information from the relevant policy documents according to the target event cause and the target event result, and obtain the target key information corresponding to the target event cause and the target event result in the relevant policy documents.
9. A terminal device, characterized in that: The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and implement the artificial intelligence-based policy document key information extraction method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer storage medium for computer storage, characterized in that: The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the artificial intelligence-based policy document key information extraction method described in any one of claims 1 to 7.